# telltale — design doc Status: v1 is cut — **v0.2.0, released 2026-08-14**, the snapshot gates held (§1). The honest-gauge rule requires every segment's data source to be named here before that segment ships; the tables below are the authority the eval harness tests against. ## Citing this document Code comments, PR bodies and the other docs cite this document by section number. Every numbered section therefore carries a stable HTML anchor, so a citation can be a link. **The anchor comes from the section number alone.** Write `s`, then the number, and replace each `.` with `-`. So §9 is `#s9`, §4a.1 is `#s4a-1`, and §7.16a is `#s7-16a`. ``` [§7.16a](docs/design.md#s7-16a) from another file [§7.16a](#s7-16a) from inside this document ``` The anchor is keyed on the number and never on the heading text. GitHub derives its own anchor from the heading text, and the headings here carry dated prose that gets amended, so a text-derived anchor breaks on every rewording. `#s7-16a` survives it. Add the matching anchor line when you add a numbered section. ## ADR index The decision records themselves are archived outside this repository, because the 2026-08-05 ruling stopped adding them here. The ADR numbers stay in the prose, so this table says what each one decided and points at the section that carries the substance. Each row is also an anchor target: ADR-002 is `#adr-002`. | ADR | What it decided | Where the substance is | |---|---|---| | ADR-001 | The honest-gauge rule. A displayed value must come from measured vendor output, and an inferred value is omitted or marked as an estimate. | [§4a.1](#s4a-1), [§3.4](#s3-4) | | ADR-002 | Go with Bubble Tea and Lipgloss, one binary, two modes. Windows is the primary target, and the statusline path initializes no TUI framework. | [§6](#s6), [§5](#s5) | | ADR-003 | Gemini CLI gets a built-in adapter. Its verification hold is released, and an addendum records the consumer-tier withdrawal. | [§3.7](#s3-7) | | ADR-004 | Antigravity CLI gets a statusline, routed on the payload's documented `product` field. | [§2.1](#s2-1) | | ADR-005 | External adoption is an explicit product goal, so the roadmap carries an adoption track. | [§8](#s8) | | ADR-006 | Antigravity also gets a HUD adapter. A re-survey found the transcript that the first verdict recorded as absent. | [§3.8](#s3-8) | | ADR-007 | Cursor gets a built-in HUD adapter, because its seam is on disk rather than behind a CLI. | [§3.9](#s3-9) | | ADR-008 | `telltale council` is the dispatch room. It spawns vendor CLIs, so it sits off the gauge data layer. | [§9](#s9) | ADR-006 is the one number above that the current text never prints; §3.8 carries its account under the re-survey heading instead. **ADR-010 and ADR-012 in this repository are not telltale's.** They belong to the `agent-ops` decision series, and the prose always names that series beside them. The two series number independently, so telltale's ADR-007 is the Cursor HUD adapter while `agent-ops` ADR-007 is a different decision. ## 1. Product shape **`telltale council` is the product. The gauges — `telltale statusline` and `telltale hud` — are the infrastructure under it.** That ranking had been stated out loud and recorded nowhere, which meant it bound nothing and every argument about what to build next started over from memory. It is written here so it stops depending on who was in the room. It does not demote the gauges: they are where each vendor's on-disk seam was surveyed and written down (§3), they are what the honest-gauge rule was built and tested against (§5), and council inherits both — it renders through the same `internal/model` vocabulary and `internal/theme` palette the two gauge paths share. They are finished, they are load-bearing, and they are not the thing this is for. **v1 is a snapshot, not a freeze — it cuts when three gates hold, not when the room goes quiet (re-cut 2026-08-08, owner's ruling).** The original hold said "until council settles," and the first attempt to operationalize that — five consecutive days with no merged PR touching council's visible surface — was falsified within a day: this project is driven daily by its owner, so a quietness clock measures abandonment, not stability. What the hold was actually protecting is narrower and checkable: a stranger who reads the launch post and installs must find the room the post described. So v1 cuts when, on the day of the tag: 1. **Nothing on the surface is half-finished or landed-but-never-driven** — every recently re-founded piece (a seat's protocol, a new command) has been used by the owner for real work for a few days; 2. **The README is verified against tip** — every claim, keybinding and badge checked against the code, with the frame-freshness tests holding the renders; 3. **No breaking change to the routing grammar, room commands or keymap is planned** — churn after the tag is welcome; a *known upcoming* contract break is not. The owner's dogfood bar (two weeks of daily use, clock from 2026-08-01) still applies and closes no earlier than 2026-08-15. Development never pauses for any of this: work merged after the tag becomes the next minor version, and the tag itself is one command (§8). The standing alternative — cut v1 as gauges only, statusline and HUD with declared vendor version pins — remains rejected, because a v1 that named the gauges would name the wrong product. Two gauge surfaces over one data layer: ``` vendor adapters ──► normalized session model ──► renderers (claude, codex) (one schema, documented) (statusline / HUD) ``` One Go module, one binary (`telltale.exe`). ADR-002 specified two modes; council (ADR-008) is the third, and it does not sit on the pipeline above: - **`telltale statusline`** (Claude Code and, since ADR-004, Antigravity CLI — routed on the payload's documented `product` field, §2.1): reads the vendor's JSON on stdin, prints one line, exits. **Bubble Tea is never initialized on this path** — a convention until 2026-08-16, gated since by `TestFastPathNeverReachesTUIFramework` (§5). Latency budget: single-digit milliseconds of telltale's OWN work, and that much is measured — parse+render is 14 µs (`BenchmarkRender`). The end-to-end cost the operator actually pays is larger and is process start, not this code: ~25 ms median per invocation on the reference workstation. Budget-conscious output (every character renders on every prompt). - **`telltale hud`** (cross-vendor): a Bubble Tea/Lipgloss watch-mode TUI listing live sessions across vendors with per-session gauges. **First-class UI surface** — a UI design section (layout grid, color/threshold system, motion rules, empty/degraded state designs) is written here BEFORE the HUD is built, and degraded-state renders are eval fixtures. Windows Terminal is the reference rendering environment. - **`telltale council`** (ADR-008, §9): the dispatch room — one brief typed once, answered by the seated vendor CLIs side by side, each column claiming only what was measured about that vendor. It spawns vendor CLIs instead of reading their session files, which is why it is off the data layer above and specified separately in §9. - **`telltale hook `** (§7.16): the vendor-hook relay — a per-turn payload on stdin, token counts to `~/.telltale/usage/`, and **nothing on stdout**, because a hook's stdout is parsed by the vendor as a hook result. Not a gauge and not a room; it renders nothing and is never run by a human. **The gauges never write, with three bounded exceptions — all under `~/.telltale/`, all numbers and keys only, never content.** `telltale council` keeps `council/room.json` (the session ids reattaching needs); the statusline relays `quota/.json` after its line is on stdout (§7.15); and `usage/.json` accumulates per-turn token counts from two writers, `telltale hook` (§7.16) and the `telltale otel` collector (§7.16a). Each store is atomic (temp+rename), best-effort, self-expiring on read, and pinned by a test that walks the serialized form field by field. No transcript, prompt, reply, path or address reaches any of the three. Anything else that wants to write from `internal/hud` or `internal/statusline` is in the wrong package. **Amended 2026-08-11: a FOURTH store exists, it carries content, and the paragraph above was never corrected for it.** The event sink (`telltale events`, §7.21) writes each hook payload VERBATIM under `~/.telltale/events/`. It does not widen the rule above, because what contains it is scope rather than redaction: it is its own foreground mode the operator starts, its server binds loopback only and refuses any other host, and nothing in the gauges reads or renders those files. So the three counted above stay numbers-and-keys, and the fourth is named as an exception instead of being folded into them. §7.21 carries the record and CLAUDE.md's boundary section carries the same exception. **Amended 2026-09-04: the numbers-and-keys list is FOUR, not three.** `telltale probe` (LEDGER, 2026-09-04) writes `probe/.json` per seat: the vendor id, the version string that binary printed, the day, the telltale build that probed, and one result plus a millisecond count for each of its three checks. It is the strictest of the four, because its writer DRIVES an agent: the brief, the reply, the session id and the workspace path are all within reach of it and none is written, and neither is the failure reason, because a vendor's own error line carries paths. It is not a gauge, so the paragraph above is unchanged in what it says about `internal/hud` and `internal/statusline`. SECURITY.md and CLAUDE.md carry the same four; a count in prose goes stale with nothing to catch it, so `internal/history/boundary_test.go` enumerates the writer packages by name and states none. The two gauge paths share exactly two packages: `internal/model` (the schema) and `internal/theme` (thresholds and value formatters). `internal/theme` is stdlib-only and holds no style type, which is what lets both surfaces share the numbers while the statusline links no TUI framework. ## 2. Statusline segments (v1) | Segment | Source (exact field) | Empty/degraded state | Status | |---|---|---|---| | Model | stdin `model.display_name` (falls back to `model.id`) | hide if both empty | **built** | | Context % | stdin `context_window.used_percentage` (input-token based per docs) | hide segment | **built** | | Cache hit % | stdin `prompt_cache.hit_ratio` (vendor-COMPUTED over the session's main-conversation requests; ×100 is a unit conversion), gated on `prompt_cache.caching_observed` | absent block, null ratio, or `caching_observed` false all hide; 0 with caching observed is a reading and renders `cache 0%` | **built** (§7.16c) | | Session cost | stdin `cost.total_cost_usd` | hide segment | **built** | | Quota pacing (5h) | stdin `rate_limits.five_hour.used_percentage` + `resets_at` (unix s) | rate_limits absent on API-key logins; each window independently absent → hide, never zero; countdown hides without `resets_at` | **built** | | Quota pacing (7d) | stdin `rate_limits.seven_day.*` | same rule | **built** | | Worktree | stdin `worktree.name` (present only in `--worktree` sessions) | hide segment | **built** | | Folder | stdin `workspace.current_dir` (fallback `cwd`), basename only — no filesystem/git calls | hide segment | **built** | Deliberately not shown **on this path**: git branch (would require an exec; the statusline path reads nothing beyond stdin — revisit only with a measured budget), permission mode (not in the stdin payload; same call the predecessor script made). The path's one write is the quota relay (§7.15): after the line is on stdout, the payload's rate-limit windows go to `~/.telltale/quota/` for the HUD — numbers only, best-effort, never ahead of the render. > Both exclusions are properties of the **stdin seam**, not of Claude Code. The > transcript carries `gitBranch` and `permissionMode` directly, so the HUD's disk path > gets both for free (§3.1). The statusline's exclusions are unaffected. Threshold colors (applies to any percentage segment): green < 60, yellow ≥ 60, red ≥ 85, from `theme.WarnPct` / `theme.CritPct`. `NO_COLOR` env strips styling. Derived displays (reset countdown `↻2h13m`) are arithmetic on `resets_at` only. Schema verification record: full stdin JSON schema captured from code.claude.com/docs/en/statusline on 2026-08-01, including per-field absence semantics (`rate_limits` Pro/Max-only and only after first API response; each window independently absent). Statusline updates are debounced at 300ms and in-flight scripts are cancelled — which is the empirical backing for the fast-exit budget. Parsing ignores unknown fields by design (vendor adds fields between versions). **Re-measured 2026-08-16 at CLI 2.1.233 by source read (§7.16b), and the payload had grown.** `context_window` now also carries `total_input_tokens`, `total_output_tokens` and a `current_usage` object, and the payload carries an undocumented `prompt_id`. The first three are modelled and **render nothing and relay nothing** — they describe one API call, and `total_input_tokens` is the window's occupancy rather than a spend, so no total may be built from them. That is the whole of §7.16b, and the tolerance of unknown fields above is why every release between 2.1.90 and 2.1.233 was unaffected by the growth. The segment table above is unchanged: no new segment was added. **Re-measured 2026-09-04 at CLI 2.1.260 by source read (§7.16c), and this time the growth DID earn a segment.** The payload gained `prompt_cache` at 2.1.251 — the vendor's own prompt-cache statistics for the main conversation — and one of its fields is a number the vendor computes rather than a count telltale would have to combine. The segment table above carries the new row; §7.16c carries the measurement and the reason the transcript adapter was left alone. **Known divergence from `theme`, deliberate and unresolved:** the statusline's local `pct` uses `%.1f`, which rounds, and its local `shortDur` has no days branch. The shared helpers `theme.Percent` and `theme.Countdown` floor and carry days respectively, and the HUD uses them. Unifying the statusline onto them changes its rendered output (99.96 would stop reading as `100.0%`, a 7-day window would stop reading as `↻120h00m`), which is a behaviour change to a shipped surface and is therefore a separate change with its own fixture updates. The thresholds are already unified; the formatters are not. ### 2.1 Antigravity CLI statusline (added 2026-08-02, ADR-004) `telltale statusline` serves a second vendor: Antigravity CLI (`agy`) hands statusline commands a JSON payload on stdin, same seam shape as Claude Code. **Routing is the documented `product` field** — agy stamps `"product": "antigravity"` on every payload (observed on all six live captures) and Claude's payload has no product field; one binary, one subcommand, no flag. Stdin is read once and handed to whichever parser the marker selects (`internal/antigravity`). | Segment | Source (exact field) | Empty/degraded state | Status | |---|---|---|---| | Model | stdin `model.display_name` (falls back to `model.id`) | hide if both empty | **built** | | Context % | stdin `context_window.used_percentage` (vendor-reported; the payload also carries `context_window_size`) | hide segment; 0 is a reading and renders `ctx 0%` | **built** | | Quota buckets | stdin `quota..remaining_fraction` + `reset_in_seconds` (fallback `reset_time`) — one segment per NAMED bucket, ids rendered verbatim, sorted for stability; used% = (1−remaining)×100, a unit conversion | bucket without `remaining_fraction` hides; absent map hides all | **built** | | Agent state | stdin `agent_state` — the first vendor-REPORTED liveness signal on any seam; `tool_confirmation_pending: true` outranks it and renders `confirm?` | hide if empty; unknown vocabulary renders verbatim in dim | **built** | | Branch | stdin `vcs.branch` (+`*` when `vcs.dirty`) — in the payload, so no exec; the no-I/O-beyond-stdin rule holds | hide segment | **built** (documented; not yet observed live — §3.8) | | Folder | stdin `workspace.current_dir` (fallback `cwd`), basename only | hide segment | **built** | Not rendered, deliberately: `cost` does not exist anywhere in the payload (nothing is priced); `email` and `plan_tier` are identity, not gauges; `transcript_path` is not displayed because it is unverified, not because it is absent — the transcript IS written on disk (§3.8's 1.1.13 re-read found it in 81 of 81 conversations; the claim that agy "never writes that file" was false and is corrected there), but whether the PAYLOAD's value points at that real file needs a live capture nobody has run, and displaying an unverified path would be narrating. **Amended 2026-08-17 — `transcript_path` is no longer unverified. It is verified WRONG.** The live capture that paragraph asks for was taken (§3.8's re-capture block). The payload's path drops the `-cli` segment: it names `~\.gemini\antigravity\brain\\…`, while the real transcript for that same session sits only under `~\.gemini\antigravity-cli\brain\\…`. The advertised directory exists and is EMPTY, so a reader that trusted the value would open nothing and could not tell a missing file from a missing session. The refusal above stands unchanged and its GROUND changes: it was caution about an unverified value, and it is now a measurement of a false one. No code ever followed the path, which was re-checked across the repository rather than assumed. Schema verification record: documented contract (antigravity.google/docs/cli/statusline) cross-checked against a six-payload live capture from a real interactive session on agy 1.1.9, 2026-08-02 (§3.8). **Re-captured 2026-08-17 at agy 1.1.13 — fifteen payloads, one live interactive turn** (§3.8's re-capture block): quota confirmed at FOUR named buckets (`3p-5h`, `3p-weekly`, `gemini-5h`, `gemini-weekly`), each carrying `remaining_fraction`, `reset_time` and `reset_in_seconds`; `agent_state` observed live, including one value the documented vocabulary omits; `context_window` grown to the §7.16b shape. `vcs` is still the one documented segment nobody has observed — both capture sessions ran outside a git repo. Fixtures are synthesized to the observed shapes. ### 2.2 Cursor CLI statusline (added 2026-08-16) `telltale statusline` serves a third vendor. `cursor-agent` reads a top-level `statusLine` object from `~/.cursor/cli-config.json` and hands the command a JSON payload on stdin — the same seam shape as the other two. Measured at **cursor-agent 2026.08.04-aaa8809** on Windows 11, live capture first and source read second; §7.16's dated amendment carries the per-surface measurement and `internal/cursorstatus`'s package doc carries the shapes. ```json {"statusLine":{"type":"command","command":"telltale statusline --vendor cursor", "padding":0,"updateIntervalMs":300,"timeoutMs":2000}} ``` **Routing is an explicit flag, and that is a finding rather than a shortcut.** This payload carries no `product` field, no `hook_event_name`, and no other vendor name — `version` holds the CLI's own build string, which is a value and not a marker. It is Claude-shaped on purpose: the vendor's bundled `statusline` skill says the spec "is aligned with Claude Code's status line", and the two overlap on `session_id`, `transcript_path`, `cwd`, `model.*`, `workspace.*`, `version` and `output_style.name`. So §2.1's affirmative-marker scheme cannot be extended to it, and guessing from structure (`render_width_chars` present, `cost` absent) would be a heuristic over a payload the vendor may grow at any release — with a silent failure mode, since a misrouted Claude payload renders a plausible line with its quota missing. `--vendor cursor` is written once into the config above and wins over the marker probe. | Segment | Source (exact field) | Empty/degraded state | Status | |---|---|---|---| | Model | stdin `model.display_name` (falls back to `model.id`) | hide if both empty | **built** | | Context % | stdin `context_window.used_percentage` — the ONE context number this vendor sources rather than computes | hide segment; a session before its first API call sends every `context_window` key as `null`, which is an unread field and not a zero | **built** | | Autorun | stdin `autorun` — no counterpart in either other vendor's payload; renders only when `true` | `false` and absent both hide (see below) | **built** | | Worktree | stdin `worktree.name`, same `⌥` mark as the Claude path | hide segment | **built** (documented; not observed live) | | Folder | stdin `workspace.current_dir` (fallback `cwd`), basename only | hide segment | **built** | **Two fields are REFUSED, and they are the reason this section is careful.** `context_window` arrives with six keys; two of them are computed by the CLI from a third and named as though they were read. From the bundle's own payload builder in `./src/ui.tsx` at the pinned build, where `Ve` is the usage reading and `Ke` the window size: ``` m = null!=Ve?Ve:null, // used_percentage f = null!=m ? Math.max(0,Math.round(10*(100-m))/10) : null, // remaining_percentage v = null!=Ke&&null!=m ? Math.round(m/100*Ke) : null // total_input_tokens ``` The vendor documents it too — its `statusline` skill calls `total_input_tokens` "Estimated input tokens (derived from used_percentage)". A token count that is really a rounded percentage, under a name that reads like a meter, is the ADR-001 violation print mode's `inputTokens` already cost this repo once (§7.16). **They are absent from `internal/cursorstatus`'s structs rather than parsed-and-ignored**, on the `internal/cursorhook` rule: the struct is the allowlist, `encoding/json` drops every field with no destination, and a field that does not exist cannot be reached by a later change that did not read the comment. `TestCursorDerivedFieldsNeverRender` feeds a fixture carrying both, populated and plausible, and asserts neither number reaches the line. Also not rendered: `context_window_size` (genuinely vendor-reported, but nothing draws it — add it back with the segment that wants it) and `current_usage` (observed only as `null`, so its populated shape is unmeasured and declaring one would be inventing a schema). There is **no quota, no rate limit and no cost anywhere in this payload**, so this path writes no quota relay at all (§7.15) — the absence is the honest answer, not a gap. **The autorun asymmetry is deliberate and is not the zero-vs-absent rule being bent.** That rule governs gauge READINGS, where 0% and "no source" are two facts a user must tell apart. `autorun` is a posture flag with a default: `false` is the ordinary state of every session, and a segment that says "nothing unusual" on every line teaches the reader to stop looking. Off and absent both render nothing; on renders a word, in yellow, because it is the state where the agent may run a command without asking. **Budget.** The vendor's `timeoutMs` defaults to 2000 with a floor of 50, and `updateIntervalMs` is clamped to >= 300. The 2000ms is its kill deadline, not an allowance: the binary is respawned on every debounced update, so ADR-002's single-digit-millisecond target for telltale's own work is unchanged, and `BenchmarkRenderCursor` measures it at 14 µs against that 300 ms floor. Two honest limits on that sentence, added 2026-08-16: a benchmark is not a gate — CI runs no `-bench`, so `BenchmarkRenderCursor` fails nothing — and parse+render is not what the respawn costs. The respawn's end-to-end price is ~25 ms median, essentially all of it process start. §5's amendment says which half CI now holds. Schema verification record: two live payloads captured 2026-08-16 from a real interactive session (the shape with every `context_window` key null — that session had made no API call), cross-checked against the vendor's bundled `statusline` skill and a source read of `./src/hooks/use-status-line.ts` and `./src/ui.tsx` at 2026.08.04-aaa8809. Fixtures are synthesized to the observed shapes; the populated-context fixture is synthesized to the documented one, because no captured payload ever carried a number there. **That gap is CLOSED, 2026-08-17**: a live interactive session at cursor-agent `2026.08.11-e8db854` rendered `ctx 12.7%` after its first reply, so `used_percentage` is observed populated at one-decimal precision. §7.16's amendment carries the capture and the build caveat. The fixture's assumed shape was correct and did not move. ## 3. HUD (v1) One row per live session, both vendors; per-row: vendor, session identity, model, context/quota gauges **where the vendor provides them**, last-activity age. The rendered grid, its responsive tiers and every degraded state are specified in §7. ### 3.1 Claude Code adapter sources — VERIFIED LIVE 2026-08-01, Claude Code 2.1.219; RE-MEASURED 2026-08-16, Claude Code 2.1.233 Read-only survey of `%USERPROFILE%\.claude\` on the dev PC: 33 project dirs, 837 sessions, 13,211 records walked. Nothing from that survey is reproduced here or in the fixtures; the fixtures are synthesized to shape only. **Discovery glob — `~/.claude/projects/*/*.jsonl`, non-recursive, UUID basename.** Measured: non-recursive = 837 files, recursive = 2021. The extra 1,184 are subagent transcripts under `/subagents/**`, plus `tool-results/` and `workflows/` sidecars; a `**/*.jsonl` glob inflates the session list 2.4x and double-counts every token. Non-`.jsonl` neighbours share the directory (`.memory-sync-manifest.json`), so the basename is checked as a UUID, not just the extension. The project-directory slug is **lossy and must never be decoded to a path**: `\` and a literal `-` both encode as `-` (`C--Users-dev-code-my-app` could be `code\my-app` or `code-my\app`), and the drive-letter case is not stable — the same tree can hold `C--Users-…-app` and `c--Users-…-app` as siblings. `cwd` is read from the record; the slug is an opaque grouping key. Because those sibling directories can hold the same session id, `Discover` also de-duplicates by id, newest mtime winning: a duplicate id would break the HUD's row matching. The tree also mutates during a sweep (a project dir vanished between enumeration and open during the survey), so `Discover` swallows ENOENT on dirs and files and continues rather than aborting. **Record fields the adapter reads** (all on `assistant` records unless noted): | Normalized field | Exact source | Absent when | |---|---|---| | Session id | `sessionId` (present on every record type) | never | | Working dir | `cwd` | metadata-only record types | | Git branch | `gitBranch` (carried as a display-only extra) | outside a repo | | Model | `message.model` | non-`assistant` records; `""` is rejected | | Tokens in context | `message.usage.input_tokens + .cache_read_input_tokens + .cache_creation_input_tokens` | no `usage` | | Last activity | file mtime | never (but see clock skew below) | | CLI version | `version` (display-only extra) | never | | Title | `custom-title.customTitle`, else an `ai-title` record | untitled sessions | | **Sub-agent count** (v1.1) | **stat only:** entries matching `.jsonl` under `/subagents/`, mtime within 15 min | never — an absent directory is a measured **zero**; only an unreadable one is absent | **The sub-agent count is `CapDerived`, and the reason is worth stating precisely.** The files are counted *exactly* — one `ReadDir`, one `Info` per entry, no file opened and no byte parsed, which is what makes it affordable on the 1 s poll. What is inferred is the 15-minute recency boundary that turns "written lately" into "a fan-out is running now". That inference is the thing the estimate marker exists to expose, so the chip renders `⑂~2` rather than `⑂2` (§7.13). The boundary is `model.DefaultLivenessThresholds.Idle` rather than a second constant: the chip sits on a row whose state dot already classifies "recent" at that boundary, and two definitions of recent on one line is how a display starts contradicting itself. Two absences are distinguished, per §4a.1: the directory **not existing** means the session never fanned out, which is a countable zero; the directory existing and the OS **refusing** is nil plus a diagnostic, because we do not know. A sub-agent transcript whose mtime is ahead of the local clock is not counted at all — the same rule the session's own mtime gets, for the same reason: a timestamp ahead of the clock is not a readable time, so it cannot be evidence of recency. **Amended 2026-08-17 — `subagentStatusLine` was evaluated as a replacement for that inference, and refused.** The estimate marker exists because of the 15-minute boundary, so a vendor surface that *reports* which sub-agents run would retire `CapDerived` honestly. Claude Code 2.1.233 ships one. It is a settings key with the same `{type:"command", command}` shape as `statusLine`, and the bundle's own schema describes it as a "Custom per-subagent status line shown in the agent panel; receives row context as JSON on stdin". Measured by a source read of the shipped bundle at **Claude Code 2.1.233** — the instrument §7.16b used, labelled here for the same reason. The payload is the shared session-basics block (the `py` helper §7.16b identified) plus `columns` and a `tasks` array. Each `tasks` entry carries `id`, `name`, `type`, `status`, `description`, `label`, `startTime`, `model` and `effort`. The command runs with a 5 s timeout, and its stdout is parsed as JSON lines of `{id, content}`. A React effect drives it: a 300 ms debounce on change, then a repeating 5 s tick while any row is live, with an overlap guard. **The count it carries is not the count this adapter defines, and that is the refusal.** Three measured reasons, in order of weight: - **The population is wider.** `tasks` holds panel rows, and the observed `type` values include `local_agent`, `remote_agent`, `in_process_teammate` and `local_workflow`. §3.1 counts transcripts in a session's `subagents/` sidecar. `len(tasks)` answers a different question under this field's name. - **The list keeps rows that finished.** The tick filters on `evictAfter !== 0`, so a row survives its own completion until eviction. A length taken from it overstates a fan-out in progress. The 15-minute boundary makes the same error, but it declares itself an estimate. - **It is the interactive UI's path, and the HUD reads disk.** The effect is a React hook over the agent panel and the terminal's column count, and print mode mounts no panel. To source it, telltale needs a new relay mode plus operator wiring in `settings.json`. It would then cover only the sessions where both are true, while `countSubagents` covers every session on disk today. So `FieldSubagents` stays `CapDerived` and `countSubagents` is unchanged. The payload's `status` field is the one genuinely stronger signal here, because a reported status beats an inferred recency window. It is reachable only through that relay, so this ruling is re-openable against that build rather than closed. **One arm is owed, and it is named rather than dropped.** A live payload capture was attempted and blocked. To exercise `subagentStatusLine`, the key must sit in a `settings.json`, and this machine's credential guard default-denies writes to that path — including the throwaway, project-local copy the probe wanted. The block was accepted rather than worked around, so the shape above rests on a source read with no live capture behind it. That is weaker than §7.16b ended up: it closed its own source read with a capture on the same day. Only `assistant` and `user` records carry `message`. `custom-title`, `last-prompt`, `mode` and `ai-title` carry `{type, sessionId, }` and have **no `timestamp` and no `cwd`** — a parser that assumes those fields exist will nil-deref. Full observed `type` set: `assistant, user, attachment, last-prompt, queue-operation, custom-title, pr-link, mode, system, ai-title, permission-mode, file-history-snapshot`. The survey verified the **shape** of an `ai-title` record but not the name of its payload key. The adapter therefore matches on the verified structure — exactly one key beyond `type` and `sessionId`, holding a string — rather than guessing a field name from memory. A record that does not match yields no title and the row falls back to its workspace name, which is absence rather than a wrong label. **Claude capability gaps (grepped, zero matches across the corpus):** - **No cost.** `cost.total_cost_usd` exists only on the statusline stdin payload. - **No quota.** `rate_limits.*` likewise stdin-only. - **No context window size**, so **context % is not derivable** — the denominator varies by model and by the `[1m]` variant. Token counts are sourced and carried as a display-only extra; the CONTEXT cell for a Claude row is absent (§6 Q7). **Honest-gauge traps pinned by fixtures:** - `input_tokens` alone is not context usage. Measured live: `input_tokens=2, cache_read=213388, cache_creation=2464`. Reading `input_tokens` renders **2 tokens** for a ~216k-token context. - `message.model` can be `""` (locally generated notices, zeroed usage). It must never reach the model cell and its zeros must not overwrite a real reading. - An mtime ahead of the local clock has **no readable age**. It is left nil and marked degraded rather than clamped, because "0s" claims the session was active this instant. The HUD renders `—`. **Liveness — and why the PID registry is not read in v1.** The honest primitive is **mtime = last activity**, rendered as an age. The survey found an undocumented registry at `~/.claude/sessions/.json` (`{pid, sessionId, cwd, startedAt (unix ms), version, kind, entrypoint, name, nameSource}`); at survey time all 7 entries mapped to live PIDs and to top-level transcripts. **The adapter does not read it.** Every use of it reduces to "a process with this id exists", which §4a.4 names explicitly as evidence a process exists rather than evidence the session is doing anything — and a liveness hint is the one value `Validate` cannot check, so the bar for emitting one is a signal that actually separates working-now from process-exists. `liveness` is therefore `CapNone` for Claude and the HUD classifies every vendor identically from `last_activity`. The registry stays recorded here as a verified observation on 2.1.219 (not a vendor contract) in case a later version adds a turn-start/turn-end signal worth reading. **Read strategy.** Transcripts routinely reach 7.7 MB. The adapter reads a bounded **head** (64 KiB — session id / cwd / git branch / title, verified present on the first record of 60/60 files sampled) plus a bounded **tail** (256 KiB, first fragment discarded as partial), scanning for the newest `assistant` record with `message.usage` and a non-synthetic model. A file smaller than the tail window is read once, not twice. Records with `isSidechain == true` are skipped defensively — 0 of 837 top-level transcripts contain one on 2.1.219, but the filter is free. **RE-MEASURED 2026-08-16 — Claude Code 2.1.233.** `telltale doctor`'s drift notice (§9.42) reported this survey stale on its first live run, and this is the re-survey it asked for. Same machine, same read-only method: **119 project directories, 1,045 top-level transcripts, 179,614 records walked**, up from 33 / 837 / 13,211. The pin moves to `Claude Code 2.1.233`; §3.10's cell inherits it from the adapter's own constant. *The corpus is mixed-version, and that is a method change, not a footnote.* **16 CLI builds wrote these records**, and 2.1.233 wrote only 1,704 of them. So "surveyed at 2.1.233" cannot mean "this is what 2.1.233 writes" — every claim below about *when* a field arrived is attributed by the record's own `version` field, not by the version of the binary installed. Two records types carry no `version` at all and cannot be attributed either way; the block says so where that bites. *A raw grep is no longer a safe method here, and the first survey's headline finding was wrong because of it.* This corpus now contains telltale's own development sessions, which **discuss these field names in their own text** — so a grep for `context_window_size` matches prose about the absence of `context_window_size`. The re-measure therefore parsed every record and looked for each token as a **JSON key at any nesting depth**. Raw-token hits: `rate_limits` 420 records, `total_cost_usd` 288, `context_window_size` 87. Hits as an actual key: **one**. **What held.** | claim | 2026-08-01, 2.1.219 | 2026-08-16, 2.1.233 | |---|---|---| | first record carries `sessionId` | 60 of 60 sampled | **1,045 of 1,045** | | `isSidechain` in top-level transcripts | 0 of 837 | **0 of 1,045** | | recursive glob inflates the session list | 2021 vs 837 (2.4x) | **2,001 vs 1,045 (1.9x)** | | `message.model` can be `""` | observed | **34 records** — the trap holds | | `custom-title` payload key | `customTitle` | **`customTitle`, 6,327 records** | | sessions with a `subagents/` sidecar | present | **107 of 1,045** | The `ai-title` payload key is no longer unverified. §3.1 above says the survey established the record's *shape* but not its key name, and that the adapter therefore matches on structure. The re-measure names it: **`aiTitle`**, exactly one key beyond `type` and `sessionId`, on 2,331 records. **The structural matcher stays as it is** — it was correct, it is now confirmed correct, and hard-coding the name buys nothing a measured structure does not already give. **What changed, and what it costs.** Every capability gap stays `CapNone`. One of them keeps the ruling and loses its stated reason: - **`quota` — the reason was wrong, the ruling was right.** The 2026-08-01 pass grepped the snake_case `rate_limits` and recorded zero matches. **The on-disk key is camelCase `rateLimits`**, it hangs off `error` on API-error records, and **2.1.219 itself wrote it** — the original grep missed a key that was already there, and this survey's own spelling hid it for two weeks. It is still not a quota source, for a better reason than absence: it was **`null` in 32 of 32 records**, and it only appears where a request FAILED, never on a normal turn. A key that is present and null is not a reading (§4a.1). - **`context_pct` — unchanged and re-confirmed.** No `context_window_size`, `context_window` or `contextWindow` key occurs at any depth. `message.context_management` exists (36 records) and is `null` in 34 of them; the other two carry an empty `applied_edits` array. No denominator. - **`cost` — unchanged.** No `cost` or `total_cost_usd` key at any depth. Still stdin-only. - **`liveness` — unchanged, and the registry was re-opened on purpose.** §3.1 recorded `~/.claude/sessions/.json` explicitly so a later build could be checked for a turn-start or turn-end signal. 2.1.233 adds two keys, **`peerProtocol` and `procStart`**. Neither is that signal. `procStart` hardens process *identity* — a pid plus its start time survives PID reuse, which a bare pid does not — but it still answers only that a process exists, which §4a.4 rules out. `CapNone` stands. **Fields that appeared, and are deliberately modelled by nothing.** Recording an arrival is not the same as reading it; per §7.16b, model-and-render-nothing needs a reason, and absence of need is itself a finding. - **`message.usage.output_tokens_details.thinking_tokens`** — the one genuinely new field. Written only by 2.1.228, 2.1.229 and 2.1.233 (3,462 records), zero at 2.1.219. It breaks down **output** tokens, and this adapter's token figure counts what entered **context**, so it feeds no cell. Not modelled. - **`message.usage.speed`, `.inference_geo`, `.server_tool_use`, `.iterations`** — all present at 2.1.219 as well, so not drift at all. None carries a window size. - **a top-level snake_case `session_id`** on some `assistant`, `user` and `attachment` records (680), beside the camelCase `sessionId`. Those records carry both. The adapter reads `sessionId`. - **a `file-history-delta` record type** (9 records), which the observed type set above does not list. **One claim above is narrowed, and it is the canary's.** §3.10 called `sessionId` the field on every JSONL record. At 2.1.233 it is not: **`file-history-snapshot` (37) and `file-history-delta` (9) carry no `sessionId` at all** — 46 of 179,614 records. Neither type carries a `version` field either, so *when* this changed cannot be attributed, and this block does not guess. The canary is unaffected and the reason is worth writing down so nobody re-widens the claim: those two types carry no `message`, no `cwd` and no title, so they feed nothing the adapter reads; `Saw()` fires on the first record carrying the field rather than requiring all of them to; and the first record of 1,045 of 1,045 transcripts still carries it, which is what the head read actually depends on. §3.10's cell is reworded to *"on every JSONL record that feeds a field"*. ### 3.2 Codex CLI adapter sources — RESEARCHED FROM SOURCE, **NOT LIVE-VERIFIED** Codex is now installed on the dev PC (2026-08-01): **Codex Desktop** (VS Code app 26.727.51351, bundling `codex-cli 0.146.0-alpha.9.2` under `%LOCALAPPDATA%\OpenAI\Codex`) plus the **npm CLI** (`codex-cli 0.146.0`). The claims below were first read from `github.com/openai/codex` at commit `1e85ca09` (2026-08-01): `codex-rs/utils/home-dir/src/lib.rs`, `codex-rs/rollout/src/{lib,recorder,compression,policy,metadata}.rs`, `codex-rs/protocol/src/{protocol,models}.rs`, `codex-rs/login/src/auth/default_client.rs`, `codex-rs/thread-store/README.md` — and then checked against the live corpus. **§3.4 carries the verified results and the itemized remainder**; the adapter is not "done" until the remainder is discharged. **Layout.** `$CODEX_HOME` (default `~/.codex`) `/sessions///
/rollout--.jsonl`, fixed depth, no recursion. The date directory is **local** time, not UTC — deriving today's directory from a UTC clock silently loses sessions across midnight and DST, so the adapter walks the tree instead of computing a path. Files older than 7 days are compressed in place to `.jsonl.zst` (zstd level 3); the adapter reads `.jsonl` only — a `.zst` file is by construction ≥7 days cold and cannot be a live row, so skipping it avoids a zstd dependency. `rollout-compression.lock` and `*.tmp` in the same tree are not sessions. `archived_sessions/` is deliberately ignored. **Envelope.** `RolloutLine { timestamp, ordinal?, #[serde(flatten)] item }` with `RolloutItem` tagged `#[serde(tag="type", content="payload")]`: ```json {"timestamp":"…","ordinal":42,"type":"session_meta|turn_context|response_item|event_msg|compacted|world_state|…","payload":{…}} ``` `EventMsg` is **internally** tagged, so its discriminator sits *inside* `payload` alongside its fields (`payload.type == "token_count"`, with `info` / `rate_limits` as siblings). This differs from the outer envelope and is the easiest thing to get wrong. **Field mapping:** | Normalized field | Exact source | |---|---| | Session id | filename uuid, cross-checked against `session_meta.payload.id` / `.session_id` | | Working dir | `session_meta.payload.cwd`, then the last `turn_context.payload.cwd` | | Git branch | `session_meta.payload.git.branch` (display-only extra) | | Model | **last** `turn_context.payload.model` (not on `session_meta`) | | **Context %** | **derived**: last `token_count` → `info.last_token_usage.total_tokens ÷ info.model_context_window` | | **Quota** | `payload.rate_limits.primary` / `.secondary` → `{used_percent (0–100), window_minutes, resets_at (unix s)}` | | Plan / CLI version / history mode | `rate_limits.plan_type`, `session_meta.payload.cli_version`, `.history_mode` (display-only extras) | | Sub-agent thread | `session_meta.payload.agent_nickname` / `agent_role` non-null → not a session | | Last activity | file mtime (matches Codex's own `updated_at`/`recency_at` derivation) | `session_meta.payload.history_mode` is `legacy` (default) or `paginated` and changes which *message* records exist. The adapter does not branch on it: `policy.rs` persists `SessionMeta`, `TurnContext` and `EventMsg::TokenCount` under both modes, and those are the only records it reads. The mode is carried as a display-only extra so a live verification pass can see which one produced a given fixture. `TokenCountEvent.info` and `.rate_limits` are both `Option`. **Judgement call, UNVERIFIED (§3.4):** a `token_count` whose `info` or `rate_limits` is null is treated as *clearing* that datum rather than leaving the previous value standing. `protocol.rs` annotates the neighbouring field with *"`None` is unavailable, not a sparse-update recovery"*, which reads as "we do not have it" rather than "unchanged"; it is also the conservative side of the honest-gauge rule, since it never shows a number the vendor's most recent statement did not contain. The 2026-08-01 live pass could not settle it — no session in the corpus emitted a mid-stream null after a populated event — so the conservative reading stands unfalsified rather than confirmed (§3.4 "still owed"). **Codex capability gaps:** no cost in USD anywhere; **no process-liveness registry** (mtime is the only signal); no session title, so rows fall back to the workspace basename; cold `.zst` sessions unreadable under minimal deps. Reading the SQLite state DB (`codex-rs/rollout/src/state_db.rs`) is a **rejected** path — it would add a sqlite dependency for metadata the JSONL already carries, and `thread-store/README.md` confirms JSONL stays canonical and readable without SQLite. #### Re-measure 2026-08-16 — `codex-cli 0.147.0`; the map holds, and two new fields are traps `telltale doctor` reported this pin drifted (0.146.0 surveyed, 0.147.0 installed), so the survey above was re-run against the live corpus rather than re-read from source. **The pin now reads `codex-cli 0.147.0` and nothing in the field map changed.** **Corpus scanned.** `~/.codex/sessions/`, walked exactly as the adapter walks it (`archived_sessions/` not visited): **330 native rollouts**, 35 imported transcripts filtered on the `external-import-turn` marker, 8,408 `event_msg` records of which **1,313 are `token_count`**. By writer version: 169 rollouts at `0.147.0` and 2 at `0.147.0-alpha.6.5` — so **171 rollouts written by the installed build**, beside 144 at `0.146.0` and 9 at `0.146.0-alpha.9.2`. The re-measure rests on rollouts 0.147.0 wrote, not on old files re-read. **What held.** Every path in the field-map table above still resolves on 0.147.0-written rollouts, at the same rate or better than on 0.146.0 ones: | path | 0.147.0 | 0.146.0 | |---|---|---| | `session_meta.id` / `.session_id` / `.cwd` | 169/169 | 144/144 | | `session_meta.git.branch` | 149/169 | 120/144 | | `session_meta.history_mode` | 169/169 | 144/144 | | `turn_context.model` / `.cwd` | 163/169 | 31/144 | | `info.model_context_window` + `.last_token_usage` | 57/57 rollouts that carry a `token_count` | 25/29 | The sub-100% cells are **session shape, not drift**: a rollout that never reached a user turn has no `turn_context`, and one that never reached a model call has no `token_count` (112 of the 169 at 0.147.0 are `codex_exec` runs of that kind). `git.branch` is absent exactly when the `cwd` is not a repository — `session_meta.git` itself is present, carrying `commit_hash` and `repository_url`. Also holding: both canaries (`envelope type`, `session_meta record`) on all 324 rollouts that carry any parseable record; `secondary` still **null in all 1,278** populated `rate_limits`; `used_percent` / `window_minutes` / `resets_at` unchanged as the only window keys; and `plan_type: "plus"` on 1,257 of them. **What changed — additions only, no rename.** 0.147.0 adds fields and moves none, which is the case §3.10 says costs this program nothing because every reader here addresses keys by name. New on `session_meta`: `model_provider` (`"openai"`, 324/324), `base_instructions` (now an object `{text}`), `context_window`, `dynamic_tools`, and `git.commit_hash` / `git.repository_url`. New on `turn_context`: `turn_id`, `workspace_roots`, `current_date`, `timezone`, `approvals_reviewer`, `permission_profile`, `comp_hash`, `personality`, `collaboration_mode`, `multi_agent_version`, `realtime_active`, `file_system_sandbox_policy`. `effort` gained `ultra` and `xhigh` beside the previously-observed `low`/`medium`/`high`. **Two of the additions are traps, and are deliberately NOT read:** - **`session_meta.context_window` is `{window_id: }` — an IDENTIFIER, not a size.** The name invites reading it as the context denominator, and 324 of 324 carry it while only 93 rollouts carry an `info.model_context_window`, so wiring it up would look like it *widened* coverage. It cannot: a window id is not a token count, and dividing by one would be an invented number of exactly the kind §4a.1 forbids. The denominator stays `info.model_context_window`. - **`turn_context.multi_agent_version` is the literal `"v2"` on all 288 turn contexts**, so it is a format version, not a sub-agent marker. Treating it as one would reject every session as a sub-agent thread. The sub-agent filter still keys on `agent_nickname` / `agent_role` alone. A third addition was checked and left alone: `collaboration_mode.settings.model` carries a model id, and it **equals `turn_context.model` on all 288** turn contexts. The model source is unchanged rather than merely still-working. **What stays `CapNone`, now cited at 0.147.0.** Cost in USD, session title, sub-agent count and process liveness are all still absent — a `rate`/`limit`/`cost`/`title`/`pid` sweep over the 0.147.0 rollouts matches nothing beyond the `rate_limits` block already modelled. The §3.3 matrix row is unchanged. **What this pass could NOT exercise, stated rather than glossed.** No sub-agent thread and no imported transcript written by 0.147.0 appeared in the corpus — `agent_nickname`, `agent_role` and `external-import-turn` are all unobserved at this version. `ErrSubAgentThread` and `ErrImportedTranscript` therefore still rest on the 2026-08-01 observation, and this block does not claim otherwise: those markers are **unobserved here, not measured gone**. Two further observations that are new and cost nothing: `history_mode: "paginated"` finally appeared (one rollout, 0.147.0) and it carries `turn_context.model` and a populated `token_count` exactly as `legacy` does — so §3.2's "the adapter does not branch on `history_mode`" is now confirmed against a real paginated rollout instead of a source read. And 6 rollouts on disk contain **zero parseable records**, which exercises for real the `sampled <= 0` branch `internal/adapter/drift` calls its load-bearing case: they produce no drift report, correctly. ### 3.3 Cross-vendor capability matrix — the asymmetry is a design fact, not a bug | Field | Claude (disk) | Codex (disk) | Gemini (disk, §3.7) | Antigravity (disk, §3.8) | Cursor (disk, §3.9) | Grok (disk, §3.9a) | |---|---|---|---|---|---|---| | session id, cwd, git branch | yes | yes | id yes; cwd via `projects.json`; branch no | id yes; cwd via the trajectory blob's `file:///` URI; branch no | id yes; cwd via `workspaceStorage//workspace.json`; branch no | id yes; cwd verbatim in `summary.json`; branch **yes, unused** (`head_branch`, only when the cwd is a repo) | | model | yes | yes | yes (per message) | yes (per generation, id + display name) | yes (`modelConfig.modelName`, one string; sometimes the literal `default`) | yes (`current_model_id`, one string) | | token counts | yes | yes | yes (per message) | yes (per generation, self-checking) | context totals yes; per-message counts present and **always 0** | yes (per turn in `updates.jsonl`, plus a context total in `signals.json`) | | context window size | **no** | yes | **no** (static table in CLI source only) | **no** (statusline payload only) | yes (`contextTokenLimit`) | yes (`contextWindowTokens`) | | context % | **not derivable** | **derived** | **not derivable** | **not derivable** | **reported** (the vendor persists its own; derived from raw counts only if it is missing) | **reported** (`contextWindowUsage`, an integer the vendor truncates) | | quota / rate limits | **no** (statusline stdin only) | yes | **no** (runtime 429 handling only) | **no** (statusline stdin only; never persisted) | **no** — plan *entitlements* on disk, no consumption record | **no** — nothing account-level anywhere in the store | | cost USD | no (stdin only) | no | no | no | **no** — `usageData` `{}`, token counts unpopulated zeros | **per turn yes, session total no** — `costUsdTicks`, unit measured; no cumulative figure exists | | process liveness | registry exists, deliberately unread (§3.1) | none | none | `steps.status` exists, structural only (never observed in-flight) | `status`/`generatingBubbleIds` exist, structural only (never observed in-flight); Hooks is the real seam | `active_sessions.json` exists and was **measured empty during a live turn**; `events.jsonl` phases outlive the process | | session title | yes | no | yes (`summary` metadata) | **no** — the only free text on disk is prompt content | yes (`value.name`, vendor-generated) | yes (`generated_title`, vendor-generated; absent on headless runs) | | sub-agent count | **derived** (`subagents/` sidecar, §3.1) | **no** | **derived** (`chats//` nest) | **no** — `parent_references` observed empty | **no** — `isSubagent`/`numSubComposers` observed zero throughout | **no** — a `spawn_subagent` tool exists, nothing about it reaches disk | Codex is `CapNone` for the sub-agent count and not merely empty. Sub-agent *threads* do exist in the Codex format — `session_meta.payload.agent_nickname` marks one, and the adapter rejects those rollouts with `ErrSubAgentThread` — but they are whole top-level rollout files carrying no link back to a parent session, so there is nothing to attribute a chip to. Declaring the field and always emitting zero would assert "this Codex session is running no sub-agents", which is not something the format lets us check. Claude's quota lives on the statusline seam; Codex's lives on the disk seam. So the HUD's quota block is Codex-sourced today, and the CONTEXT column carries a Codex number beside a Claude em dash. **Nothing sources cost**, so the COST column auto-hides in every real v1 frame — see the `v1-capabilities` render in §7.3. Cursor is the first vendor to put a context percentage on disk as a number *it* computed, which makes it the only unmarked bar in that frame. Everything else about it is the asymmetry again from the other side: it is also the first vendor whose store holds live credentials, so the adapter's most load-bearing property is the list of things it does not read (§3.9, decisions/007). Grok is the second to report a percentage, and the first to write a **dollar figure** to disk at all — and the COST column still auto-hides on its rows, because what it writes is one turn's cost and never the session's (§3.9a). "Nothing sources cost" became "nothing sources a session cost", which is a narrower sentence and the same column. **Percentage comparability.** Codex's own `TokenUsage::percent_of_context_window_remaining` subtracts `BASELINE_TOKENS = 12000` from both numerator and denominator; Claude's `context_window.used_percentage` is raw input-token-based. **They are not the same statistic.** The adapter therefore does *not* reproduce Codex's baseline-normalized figure: it computes a plain `last_token_usage.total_tokens ÷ model_context_window`, declares it `CapDerived`, and the HUD marks it with an estimate marker. See §6 Q7 for the resolution and its alternatives. ### 3.4 Live verification (ADR-001) — first pass run 2026-08-01; remainder itemized **Environment:** Codex Desktop app 26.727.51351 bundling `codex-cli 0.146.0-alpha.9.2` (every live rollout in the corpus was written by it, `originator: "Codex Desktop"`, `source: "vscode"`), with npm `codex-cli 0.146.0` installed alongside. The source read above was taken at CLI `2.1.219`; no contradiction between the two surfaced except where noted below. **Confirmed:** - `sessions///
/` is the **local** date: events stamped `2026-08-02T00:12Z` (UTC) sit under `08/01`. Walking the tree instead of computing today's path was right. - `session_meta` writes **both** `id` and `session_id`, identical values. - `history_mode` is `"legacy"` on every fresh thread. - `model_context_window` is populated (`258400` for `gpt-5.6-terra`), so the derived context percentage works as designed. - `ordinal` is **not emitted** — the envelope is `{timestamp, type, payload}` only. Fixtures 0002/0003 keep their `ordinal` deliberately (the field must stay tolerated); fixtures 0006/0007 pin the observed no-`ordinal` shape. - `rate_limits`, live values: **free plan** = `primary` only with `window_minutes: 43200` (a 30-day window), `secondary: null`; **plus plan** = `primary` with `window_minutes: 10080` (7 days), `secondary: null` so far, `plan_type: "plus"`. The "record real values instead of hard-coding 5h/7d" instinct was right — neither plan matches the guessed pair, and labels derive from `window_minutes` alone. Newer fields (`limit_id`, `credits{}`, `plan_type`, `rate_limit_reached_type`) are parsed loosely; `credits.balance` has been observed as both `null` and the string `"0"`, so nothing in it is typed strictly. - Go can `os.Open`, head-read, and tail-read a rollout **while a live codex process holds it** (verified against an active session; Windows sharing mode is permissive). **Learned, not on the checklist:** 1. **Imported external-agent transcripts.** Desktop onboarding imported 35 Claude sessions into `sessions//` as rollout files. Markers: `session_meta` lacks `thread_source` (native threads carry `"user"`); every turn's `task_started.turn_id` is `external-import-turn-` (inside the head window in all 35 observed files); the single `token_count` is synthetic (zero components, non-zero `total_tokens`, null window, null `rate_limits`). The adapter rejects these with `ErrImportedTranscript` on the affirmative `turn_id` marker only — absence of `thread_source` is not used, so pre-`thread_source` CLI rollouts are unaffected. Rendering an imported Claude transcript as a Codex row is a cross-vendor double count; the filter is not optional. 2. **`archived_sessions/` semantics confirmed the hard way.** The first inspection pass found every real session in flat `archived_sessions/` and only imports in `sessions//`, which read as "Desktop sessions are invisible to the adapter." A later live session disproved that: Desktop writes live rollouts under `sessions//` and threads move to `archived_sessions/` when archived — the Desktop auto-archives its onboarding threads, which is what emptied the first hour. Ignoring `archived_sessions/` remains correct. 3. **Windows mtime does not reliably advance mid-session.** On an active session the newest records were stamped ~100 s *after* the file's mtime: NTFS defers the mtime update while the writer holds the handle. `LastActivity` from mtime therefore under-reports on live sessions (never over-reports). ~~The ruling is §6 Q8~~ — **ruled and implemented 2026-08-01**: `LastActivity = max(mtime, newest record timestamp)`, both adapters; see §6 Q8 for the rules. 4. Desktop threads run in per-thread scratch workspaces (`Documents\Codex\\`), so the workspace-basename fallback shows the thread slug, not a repo name. Cosmetic, vendor-truthful, unchanged. **Discharged 2026-08-01, same evening:** - ~~A rollout written by the **standalone CLI**~~ — observed live: two sessions from the npm CLI at 0.146.0 write `session_meta` with `originator:"codex-tui"`, `source:"cli"`, `thread_source:"user"` — same record shape and tree layout as the Desktop writer, so no adapter change (nothing keys on originator). The observation also reproduced §6 Q8 a second time: the first CLI rollout's mtime settled roughly twenty minutes after the session ended, when the writer released the file. **Still owed** (re-scannable on demand: `tools/scan-passive-tail.py`): - Null `info`/`rate_limits` mid-stream, "cleared" vs "unchanged" (§3.2): the corpus contained **no mid-stream nulls**, so the conservative "clearing" reading stands unfalsified rather than confirmed. - An **API-key login** capture (rate_limits expected absent), and whether a paid plan ever populates `secondary`. Capture path when wanted: `codex login --with-api-key` (reads the key from stdin), run one short session, re-scan, then plain `codex login` to return to the ChatGPT plan. - The 7-day `.zst` compression pass — unobservable until the corpus is a week old. *Re-scan 2026-08-02:* 5 native rollouts (35 imports filtered), including one new Desktop session — still zero mid-stream nulls; plus-plan `secondary` still null across all 32 populated `rate_limits`; no API-key-signature session; no `.zst` anywhere under `sessions/`. Oldest native rollout is 2026-08-01, so the `.zst` pass stays unobservable before ~2026-08-08. *Re-scan 2026-08-11:* 316 native rollouts, 1,251 populated `rate_limits`. All three owed items stay negative. Two of them now rest on 316 rollouts rather than the 5 of 2026-08-02, and a fourth observation changes what the API-key capture must look for. - **Mid-stream nulls: still zero**, now across 316 rollouts rather than 5. The conservative "clearing" reading stays unfalsified rather than confirmed. - **`secondary`: still null** in all 1,251 populated `rate_limits`. A plus plan does not populate it. - **`.zst`: still zero files** under `sessions/` — and this item is no longer blocked on time. The oldest native rollout is 2026-08-01, so it was 10 days old at the scan, and Codex wrote new rollouts on 9 later days. The pass is **measured absent on this box, not unobservable**. The prediction above expired on ~2026-08-08. - **A `rate_limits` object can report "no windows" without being absent.** 21 `token_count` records across 4 native `codex_exec` sessions (`cli_version` 0.146.0, 2026-08-07 and 2026-08-08) carry a `rate_limits` OBJECT whose `primary`, `secondary`, `plan_type`, `limit_name`, `individual_limit` and `spend_control_reached` are all null, beside `limit_id:"premium"` and a `credits` block. Each of those sessions is null from its FIRST `token_count`, so this is not a mid-stream clear and it does not settle §3.2. **What it costs the owed capture.** `tools/scan-passive-tail.py` detects the API-key signature as `rate_limits is None` alone, so a session of this shape passes it unseen and reports as negative. Whoever runs the capture must check both signatures: an absent `rate_limits`, and a present one whose windows are all null. **The cause is deliberately not stated here** — nobody recorded which auth mode those four sessions ran under, so a claim that they are API-key sessions would be an inference, and §4a.1 forbids one dressed as a reading. *Re-scan 2026-08-16* (during the §3.2 re-measure to `codex-cli 0.147.0`; that block carries the field-map results, this line carries only the owed items). 330 native rollouts, 1,313 `token_count` records, 1,278 populated `rate_limits`. **All three owed items stay negative**, and the newer corpus does not move any of them: - **Mid-stream nulls: still zero**, now across 330 rollouts and 1,313 `token_count` records. The conservative "clearing" reading stays unfalsified rather than confirmed. - **`secondary`: still null** in all 1,278 populated `rate_limits`. A plus plan still does not populate it. - **`.zst`: still zero files** under `sessions/`, with the oldest native rollout now 15 days old. Measured absent on this box, as the 2026-08-11 entry already ruled. - **API-key capture: still not taken**, and this pass checked **both** signatures the entry above demands. Signature A (`rate_limits` absent on every `token_count`): **zero sessions**. Signature B (a `rate_limits` object whose windows are all null): **4 sessions** — the same four from 2026-08-07/08 at `cli_version` 0.146.0, and **no new ones at 0.147.0**. So the shape has not spread, and the auth mode behind it stays unrecorded and unclaimed. ### 3.5 Framing rule — now measured, not assumed (see §4) The §4 hazards were quantified against the live Claude corpus: - **64 KiB cap: firing.** 107 of 13,211 records exceed 64 KiB; the longest single line is **1,004,230 characters**. `bufio.Scanner` at its default cap returns `bufio.ErrTooLong` on ~0.8% of records, and an unchecked `Err()` silently truncates the file — reading as "no more sessions". Adapters use `bufio.Reader.ReadBytes('\n')` via `internal/jsonl`, which is the one tested implementation of this rule. - **U+2028/U+2029: not observed** (0 raw `E2 80 A8`/`E2 80 A9` bytes across the 40 newest transcripts). That is absence of evidence, not absence of hazard — the records carry model-authored text and both characters are legal unescaped inside a JSON string. The §4 byte-level rule is unchanged and both fixtures embed the character to pin it, with a `.gitattributes` entry plus a byte assertion in each adapter's tests so a checkout rewrite fails the build instead of silently disarming the test. - **Trailing partial line:** all 7 live transcripts ended on `0x0A` at survey time, so writes look line-atomic — a sampled observation, not a guarantee. The hold-until-`\n` rule stands, and both fixtures end in a deliberate truncated record with no trailing newline. - **Windows concurrent read:** all 7 live transcripts opened with share-read/write while their processes were running. Go's `os.Open` already requests `FILE_SHARE_READ|WRITE|DELETE`, so no special handling is needed — recorded here so nobody "fixes" it later. ### 3.6 Degradation rule A vendor field the adapter cannot read renders as `—` (absent), never as a zero or a stale value presented as fresh. A record that parses only partially degrades the fields it could not source to `—` and keeps the rest. A truncated trailing line is not a record. The exact renders are §7.7. ### 3.7 Gemini CLI seam — source-verified 2026-08-02; first live pass itemized **Environment:** gemini-cli 0.53.1 installed via npm 2026-08-02; the persistence layer read at tag v0.53.1 (`packages/core/src/services/chatRecordingService.ts` + `chatRecordingTypes.ts` for the writer and record shapes, `config/storage.ts` for the tree, `config/projectRegistry.ts` for the slug registry). This is the writer's own source, not its docs — the same standard as the Codex `rollout` read (§3.2). **Layout (from source):** - Sessions: `~/.gemini/tmp//chats/session--.jsonl`. The filename embeds only the session id's first 8 characters; the full id is on the first record. `~/.gemini/projects.json` maps absolute project paths (lowercased on Windows) to slugs; the slug scheme replaced sha256-hash directory names in 0.5x, and the registry self-heals from `.project_root` markers. - Sub-agent transcripts nest at `chats//.jsonl` — a structural parent link, which is why Gemini declares `subagents` (derived) where Codex cannot (§3.3): Codex's sub-agent threads are top-level files with no path back to a parent. - Legacy pre-JSONL sessions are single-document `*.json`; the adapter skips them. **Record shapes (from source):** the first line is metadata (`sessionId`, `projectHash`, `startTime`, `lastUpdated`, optional `kind`/`directories`); message records carry a string `id`, `timestamp`, `type`, and on `type:"gemini"` a `model` and a per-message `tokens` summary (`input` = promptTokenCount, `cached` a subset of it, `output`, `total`); `{"$set":{...}}` records patch metadata (including `summary`, the session title, and whole-array `messages` checkpoints that can put megabytes on one line — the §4 framing rule is earning its keep here); `{"$rewindTo":id}` truncates. **Messages are upserts**: the writer re-appends the full record under the same id when tokens or tool calls settle, so a linear last-wins pass needs no dedup map. **Traps encoded in the adapter:** - The writer **deletes** a session file on exit when it holds no resumable content, so a file vanishing between Discover and Read is normal operation (`ErrSessionGone`, row dropped silently). - Nothing quota-shaped is persisted — rate limiting exists only as runtime 429 handling (`googleQuotaErrors.ts`, `retry.ts`). `quota` is CapNone, not empty. - No context-window size reaches disk; the CLI's own percentage divides by a static per-model table compiled into its source. An assumed denominator is an invented gauge, so `context_pct` is CapNone — the §4a.7 sketch guessed "derived" here, and the source read falsified the guess. - `workspace` is read verbatim from the vendor's registry entry (REPORTED, a lookup not a computation), with a fidelity caveat: the vendor lowercases the recorded path on Windows. - The adapter replays the writer's grammar, not just its records: `$rewindTo` truncates the ordered message log (a rewind to an id outside the read windows conservatively clears it), and a `$set` messages checkpoint clears and rebuilds it — both mirroring the vendor's own loader. Independent review (2026-08-02) caught the first cut ignoring both; a rewound-away 215k-token reading would have kept rendering. - **Bounded-read limitation, stated:** the head/tail windows share the seam behaviour of every adapter — a record crossing the boundary is read by neither window, and a single line larger than the tail budget (256 KiB) is outside the read entirely. On Gemini that line can be a whole-conversation checkpoint, so a giant checkpoint's values are invisible until the next ordinary record re-establishes them. Accepted as the same tradeoff the other adapters carry; the live pass below sizes real checkpoints to check whether the budget needs raising. **Market note (2026-08-02, post-merge):** Gemini CLI stopped serving consumer tiers (free/Pro/Ultra) on 2026-06-18; it remains live for Gemini Code Assist Standard/Enterprise licenses and paid API keys, with Antigravity CLI (`agy`) as the consumer successor (ADR-003 addendum). This adapter therefore covers the enterprise/API-key flavour. (A live session was nonetheless produced on this machine 2026-08-03 — the auth flavour behind it was not investigated; recorded as an observed fact only.) **First live pass — RUN 2026-08-03 and PASSED** (gemini-cli 0.53.1, the same version the source read pinned; one real session, ~1.6 MB, 50 records, written live during the check). Adapter output against it: discovered 1; model `gemini-3.5-flash`; workspace `c:\users\sanle` via `projects.json` (lowercased-path registry confirmed); LastActivity rode the mtime side of the Q8 fold (the file was being touched after its newest record timestamp — the live-write pattern the fold exists for); subagents derived 0; zero diagnostics; Validate green. Name is absent and honestly so — the header carries no title field; the HUD label falls back to the workspace basename. The itemized checks resolve as follows (original list kept below for the record): - Metadata line is the first line: **confirmed.** Main sessions **carry `kind:"main"`** — the fixture's omitted-field assumption was falsified; the adapter is unaffected (only `"subagent"` branches) and the healthy fixture now carries `kind:"main"` to match reality. - Filename shape and prompt registry entry: **confirmed** (entry present, lowercased). - Upserts against real traffic: **confirmed** — 27 message records, 20 distinct ids (7 in-place updates). Per-message `tokens` **confirmed live** with shape `{cached, input, output, thoughts, tool, total}` — unused by the adapter (context stays CapNone per §4a.7's falsification), recorded for fidelity. Delete-on-exit for non-resumable sessions: not observable (this session persisted). - Checkpoint sizes vs the read budgets: the two live `$set` messages checkpoints were **7.3 KiB and 14.5 KiB — neither approached the 64 KiB scanner cap**; the "expected: yes, on any long session" guess did not materialize (checkpoints snapshot compactly). The budget pressure came from elsewhere: **single message lines up to 746 KiB** were observed, dwarfing both budgets. The newest checkpoint sat wholly inside the 256 KiB tail and the bounded read produced correct output — a tail that starts mid-line resyncs at the next newline, exactly the framing design. - `$rewindTo`: **not exercised by this session** — remains fixture-verified only. **ADR-003's verification hold is RELEASED**: the launch post may claim the Gemini adapter live-verified. Original itemized list (as written before the pass): - Confirm the metadata line is the FIRST line in every live file, and whether main sessions carry `kind:"main"` or omit the field (the fixture assumes omitted). - Confirm filename shape and that `projects.json` gains the entry promptly (the registry can lag a fresh project; the adapter treats a missing entry as absence). - Confirm the upsert pattern and per-message `tokens` against real traffic, and the non-resumable delete-on-exit behaviour. - Observe whether a live `$set` messages checkpoint exceeds the 64 KiB scanner cap in practice (expected: yes, on any long session), and whether any exceeds the 256 KiB tail budget (which would put whole checkpoints outside the bounded read — see the limitation above). - Observe a real `$rewindTo` and confirm the truncate-or-clear replay against the loader's behaviour on the same file. ### 3.8 Antigravity CLI seam — surveyed live 2026-08-02; statusline is the seam **Environment:** agy 1.1.9, installed 2026-08-02. Closed source (the GitHub repo, google-antigravity/antigravity-cli, is docs and examples only), so this survey is the Claude Code method: documented contracts cross-checked against live observation, no source read possible. **Disk verdict (first survey — superseded the same day by the re-survey below).** Interactive and headless sessions both write `~/.gemini/antigravity-cli/conversations/.db`: SQLite, protobuf blobs in every payload column (`steps.step_payload`, `gen_metadata.data` — inspected read-only). The first survey judged the blobs unparseable without guessing (and the repo had already rejected a SQLite dependency once, §3.2), and found no transcript: the docs and the live payload both advertise `~/.gemini/antigravity/brain//.system_generated/logs/transcript.jsonl`, that path did not exist, and `antigravity-cli/brain/` held only empty `scratch/` dirs at survey time. Verdict then: no honest HUD adapter; agy shipped as a statusline-only vendor (ADR-004), with a standing watch item on the transcript. **Re-survey (2026-08-02, later the same day; ADR-005 decision 5, prompted by ccusage issue #1402): the disk seam is OPEN, and the watch item is RESOLVED — the transcript is real.** Corpus: four conversations, all agy **1.1.9** — the same version the first survey ruled on; this is a correction, not a version change. - `brain//.system_generated/logs/transcript.jsonl` **exists for all four conversations**, non-empty, plus an undocumented untruncated sibling `transcript_full.jsonl` (`transcript.jsonl` marks its cuts with `truncated_fields`). Written default-on, with `enableTelemetry: false`, under `antigravity-cli/` — the docs' advertised `antigravity/` tree still does not exist. Why the first survey saw only empty dirs is unresolved (a flush-timing artifact vs. a missed `.system_generated` subdir); both observations stand as recorded. The transcript is plain JSONL, one line per `steps` row (verified exact, 38/38 across the corpus): step_index, source, type, status, created_at (RFC3339 UTC, second resolution), content, thinking, tool_calls, exit_code. - `gen_metadata.data` **decodes with a stdlib protobuf wire walk** (no schema, no deps). Per-generation token counts live at `#1.#4`: uncached input (`#2`), total output (`#3`), cache-read (`#5` — inferred from position and magnitude, lower confidence), thinking (`#9`), answer (`#10`), and the per-generation response id (`#11`, the dedup key — the *top-level* `#4` UUID is constant per conversation and must not be used). Model id at `#1.#19` (`gemini-3.6-flash`) and display name at `#1.#21`, matching the live statusline string verbatim. **Self-check: `thinking + answer == output` held in 15/15 decoded generations** — the arithmetic identity that promotes this from field-guessing to a schema. An adapter must assert it at read time and degrade to absent rather than render a number that fails it. - What an adapter could report as **measured**: Name, Model, Workspace (trajectory-blob URI), LastActivity (max transcript `created_at`; the trajectory-blob timestamp is session *start* — do not use it). **Partial**: context used-tokens (numerator measured; the 1,048,576 window denominator appears nowhere on disk — it must come from the statusline payload or a constant with stated provenance). **Absent**: cost (consumer auth; no pricing on disk), quota (server-refreshed in memory, never persisted). **Structural only, never observed live**: liveness (`steps.status` — all 38 observed rows `DONE`; no in-flight session sampled yet) and subagents (`parent_references` empty, `has_subtrajectory` all zero, `define_subagent`/ `invoke_subagent` present in the tool registry). - Build cautions, recorded before any adapter work: **WAL sidecars are load-bearing** (a `-wal` larger than the `.db` was observed — copy `.db`+`-wal`+`-shm` together, or open the live file strictly read-only and let SQLite replay; never open it for write). `conversation_summaries.db` is a stale index (1 row for 4 conversations) — enumerate `conversations/*.db`, never trust it. `cache/last_conversations.json` is written at session start and points at the *previous* session for a reused workspace. Transcript `content`/`thinking` and the request blobs are PII (full prompt text, file contents; the account email appears in `cli.log`) — an adapter reads structural fields only and never surfaces those into the HUD or any log. All protobuf field numbers are reverse-engineered and unversioned; the corpus is one day, one model — field stability across models is untested. **Statusline seam — verified live.** Capture method: a temporary statusline command (`telltale-capture.cmd`) appending each stdin payload to a file; six payloads captured across one interactive session. Confirmed against the documented schema: `product: "antigravity"` on every payload (the routing marker); model with display_name/effort; full context accounting (`context_window_size` 1,048,576 observed, `used_percentage`, per-request `current_usage` incl. cache reads); **named weekly quota buckets** (`gemini-weekly`, `3p-weekly` on the Starter tier) each carrying `remaining_fraction`, `reset_time`, `reset_in_seconds`; `agent_state` observed transitioning `tool_use` → `idle`; `tool_confirmation_pending` observed `true` while a permission prompt was on screen. Documented but not yet observed live (the capture session ran outside a repo): `vcs`, `artifact_count`, `task_count`, `execution_mode` — the parser carries them as optional; the branch segment's live confirmation is the itemized remainder here. **Adapter built, 2026-08-02 (`internal/adapter/antigravity`).** What the adapter took from this survey and what it left: - **Took:** Model (`#1.#21` display name, `#1.#19` id fallback), Workspace (the trajectory blob's URI, converted to a native path), LastActivity (the Q8 fold over the transcript's newest `created_at` and the mtimes of the transcript, the database and its sidecar — the sidecar because on a live conversation that is the file being written). All three REPORTED; nothing is derived. Name was sourced from the conversation id at first shipping (shortened for the grid, the only label on disk that is not somebody's prompt) — **revised 2026-08-12 (ruled):** the HUD row's Name showed the id prefix instead of the workspace basename every other vendor with no on-disk title falls back to, so Name moved from REPORTED to `CapNone` and the adapter no longer writes it; the HUD sources the row's label itself, the same fallback a Gemini row takes with no summary of its own. - **Left:** context % (the numerator is measured and the denominator is not — the token totals are display-only extras instead), cost, quota, liveness, subagents and (as of the 2026-08-12 revision) name, all `CapNone` for the reasons itemized above. The liveness and subagent deferrals are pending live observation, not pending effort. - **The identity is asserted, not assumed:** every generation must satisfy `thinking + answer == output` before its numbers count. A generation that fails contributes nothing and the row says a self-check failed. Across the live corpus on the day the adapter landed the identity held **16/16** (one more generation than this survey's 15, a conversation having advanced in between). - **Zero dependencies added.** The `.db` and `-wal` bytes are read by `internal/sqlite`, a read-only reader written for this seam: header, `sqlite_master`, table b-trees, overflow chains, and a WAL overlay applying SQLite's own recovery semantics. The rejected alternative and the reasoning are decisions/006. - **Live verification, same day:** all five local conversations discovered and read; model `Gemini 3.6 Flash (High)` on every one; three workspaces resolved to `C:\Users\sanle` and two absent (those conversations genuinely carry no URI — absence, not degradation); a real 284 KiB `-wal` parsed and accepted with no diagnostic; every session passed `Validate` with an empty degraded set. **Fidelity notes:** the docs' storage path (`antigravity/`) and the real one (`antigravity-cli/`) disagree — the re-survey confirmed the real tree (the docs path has never existed on this machine); `model.id` equals the display string ("Gemini 3.6 Flash (High)"), not a machine id; the payload carries the signed-in email and plan tier, so real captures are PII and never enter `testdata/`. **Re-verification, 2026-08-15 — the pin moves to `agy 1.1.13` and the field map does not.** agy self-updates, so the version below was read from `agy --version` at read time, never from a release note. The corpus is 20× the one §3.8 ruled on: **81 conversations, 4,284 transcript records, 1,926 decoded generations.** The adapter was run over the live tree unmodified and every field it declares still sources. | field | §3.8 (agy 1.1.9) | re-read (agy 1.1.13) | verdict | |---|---|---|---| | model | 5/5, one id string | **81/81** | sources; id string DRIFTED, see below | | workspace | 3 of 5, 2 absent | **23 of 81**, 58 absent | sources; absence confirmed genuine | | last_activity | measured | **81/81** | sources, no drift | | tokens (extras + §7.17 sum) | identity 15/15 | **identity 1,926/1,926** | sources, no drift | | liveness | `CapNone` — all 38 rows DONE | `CapNone` — **164 RUNNING now seen** | stays `CapNone`, for a NEW reason | | subagents | `CapNone` — never observed | `CapNone` — **3 `INVOKE_SUBAGENT` steps** | stays `CapNone` | **0 rows degraded, 0 diagnostics, 0 unparseable transcript records** across the whole corpus — including conversations whose `-wal` sidecar is 25× the `.db` it belongs to, which is the sidecar contract still holding at scale. - **The two vendor release-note claims are NOT corroborated at this seam, and stay vendor claims.** 1.1.13 claimed a transcript-corruption fix during compaction: 0 of 4,284 records failed to parse, so there is no corruption here to have been fixed and nothing measurable changed. 1.1.12 claimed a Windows drive-letter fix in a transcript path converter: all 23 workspace URIs present read `file:///C:/Users…` — the same form §3.8 recorded — and `pathFromFileURI` converted 23 of 23 to `C:\Users\sanle`. Whatever the vendor fixed, it was not the shape this adapter reads. - **Workspace absence is absence, not a broken converter.** This is the check that separates the two, and it was run rather than assumed: of 81 trajectory blobs, **23 contain a `file:` substring anywhere in their bytes and 58 contain none.** The 58 carry no URI to convert, so the field is correctly absent (§4a.1's zero-vs-absent distinction), and the converter has no silent failure hiding behind them. - **`status` gained `RUNNING`, and it is still not liveness — the new evidence makes the refusal stronger.** §3.8 declined the field because every one of 38 rows read `DONE`. The re-read found the missing state: **164 `RUNNING` rows in 27 of 81 transcripts**, 158 of them on `RUN_COMMAND`. But the oldest is dated **2026-08-03, thirteen days before the read**, in a conversation nothing has touched since. `RUNNING` therefore means "no terminal status was ever written for this step", which is equally true of a live command and of one whose process was killed. Wiring it would have reported dozens of long-dead sessions as working — the exact dishonest-gauge failure ADR-001 exists to prevent. The field stays `CapNone` and the HUD keeps classifying age from `last_activity`. - **`INVOKE_SUBAGENT` is now observed (3 steps) and still buys no count.** It is the first evidence the feature is used at all. A step recording that a subagent was invoked at some past moment is not a number of subagents running now, which is what the field means, so the `CapNone` stands. - **The model id gained a second spelling.** `gemini-flash-3.6-high-control` (37 rows) now appears beside `gemini-3.6-flash` (41 rows) for the one display string "Gemini 3.6 Flash (High)", and **3 rows carry the display name at `#1.#21` with nothing at `#1.#19`** — the first corpus to exercise the adapter's either-half-will-do fallback. Both strings are read verbatim; normalizing them would be inventing a vocabulary. - **New transcript keys, none of them read:** `error` (23), `error_code` (9), alongside the `truncated_fields` (345) §3.8 already described. The `step` struct is an allowlist, so added keys cost nothing — and no key the adapter *does* read has gone missing. - **A `logs/chunks/` tree appeared (first seen 2026-08-14, on the 4 newest conversations)**: `chunks/transcript/00000000.jsonl` plus a `chunks/transcript_full/`. This is the likeliest mechanism behind the compaction claim. It changes nothing today — the flat `transcript.jsonl` is still written and was **byte-identical (md5) to its single chunk on all 4**. All four hold exactly one chunk, so the case that would matter — whether the flat file stays complete once a second chunk exists — is **unobserved**, and is the standing watch item this block leaves behind. - **The docs' `antigravity/` tree now EXISTS and is still not the data.** §3.8 recorded that it had never existed. At 1.1.13 `~/.gemini/antigravity/` holds a `bin/`, a `builtin/skills/` tree and crash logs, while its `brain/` and `conversations/` are **empty** and all 81 conversations remain under `antigravity-cli/`. An adapter that picked its root by probing which path exists would now pick the empty one; this one roots by name. **`/quota` cross-check: MEASURED, and deliberately NOT built.** The probe was gated on proving it is free before anything could be wired to it, because a probe that quietly started a turn would plant phantom rows in telltale's own gauge. Two `agy -p "/quota"` invocations (PowerShell — Git Bash mangles the leading slash) against a quiescent store, each bracketed by a full snapshot: - **No residue in the store the HUD scans.** 81 conversations / 81 brain dirs / 81 HUD rows / 1,926 generations / identical token totals, before and after both runs. The probe writes only `cli.log`, `log/`, `last_check.timestamp`, `updater/update_status.json` and a **0-byte** `crashes/crash__.log` sink — none of them in the adapter's scan path. - **No token cost.** Both runs returned identical remaining figures (87% / 98% / 100% / 100%) and neither added a generation. A spent turn writes a conversation and a `gen_metadata` row, as every other agy turn in this corpus did. **So the gate passes and the feature is still wrong to build**, on evidence the measurement itself produced: the probe takes **2.9–3.8 s** and its numbers come from the server. Wiring it into a gauge would break two rules that are not about cost at all — the statusline path "reads nothing beyond stdin" (§2, revisitable "only with a measured budget", and 3 seconds is not one), and the gauges make no network calls (`CLAUDE.md`'s read/write boundary). It would also spend that round trip cross-checking a number the vendor already hands us free on stdin, which is where the quota buckets come from today. Nothing is wired; this paragraph is the record of why. **`/config` in `doctor`: NOT built, and the measurement that would license it was not made.** `doctor`'s charter is a cheap local `--version` parse that "never starts a turn, spends quota, reads a credential, writes under `~/.telltale`, or calls the network" (`cmd/telltale/main.go`). The `/quota` timings above show an agy slash command is a different animal: seconds, not milliseconds, and server-backed. Its behavior with the network or auth absent was **not measured** — doing so means disabling networking or de-authenticating the operator's account — and an unmeasured network-backed probe in `doctor` would red-cross exactly when the operator is offline, which is the dishonest cell that mode exists to avoid. `doctor` is left alone. **Amended 2026-08-16 — the multi-chunk case is PINNED, and it is still not MEASURED.** The re-verification block above left a standing watch item: whether the flat `transcript.jsonl` stays complete once a second chunk exists. That is still unobserved. No live multi-chunk conversation has been captured on any machine, and this amendment does not claim one. What changed is that the adapter's side of the question now has tests instead of a comment — `internal/adapter/antigravity/multichunk_test.go`, over a **synthesized** two-chunk fixture built by `testdata/gen_fixtures.py`. **The distinction is the whole point of writing this down.** Every other fixture in that directory reproduces a shape this section MEASURED. This one reproduces a shape nobody has seen, extended from the measured single-chunk shape along the vendor's own naming (`logs/chunks/transcript/0000000N.jsonl`, `logs/chunks/transcript_full/`), with the one relationship the corpus did measure — the flat file is byte-identical to its chunk — carried forward to two chunks. So the fixture pins **the adapter's contract with itself**, never a vendor claim, and nothing here is admissible as evidence about what agy writes. **What the pin found: the adapter is correct under that contract, and no code changed.** Six properties now have tests, and each was mutation-checked — the guard it names was deliberately removed and the test failed with the intended message, because a green test that cannot fail pins nothing. - **The flat file is what gets read.** The chunk tree's `transcript_full` files carry a poison step dated 23:59 against the flat file's newest at 09:11:39, dated in the past so the future-skew guard cannot silently swallow it. An adapter that followed the chunk tree moves `last_activity` by fourteen hours and the test says so by name. Note the limit honestly: `chunks/transcript/` is byte-identical to the flat file, so a switch to *that* is invisible to any test — which is precisely why the switch must not be made on the strength of a directory name. - **A flat file frozen at the first chunk costs precision, not the row.** This is the worst plausible form of the unobserved case: the vendor stops appending to `transcript.jsonl` and writes only into chunk 1. The §6 Q8 mtime fold carries it — the database is the file still being written — so `last_activity` stays correct, nothing degrades, and nothing is invented. That is an honest outcome rather than a lucky one: the adapter never claimed the transcript was its only clock. - **Damage inside the second chunk degrades the reading, not the row** (ADR-001's partial-read rule). Both branches are exercised: JSON that does not parse, and JSON that parses carrying a timestamp that does not. Three torn records, counted once, stated once in `Diagnostics`, and the newest surviving step still dates the row. - **The read budget's blind middle is absence, not corruption.** A transcript big enough to chunk is the first one this adapter splits into a head and a tail read at all, and the ~120 records it skips must never be reported as unparseable. An adapter accusing the vendor of damage that is not there is the same class of dishonesty as rendering a number it did not read. - **The head/tail overlap guard now has a test, and it had none before.** Every other fixture transcript on this vendor is under 1 KiB — swallowed whole by the tail read, never reaching `jsonl.Head`. The narrowing branch executes only between the tail budget and the full budget, so the test truncates a copy into that band and damages one record in the contested bytes. With the guard the count is 1; with it removed the measured result is **2 unparseable transcript records skipped** for one torn record. Duplication is invisible to every other signal here, because the newest-timestamp fold is a maximum and survives being fed a record twice; the unparseable counter is a sum, and it is the one place a doubled read surfaces. - **The PII boundary holds over a transcript 400× the size of the others**, chunk tree included. **What is still owed is the capture, and it is unchanged by any of this.** A synthetic fixture can prove the adapter is self-consistent across the behaviors the vendor might have. It cannot say which one agy has, and the four conversations that grew a chunk tree still hold one chunk each. The instrument this seam wants is a live multi-chunk conversation, and nothing above substitutes for it. **Re-capture, 2026-08-17 — the statusline pin moves to agy 1.1.13, and one field is FALSIFIED.** The statusline seam block above is pinned at 1.1.9 and six payloads. That record is superseded here and kept as the older reading. Capture method: the same temporary statusline command as 2026-08-02, appending each stdin payload to its own file; **fifteen payloads across one live interactive turn**, in Windows Terminal. The session ran outside a git repo again, so `vcs` is still unobserved. **The four-bucket prediction is CONFIRMED, with ids rather than labels.** §3.8's `/quota` cross-check reported four windows under human labels and refused to call that an observation of bucket ids, because the labels are not the ids. The payload now names them: **`3p-5h`, `3p-weekly`, `gemini-5h`, `gemini-weekly`** — a five-hour and a weekly window for each of two model families. **All four carry `remaining_fraction`, `reset_time` AND `reset_in_seconds`**; the 1.1.9 record claimed those three keys for two buckets, and they hold for four. Nothing in the renderer changed, which is what §2.1's verbatim-id rule was built to buy. The `week` page's fixture already carried all four names, so the fixture was right before the capture was taken. Two quota readings are recorded because both are honesty cases rather than trivia: - **Three of the four buckets reported `remaining_fraction` exactly 1**, serialized as the bare literal `1` rather than `1.0`. That renders `0%` used — a MEASURED zero, drawing a full empty track, not an absence. §4a.1's zero-versus-absent distinction, arriving off the wire this time instead of out of the renderer. `week.txt`'s `3p-weekly ─── 0%` row is the shape this produces, and it was synthesized correctly. - **`remaining_fraction` never moved across the fifteen fires.** Only `reset_in_seconds` counted down. The quota a line draws is therefore not the turn it is drawing. This is a vendor property, it is not hidden, and nothing here compensates for it. **`agent_state` is observed live, and the documented vocabulary is incomplete.** The turn moved `authenticating` → `idle` → `working` → `idle`. `idle` and `working` are on the documented list; **`authenticating` is not**. The renderer draws an unknown state verbatim in dim, so an unlisted value costs nothing — which is the point of having built it that way. The list is documented, not closed, and this is the measurement that says so. `thinking`, `tool_use` and `initializing` did not appear on this turn. **The payload GREW since 1.1.9, into the same shape Claude's block has.** `context_window` now carries `total_input_tokens`, `total_output_tokens`, `context_window_size`, float `used_percentage` and `remaining_percentage`, and a populated `current_usage` object (`input_tokens`, `output_tokens`, `cache_creation_input_tokens`, `cache_read_input_tokens`) — the §7.16b block's shape, on a second vendor. Also present and not previously recorded here: `conversation_id` (**equal to `session_id` byte for byte on all fifteen fires**), `exceeds_200k_tokens`, `sandbox.enabled`, `terminal_width` and `model.effort`. The parser was checked field by field against this capture: it models fourteen of the sixteen top-level keys. `email` is excluded deliberately and that exclusion is tested. The one field genuinely unmodelled by silence was `exceeds_200k_tokens`, and `internal/antigravity/stdin.go` now carries the agy-side reason instead of inheriting a ruling made about Claude's payload. **Nothing new was wired to a renderer or a relay** — §7.16's display hold binds this seam. **The numbers populate on the LAST fire only, and that is a zero-versus-absent collapse in the vendor's own serializer.** Through fourteen of fifteen fires the vendor sent `used_percentage: 0` with the token counters at `0`; every real number arrived at once on the final fire, after the turn. On the very first fire it sent `used_percentage: 0` alongside `context_window_size: 0`, which is a placeholder that no pointer type can tell from a measured zero — the vendor writes a literal `0` instead of omitting the key. So the statusline honestly draws `ctx 0%` for the length of a turn and then jumps. This is recorded, not compensated for: inferring "not yet known" from a zero the vendor stated would be exactly the invention §4a.1 forbids, and the collapse happens upstream of anything this repo controls. **`transcript_path` is FALSIFIED, and §2.1's refusal now has a measured reason.** §2.1 declined to display the field because the payload's value was *unverified* — the transcript is real on disk, but nobody had captured a payload to check whether its path pointed at it. This capture checked. **It does not.** The payload roots the path at `~\.gemini\antigravity\brain\\.system_generated\logs\transcript.jsonl`, and the real transcript for that same session exists only under `~\.gemini\antigravity-cli\brain\\`. The payload drops the `-cli` segment. Both roots exist at 1.1.13 and that is what makes the trap sharp: `antigravity/brain/` is present and EMPTY while `antigravity-cli/brain/` holds every conversation, so a reader that trusted the payload would open nothing and could not distinguish a missing file from a missing session. This is the docs-versus-data tree disagreement §3.8 recorded in 2026-08-02, still unfixed at 1.1.13 and now measured on the statusline seam as well as on disk. **No code trusted it, and that was checked rather than assumed.** Every use of `transcript_path` and `TranscriptPath` in the repository was read. No non-test code dereferences the field — it is parsed and held, never opened, stated or rendered — and `TestTranscriptPathIsHeldButNeverAPath` already pins that. The HUD adapter reaches the transcript by its own root by name (`internal/adapter/antigravity`), never by trusting a path handed to a gauge, which is the decision that makes this vendor bug a non-event here. So nothing is fixed, because nothing is broken; §2.1's refusal is upgraded from caution to measurement and the field stays parsed-and-unused, because deleting it would delete the evidence. One further shape is recorded for whoever reads a raw capture next: before a session id exists, the vendor joins an EMPTY id segment rather than omitting the key, so the path collapses to `…\brain\.system_generated\logs\transcript.jsonl`. **One inventory cell went stale with the re-verify, and is now fixed.** §3.10's canary table read `agy 1.1.9` for the Antigravity row after the adapter moved to `agy 1.1.13`; the row was corrected in the same-day ledger pass. The weakness this exposed stands recorded: `TestTheCanaryInventoryMatchesThisAdapter` substring-matches the pin anywhere in `design.md`, so a dated block quoting the new pin turns the guard green while the table it guards stays wrong. Scoping that assertion to the table is unowned. ### 3.9 Cursor (Composer) seam — surveyed live 2026-08-02; the store is open, and it holds credentials **Environment:** Cursor 3.14.7, Windows. Closed source, and the store format is **undocumented and unversioned** — there is no changelog for the 3.12–3.14 line and no schema anywhere. So this is the Claude Code / Antigravity method again: a read-only live survey, every field cross-checked, nothing claimed that was not observed. **Store inventory.** One SQLite database backs every Composer session: ``` %APPDATA%\Cursor\User\ globalStorage\state.vscdb 9.3 MB at survey time globalStorage\state.vscdb-wal 4.6 MB — LIVE, Cursor was running workspaceStorage\\workspace.json workspaceStorage\\state.vscdb 4096 bytes workspaceStorage\\state.vscdb-wal 300–540 KB ``` Three tables. `composerHeaders(composerId, workspaceId, createdAt, lastUpdatedAt, isArchived, isSubagent, recency, checkpointAt, value)` is one row per session, `value` being a JSON blob. `cursorDiskKV(key, value)` is key/value: `composerData:` holds per-session state, `bubbleId:*` and `ofsContent:*` hold the message payloads. **`ItemTable(key, value)` holds the credentials** — see below. **Field map.** | Field | Verdict | Source, and what was measured | |---|---|---| | name | **MEASURED** | `composerHeaders.value.name` — a vendor-GENERATED session title ("Multi-vendor orchestration"), same class as the Claude summaries the HUD already shows. Absent on a session the vendor has not titled yet; the composerId's first eight characters then. | | model | **MEASURED** | `composerData:` → `modelConfig.modelName`. Observed `composer-2.5`, `grok-4.5`, `gpt-5.6-sol`, and the literal `default`. One string, no display name beside it. | | workspace | **MEASURED** | `composerHeaders.workspaceId` → `workspaceStorage//workspace.json` → `.folder`, a `file:///c%3A/...` URI (lower-case drive letter, percent-encoded colon). `value.workspaceIdentifier.uri.fsPath` carries the same path and confirms it. | | context % | **MEASURED** | `composerData.contextUsagePercent`, a float the vendor persists: 37.05, 29.008, 12.38, 10.99 observed, alongside raw `contextTokensUsed`/`contextTokenLimit` (94854/256000, 24763/200000, 44k/1M). The header row's `value.contextUsagePercent` mirrors it and agreed **exactly** on all four rows carrying both. | | cost | **ABSENT** | `usageData` was `{}` in all 8 blobs, and `tokenCount.inputTokens`/`outputTokens` read **0 in all 310 message rows**. The schema is present and never populated. | | quota | **ABSENT** | No consumption record on disk anywhere. What IS there is plan ENTITLEMENT (`credit_dollars: 25`, `included_usage_dollars: 40`) — what the plan grants, not what has been spent. | | last_activity | **MEASURED** | `lastUpdatedAt` / `recency` / `checkpointAt`, epoch **milliseconds**. Not every row has all three (`lastUpdatedAt` was NULL on 4 of 9). | | liveness | **PARTIAL, never observed in flight** | `composerData.status` (`completed`, `aborted`, `none`), `generatingBubbleIds`, `hasBlockingPendingActions`. All read terminal or empty across the corpus; no session was ever sampled mid-generation. | | subagents | **ABSENT (structural only)** | `isSubagent` 0, `numSubComposers` 0, `subComposerIds` `[]` on every row. The fields exist; the observation does not. | **Amended 2026-08-11: a larger survey HARDENED the `cost` row and CORRECTED the reason behind the `quota` row. The corrected reading lived only in §7.16 until now.** The two rows above rest on the 2026-08-02 corpus — 8 `composerData` blobs and 310 message rows. A re-verification on 2026-08-08 (§7.16) read a bigger one and found the same thing harder: `usageData` was `{}` in **19 of 19** blobs, `tokenCount` was zero in **1,622 of 1,622** message rows, and 78 `turn_ended` records and 51 transcripts carried status and no numbers. `cost` **ABSENT** therefore gets stronger, not weaker, and the original counts above stay as the smaller measurement they were. The `quota` row's VERDICT stands and its REASON does not. The row reads `credit_dollars: 25` and `included_usage_dollars: 40` as plan ENTITLEMENT — what the plan grants. The 2026-08-08 survey measured those constants as **Statsig experiment values stamped `is_user_in_experiment:false`**, and they were the only account figures anywhere on that disk. So they describe an experiment this account is not in, not what this account was granted. §7.17's absence table already renders Cursor as `no quota anywhere · its store holds experiment values, not usage` on exactly that measurement. What follows from both rows is the same: nothing about consumption reaches the store as a byproduct of a turn, so the number is FETCHED rather than found. **Cursor Hooks is the real token seam** — the vendor's own documented `afterAgentResponse` step hands a command hook the turn's token counts on stdin, and §7.16 is the record of it, including why print mode's derived `inputTokens` was refused. **Most rows are not sessions.** 9 header rows, of which **5** were: the empty-state draft (`composerId` literally `empty-state-draft`, `isDraft` true), two pre-created composers a new window makes before anyone types (no title, no `lastUpdatedAt`), and two archived threads. Filtering on `isDraft`, `isArchived`, `isSubagent`, `workspaceId == "empty-window"` and the draft sentinel left 4 real sessions, which is what the HUD shows. **Build cautions, recorded before any adapter work:** - **The WAL is where the data is.** This is stronger than the usual "read the sidecar too" (§3.8): every workspace-level `state.vscdb` was **4096 bytes — one empty page — with 300–540 KB in its `-wal`**. A reader that opens only the `.db` there does not get stale data, it gets an empty database. The global store was 9.3 MB with a 4.6 MB live sidecar. Read both as bytes; never open or lock the file Cursor owns. - **THE STORE HOLDS LIVE CREDENTIALS.** `ItemTable` carries `cursorAuth/accessToken`, `cursorAuth/refreshToken`, `mcpOAuth.secret.*` and git-IPC auth tokens; `composerData` blobs carry `blobEncryptionKey` and `speculativeSummarizationEncryptionKey`. This is the first vendor where "read the store" and "read the user's tokens" are the same sentence unless an adapter is explicit about what it will not touch. The allowlist is decisions/007 and it is asserted by a test, not promised in a comment. - **`ItemTable['composer.composerHeaders']` is a legacy JSON mirror and it is STALE.** At survey time it named **3** composers to the table's **9**, and all three were rows the filter drops — so an adapter reading the mirror would report *zero* sessions on a machine running four. Read the table. (The mirror is in `ItemTable` anyway, so the credential rule forbids it independently.) - **Timestamps are mixed.** The header columns are epoch milliseconds; ISO-8601 UTC strings live in the same store at `composerData.fullConversationHeadersOnly[].createdAt` (per-message, structural). That path is a finer-grained activity signal than `lastUpdatedAt` and the adapter deliberately does **not** read it: it is outside the allowlist, and widening the allowlist for precision is exactly the trade this seam should not make. The timestamp reader accepts both encodings anyway, because an unversioned INTEGER column is not promised to stay one. - **One store, many sessions**, so the store's file mtime dates the STORE. Folding it into `last_activity` — which is what §6 Q8 prescribes for every other vendor — would mark every Cursor row live whenever Cursor wrote anything, forever. The Q8 fold runs over the per-row timestamps only, and degrades when none is readable. This is the one deliberate departure from the Q8 shape, and the reason is that the shape assumes one file per session. - **Version fragility.** No changelog, no schema version in the file, no documentation. The adapter addresses columns by NAME (read out of the CREATE statement `sqlite_master` stores) rather than by position, and a store missing `composerHeaders` or one of the columns the field map needs is reported **unreadable with the reason** rather than as zero sessions — a wrong "your agents are idle" is worse than a visible "I cannot read this". **The needs-input seam, for later.** Cursor's *documented* surface is Hooks (cursor.com/docs/hooks): the base payload carries `conversation_id`, `model`, `workspace_roots` and `transcript_path`, and `preCompact` carries context numbers. That is a supported, versioned contract and it is where a liveness/needs-input signal should come from — not from reverse-engineering `status` out of the store. Recorded as the watch item; §8 carries it. Separately, the `cursor-agent` CLI keeps its own store, which is **not installed on this machine** and therefore an unverified surface, out of scope. (That last sentence held until 2026-08-29. `cursor-agent` is installed here now and its session manifest IS read; see this section's 2026-08-29 addendum. The Cursor Hooks watch item above is untouched.) **Adapter built, 2026-08-02 (`internal/adapter/cursor`).** Name, model, workspace and last_activity REPORTED; context % declared DERIVED and marked per read only when the adapter computed it; cost, quota, liveness and subagents `CapNone`. `internal/sqlite` gained two additions rather than being worked around: `Columns` (split the column list out of the stored CREATE statement, so columns are addressed by name) and `Rows` (stream a table to a callback that can stop, so filtering a key/value table by prefix retains nothing). **Live verification, same day:** 5 sessions discovered and read against the real store **with Cursor running** and a 4.6 MB live sidecar; workspaces resolved to `agent-ops` and `faithfulness-judge`; models `grok-4.5`, `composer-2.5`, `gpt-5.6-sol`; context percentages 4.42 / 12.93 / 37.05 all vendor-REPORTED, none marked derived; one session with no `composerData` row rendering an absent model rather than an empty one; every session passed `Validate` with an empty degraded set; no cost, no quota, no sub-agent count anywhere. First `Discover` 78 ms, second 0 ms (the store had not moved). What this does **not** cover, itemized: no in-flight session was sampled, no fan-out was observed, the derived-percentage path did not fire on live data (no real session was missing `contextUsagePercent` while carrying raw counts), and the corpus is one machine, one day, one Cursor version. **Observed 2026-08-17: a live `cursor-agent` CLI session drew no HUD row — and that is the DESIGN, not a defect.** While the operator drove the interactive session that paid §7.16's context capture, no HUD row appeared for it. The reason is structural and already on the record: this adapter reads exactly one store, `%APPDATA%\Cursor\User\globalStorage\state.vscdb`, which is the **IDE's** Composer store, and the `cursor-agent` **CLI keeps its own** (§7.16's three-excluded-fields block says so in as many words, and it is why a `conversation_id` is not stored — it would dangle a join that does not exist). A CLI session was therefore never a row this adapter could draw. The observation confirms an existing claim rather than opening a question, and it is written down because "the gauge showed nothing" is the kind of report that gets re-investigated every few months unless the expected answer is recorded beside it. **This paragraph describes the build it was written at and no longer describes the HUD: a CLI session draws a row as of 2026-08-29. See this section's addendum below.** **What the same observation does NOT settle, and it is the measurement worth taking.** The operator's Cursor configuration carries `ghostMode: true` and `privacyMode: 2`. Neither setting can explain the missing row above, because that row was never expected. They bear on a different and live question: **whether those settings suppress session writes to the IDE Composer store this adapter DOES read.** If they do, then every Cursor row in the HUD is conditional on a vendor privacy setting nobody has measured, and a reader would see an empty Cursor section with no way to tell "no sessions" from "sessions not written". That is a zero-versus-absent failure one layer below the renderer, where §4a.1 cannot reach it. Stated as an open item rather than a finding, because nothing here was measured: the hypothesis is untested, and it must not be repeated as though it were observed. The measurement is cheap and needs the IDE rather than the CLI — drive a Composer session in the IDE with those settings on, then off, and compare what `composerHeaders` gains in each arm. Until somebody runs it, this paragraph is a hypothesis with its reason attached and nothing in the adapter changes. **Amended 2026-08-29 — the CLI store is readable on this machine, and the 2026-08-17 gap is CLOSED. `cursor-agent` CLI sessions draw HUD rows now.** The paragraph above stays correct about the build it described. Its premise moved: `cursor-agent` is installed here, and a 2026-08-29 read-only survey of its trees measured a per-session manifest the 2026-08-02 survey could not have seen. `internal/adapter/cursor` reads it, in `chats.go`, beside the Composer store it already read. **First, the claim that did NOT survive the survey.** A JSONL transcript tree exists at `~/.cursor/projects//agent-transcripts//.jsonl`, and a vendor-seam refresh candidate described it as Claude-compatible JSONL carrying per-turn token counts. Both halves are false, measured structurally over 71 files and 1,951 records with 0 unparseable. The envelope carries exactly three key sets — `message`+`role` (1,852), `status`+`type` (87), `error`+`status`+`type` (12) — and none of `sessionId`, `timestamp`, `cwd` or `version`, so `internal/adapter/claudecode` cannot be reused: its canary is `sessionId`. A key sweep for `usage|tokens|input_tokens|output_tokens|totalTokens|cacheRead|cost` over all 71 files returns zero matches. That CONFIRMS §7.16's 2026-08-08 measurement on a corpus 1.4x larger at a CLI build one week newer, so §7.16's held display is untouched and cost stays `CapNone` for this vendor. The transcript tree is not read and no adapter opens it. **What IS built, and what it measures.** The store is one plain-JSON manifest per session: ``` ~/.cursor/chats///meta.json ``` Observed key sets over 43 manifests, at `cursor-agent 2026.08.11-e8db854`: 40 carry `schemaVersion`, `createdAtMs`, `hasConversation`, `updatedAtMs`, `cwd`; 3 add `title`. `schemaVersion` is `1` on 43 of 43. **This is the first Cursor surface anywhere that declares its own format version**, against the "Version fragility" caution above, and the value is PINNED in `chats.go` and in a fixture. | Field | Verdict | Source, and what was measured | |---|---|---| | name | **MEASURED** | `title`, on 3 of 43 manifests. An absent title is therefore genuine absence, not a failed read. The row then takes the `cwd` basename, which is the fallback `internal/hud`'s own `sessionLabel` applies and `internal/adapter/pi` applies at the adapter. | | workspace | **MEASURED** | `cwd`, a native path (`C:\...` on 43 of 43 here), taken verbatim. It is NOT the `file:///c%3A/...` URI the Composer store's `workspace.json` carries, so no conversion runs. | | last_activity | **MEASURED** | `updatedAtMs`, epoch milliseconds, folded with the manifest's own mtime per §6 Q8. | | model | **ABSENT** | No key of any kind. The Composer store's `composerData` names one; this store does not. | | context %, cost, quota | **ABSENT** | See the token sweep above. Zero matches in 1,951 records. | | liveness | **ABSENT** | No in-flight session was sampled, same as the Composer half. The HUD classifies age. | **The Q8 fold runs here, and that is not a contradiction of the build caution above.** The Composer half deliberately excludes the store's file mtime because ONE file backs every session there. The CLI gives every session its own manifest, so that file's mtime dates that session, and the fold is the ordinary shape every other adapter uses. The two agreed within 0.2 s on 43 of 43 manifests, so the fold is a guard rather than a source: it is what keeps the age honest if the vendor ever rewrites the file without restamping the key. **`hasConversation:false` is an empty shell and draws no row. The ruling is measured.** All three such manifests held nothing but themselves — no `store.db`, no `prompt_history.json` — and their `updatedAtMs` stood 263–387 ms past their own `createdAtMs`. They are the directory the vendor stamps when a session is created and nobody types. That is the same class as the Composer store's `empty-state-draft` and its pre-created composers, so the same filter applies. An ABSENT `hasConversation` key is NOT read as a declared false and keeps its row: a vendor that stopped writing the key must not empty the HUD. **Two more shapes are skipped in silence, and the counts are recorded here so the row count is not read as a defect.** 22 of the 65 session directories hold `store.db` with no `meta.json` at all, because the manifest is newer than the tree; a row for one would date a session from a directory mtime whose meaning was never measured. Together with the 3 shells, 40 of 65 directories drew rows on the survey machine. A manifest whose `schemaVersion` this adapter does not read is the opposite case and is REPORTED every time: it draws no row (dropfile's rule — the keys may no longer mean what the reader thinks) and the skip is counted into a diagnostic that rides on every Cursor row, Composer rows included. A silent skip there would turn a vendor format bump into "you have no CLI sessions", which is a wrong answer rather than a missing one. **One adapter, two stores, and every row says which one it came from.** Both stores describe Cursor sessions, so both feed `model.VendorCursor`; the Adapter composes a sibling reader rather than shipping a second adapter, because a vendor id is what the HUD's identity column and the `--vendor` flag address, and two adapters sharing one id would give the registry two answers. The distinction IS measurable, so it is displayed: every session carries an Extra labelled `source` naming its store, and a CLI row's id is prefixed `cli:`. The labels are symmetric — a Composer row carries one too — so that the ABSENCE of a label never becomes the thing that identifies a store. The alternative was a second vendor id (`cursor-cli`); it was rejected because it buys one distinction and pays a second vendor line, a second `--vendor` value and a second doctor pin for what the operator experiences as one tool. The composition has one honest cost and `chats.go` states it on the row rather than hiding it: `model.Capabilities` is static per adapter, so a CLI row's empty model and context cells read as "absent now" when the truthful reading is "this store has no such field". Each CLI row carries a second Extra, `not in this manifest`, naming those two fields. **The credential rule is unchanged and narrower here.** The tree sits beside `store.db` (4096 bytes with a 300 KB–1.7 MB `-wal`, the WAL trap in its strongest form), `prompt_history.json`, and the config and cache files in `~/.cursor`. The reader opens ONE file name, `meta.json`, at a fixed depth of two directories below `chats/`. It never opens `store.db`, never recurses, and never looks at a sibling for any purpose. `chats_test.go` plants credential-shaped and prompt-shaped markers in those neighbours and asserts none reaches a displayed field, which is the same standing test the Composer reader carries. `conversation-search.db` — a new FTS5 index over conversation `body` text that appeared beside `state.vscdb`, with a `source` column admitting `'cloud-cache'` — stays unopened under the rule that keeps grok's `session_search.sqlite` closed. **What this amendment did NOT measure.** `~/.cursor/acp-sessions//meta.json` carries the same shape (49 manifests, same three key sets, 3 of them shells) and is deliberately NOT read: those are `cursor-agent acp` sessions, which is the server council's own Cursor seat runs (§9.36), and whether telltale should draw rows for its own council seats is a separate question this change does not answer. The `~/.cursor` root is measured on Windows only; see `PARITY.md`. No transcript, no `store.db` blob and no `ItemTable` row was read. ### 3.9a Grok CLI seam — surveyed live 2026-08-09; the first vendor that writes money down **Environment:** grok 1.0.0 (3cd0d0cbce), Windows 11, signed in against grok.com, model `grok-4.5`. Closed source and undocumented — `~/.grok/README.md` is user-facing product prose, not a format spec — so this is the Claude Code / Cursor method again: a read-only live survey over the real store, every claim measured, nothing carried over from `--help`. **Numbered against the sections above rather than after §3.10** so the survey sits with the surveys and the inventory keeps summarizing everything before it. The council seat (§9.39) already drives this vendor, and the HUD could not see it at all. This closes that; the seat's parser and this adapter now read the same wire from two sides, which is what let the cost unit below be pinned rather than assumed. **Store inventory.** Sessions are DIRECTORIES, not files: ``` ~/.grok/ sessions/ session_search.sqlite a FILE at the root — full-text index / prompt_history.jsonl a FILE at workspace level / ONE SESSION summary.json signals.json chat_history.jsonl events.jsonl updates.jsonl prompt_context.json system_prompt.txt resources_state.json rewind_points.jsonl announcement_state.json terminal/ *.lock active_sessions.json auth.json worktrees.db models_cache.json logs/ memtrace/ ``` **The variance is a finding, not noise.** Across 30 session directories in 8 workspaces, only five files were present in all 30: `summary.json`, `chat_history.jsonl`, `events.jsonl`, `prompt_context.json`, `system_prompt.txt`. `updates.jsonl` was in 29, `signals.json` in 23, `resources_state.json` in 13, a `terminal/` subdirectory in 5. The adapter therefore sources every required field from `summary.json` — the invariant — and treats a field that lives anywhere else as *absent now* when its file is missing, which is a first-class state here and not a failure. **Field map.** | Field | Verdict | Source, and what was measured | |---|---|---| | name | **MEASURED** | `summary.json` → `generated_title` (17 of 30), falling back to `session_summary`, which held the identical string on every session carrying both. A headless `--single` run has `session_summary: ""` and **no `generated_title` key at all** — verified by running one — so an unnamed grok row is absence, not a failed read. | | model | **MEASURED** | `summary.json` → `current_model_id`. `grok-4.5` on 30 of 30. Note the per-model usage blocks key off `grok-4.5-build` instead; the adapter reports the id the vendor puts in the session's own model field, not the billing variant. | | workspace | **MEASURED** | `summary.json` → `info.cwd`, absolute and native (`C:\Users\sanle\code\telltale`), on 30 of 30. The parent directory name carries the same path percent-encoded and is the **fallback** — see the round-trip note below. | | context % | **MEASURED** | `signals.json` → `contextWindowUsage`, an integer percentage the vendor computes, beside the raw `contextTokensUsed` / `contextWindowTokens` (500000 on every session). It TRUNCATES: 39656/500000 = 7.93 is written `7`, 22675/500000 = 4.535 is written `4`. The adapter reports the vendor's integer and does not recompute — grok is the second vendor after Cursor with no assumed denominator, and the more precise float would be a number the vendor never said. These three keys are still the source, and `signals.json` holds 61 more beside them — see the 2026-08-29 drift census below. | | cost | **PER-TURN ONLY, so the field stays `CapNone`** | `updates.jsonl` → each `turn_completed` record's `usage.costUsdTicks`. The unit is measured twice over (below). It is **not cumulative**: one session's three turns read 455412000, 820464000, 747416000 ticks, the third smaller than the second. `"[a-z_]*cost[a-z_]*"` over every `.json`/`.jsonl` in the store matched `costUsdTicks` and **nothing else** — no session total exists anywhere. Summing needs every turn record and `updates.jsonl` reached **818 KB** in one session, past any bounded tail; a tail-window sum is a lower bound, and a lower bound in a column headed COST is a derived number wearing a read one's clothes. The **last turn's** cost is carried as a labeled Extra instead. **This row named one key and the record carries eleven** — see the 2026-08-29 re-measure below, which is a correction to this row's completeness and not to its verdict: the sweep spelled here could not match `inputTokens` and its siblings. | | quota | **ABSENT** | `"[a-z_]*(rate\|limit\|quota)[a-z_]*"` over the whole store returned only tool-configuration keys (`output_byte_limit`, `head_limit`). No window, no ordinal, no reset time reaches disk. That is a statement about the *disk*; the network half of the same question is measured and closed separately below. | | last_activity | **MEASURED** | `summary.json` → `last_active_at` (29 of 30) then `updated_at` (30 of 30) then `created_at`, folded with the file's mtime per §6 Q8. `summary.json` is rewritten every turn, which is also why it — not the session DIRECTORY, whose mtime moves only when an entry is added — is the freshness hint `Discover` returns. | | liveness | **ABSENT, and this one was probed rather than reasoned about** | See below. | | sub-agent count | **ABSENT** | grok ships a `spawn_subagent` tool (it is in the tool list every headless run prints), and `"subagent[A-Za-z_]*"` matched nothing on disk outside the system prompt's own description of it. No nest, no count, no parent link. Same ruling as Codex (§3.3): declaring the field and emitting zero would assert something the format cannot check. **The sweep spelled here read file contents and a `subagents/` DIRECTORY exists — see the 2026-08-29 drift census below.** The verdict is unchanged and its stated reason is not. | **The cost unit, pinned twice.** `costUsdTicks` is fixed-point USD at **1e10 ticks to the dollar**, and neither half of that was inferred from the name. First, grok's headless wire prints both forms of the same number on its `end` event, and three live runs on this box on 2026-08-09 gave `0.0306488 / 306488000`, `0.0315248 / 315248000` and `0.0382104 / 382104000` — exactly 1e10, three times. Second, the disk field is spelled differently (`costUsdTicks` vs the wire's `total_cost_usd_ticks`), so it had to be shown to be the same quantity and not merely a similarly named one: those three runs' on-disk values were read back and matched the wire's tick counts **value for value**. Without the second step this would be a plausible unit rather than a measured one. **`active_sessions.json` claims liveness and does not deliver it — measured.** With 30 sessions in the store the file held the two bytes `[]`. It still held `[]` **while a headless turn was mid-flight**, sampled with `grok.exe` confirmed running by PID and with the file's own mtime freshly stamped by that run — so the vendor had written the file and written nothing in it. A registry that is empty during a live session cannot tell "nothing is running" from "the thing running is not the kind it tracks", and §4a.4 already names process-existence as the one case where an adapter can lie to the HUD undetectably. Not read. `events.jsonl` carries the other tempting signal — `phase_changed` (1765 of them in one session, spelling `waiting_for_model`, `streaming_reasoning`, `streaming_text`, `tool_execution`, `permission_prompt`), `turn_started` / `turn_ended` with an `outcome`, and a `permission_requested` that is a genuine needs-input state. It is left unread in v1 for the reason the corpus itself demonstrates: the newest session ends on an **unresolved `permission_requested`** written minutes before grok exited, so "the last event is a prompt" and "a dead session was killed at a prompt" are the same bytes. A hint that stays true forever after the process is gone is worse than no hint. This is the strongest needs-input seam any vendor has offered so far and it is recorded here as the watch item, not spent. **The percent-encoding round-trips, and that is why this adapter may decode it.** `C%3A%5CUsers%5Csanle%5Ccode%5Ctelltale` is `C:\Users\sanle\code\telltale`: `:` as `%3A`, `\` as `%5C`, with letters, digits and a literal `-` passing through unescaped. Every one of the 8 workspace directory names decoded to exactly the `info.cwd` its sessions recorded, drive-letter case included. That is the opposite of §3.1's ruling for Claude Code, whose project slug maps both `\` and a literal `-` onto `-` and is therefore lossy — decoding *that* would invent a path. Grok's encoding is injective, so decoding it invents nothing, and the adapter uses it as the workspace fallback. It stays a fallback: the vendor's own record of its cwd outranks a key we reconstructed. **What is deliberately not opened, and why it is stated rather than assumed:** - **`~/.grok/auth.json`** holds the OAuth token. The adapter resolves no path outside the sessions tree at all. - **`sessions/session_search.sqlite`** is an FTS5 index whose `session_docs` table carries `(session_id, cwd, updated_at, title, content, content_hash)` — **`content` is transcript text**. It would answer "what is this session about" in one query, and that is exactly the trade this repo does not make: `summary.json` already has the vendor's own label. - **`prompt_context.json`** inlines the user's `CLAUDE.md`/`AGENTS.md` verbatim (39 KB on one session). Same rule. A test plants a marker in both files and asserts nothing carrying it reaches any displayable field, the way §3.9's credential allowlist is enforced. - **The `.lock` sidecars** beside `summary.json`, `chat_history.jsonl`, `updates.jsonl` and `rewind_points.jsonl` are never opened and never created. The gauges read. **The quota question has a network half, and it is closed three times over — probed 2026-08-09**, the same local day as the survey above (the HTTP `Date` headers quoted below read 2026-08-10 UTC; this box runs UTC−4). The table's `quota` row says nothing reaches *disk*, which invites the obvious follow-up: the CLI talks to a server, so ask the server. `~/.grok/README.md` ("Using auth.json for API Access") even documents the call. This block exists so the next person to ask stops here instead of re-deriving it. Three findings, then three rules. *The documented recipe does not run on this build.* Its `jq` path `."https://accounts.x.ai/sign-in".key` matches nothing in this box's `auth.json`, whose one entry is keyed `https://auth.x.ai::` (`auth_mode: "oidc"`, a six-hour `expires_at`); sent as written it returns **401**, `www-authenticate: … reason=no auth context`. And `POST /v1/chat/completions` gates on a header the recipe never mentions — omit it and the proxy answers **426 Upgrade Required**, body `"Your Grok CLI version (none) is outdated"`. The header is `x-grok-client-version`, read out of `grok.exe`'s string table beside `x-grok-client-identifier`, `x-grok-client-mode`, `x-grok-session-id` and the documented `x-grok-model-override`. The README is product prose here too, exactly as this section's header warns. *The free half of the question is answered, and the answer is no.* `GET /v1/models` — the same request whose result the CLI already caches to `models_cache.json`, so it bills nothing — returns **200** carrying `etag`, `strict-transport-security`, `cf-cache-status`, `CF-RAY`, `alt-svc` and `Server: cloudflare`, and **not one `x-ratelimit-*` header of any kind**. The one proxy response this vendor already persists has no quota in it. *The billed half is `not checked`, which per §9.42 carries a reason and never a value.* Whether `POST /v1/chat/completions` returns those headers was **not measured**: it is the only probe here that spends a turn, and it stopped at this machine's own tool-permission boundary rather than being run. Two attempts to get it from the vendor's own logging failed for a reason worth recording, because it looks like evidence and is not: `~/.grok/logs/ unified.jsonl` is **not an HTTP log** — every record's `src` is `shell` or `grok-pager` — so its silence about rate-limit headers was never evidence about the wire, and a live `grok -p` run driven with `--debug-file` produced no file at all. *The disk sweep was re-run with a case gap closed.* The original survey's `"[a-z_]*(rate|limit|quota)[a-z_]*"` is **lowercase-only** and would have missed a camelCase `rateLimit` — which matters precisely here, because this is the vendor that writes `contextWindowUsage` and `contextTokensUsed`, so camelCase is its house style and the original regex had a live blind spot. Re-swept as `"[A-Za-z_-]*([Rr]ate[Ll]imit|[Qq]uota|[Rr]emaining)[A-Za-z_-]*"` over the whole sessions store, including a session directory created by a fresh live turn: the sole hit is `agents_remaining` (17 occurrences), a **sub-agent budget**, not an account one. The `quota: ABSENT` verdict survives the stronger sweep. *And the vendor's own monitoring surface settles it.* `docs/user-guide/24-monitoring-usage.md` documents an external OpenTelemetry stream — the one place grok is *designed* to report what an account is doing — and its attribute keys are **a closed enum**, with an export-time validator that drops any record carrying a key outside it. What it carries is `grok_code.token.usage` (`input` / `output` / `reasoning` / `cache_read`, by model) and a `grok_code.api_request` event with the same four counts. What it carries **nowhere** is a window, a reset, a remaining percentage or a limit of any kind; `subscription tier` sits on its explicit never-exported list, and the doc states outright that there is no cost metric ("join `grok_code.token.usage` with your own price sheet"). So the absence is not an oversight in a session file. The vendor built a schema for exactly this question, put **spend** in it, and put **no quota in it at all** — which is §7.15 versus §7.16 in the vendor's own hand: a count with none, never a reading against a limit. That stream is also the honest answer to "then what *could* be measured". If a grok **spend** row is ever wanted it is the seam — exported to a local collector whose output telltale would read as a *file*, leaving §4a.5's no-network-calls contract intact, since the push is grok's and not ours. It is the cursor relay's shape (§7.16), and it is emphatically not a quota window. Cost: a double opt-in, an endpoint, and a collector to run; the `console` exporter is no shortcut because the doc says it is suppressed in the `agent` and `headless` entrypoints. Recorded as the seam, not spent. **The seam was spent on 2026-08-10 — `telltale otel grok` is the collector (§7.16a).** Before it was built, the doc's schema table was checked against the wire the way this section's header demands: a dump collector on 127.0.0.1:4318, and a live headless `grok -p "hi"` from grok 1.0.0 (3cd0d0cbce) with the double opt-in set. What arrived, versus what the doc says: - Transport as documented: OTLP http/protobuf POSTs to `/v1/logs` and `/v1/metrics`, `Content-Type: application/x-protobuf`, uncompressed, from `OTel-OTLP-Exporter-Rust/0.32.0`. The batches flushed before the headless process exited, at the default export intervals — a short `-p` run loses nothing. The fleet-policy startup suppression the doc warns about was not observed to delay anything on this signed-in box. - `grok_code.api_request` events carry `input_tokens`, `output_tokens`, `reasoning_tokens` and `cache_read_tokens` as int attributes — all four on one record — beside `model`, `duration_ms`, `stop_reason`, `session.id`, a per-session monotonic `event.sequence`, `user.id` and `team.id`. The event name arrives in the LogRecord's `event_name` field, not as an attribute. - The `grok_code.token.usage` metric (delta temporality, by `model` and `type`) carried **the same four counts value-for-value** as the same turn's api_request event: 20323/56/42/2560 on both sides of one capture. One number, two envelopes. - `turn_completed` **on the stream** carries outcome and duration and **no token counts**, as the table says. The qualifier was added 2026-08-29 and it matters: the record of the same name that grok persists to `updates.jsonl` carries nine of them, which is the re-measure block below. One event name, two envelopes, opposite contents. - One departure from the doc's letter: `OTEL_METRICS_INCLUDE_SESSION_ID` defaults on, but the token.usage data points carried no `session.id` — only `session.count`'s did. Nothing here reads metrics, so nothing turns on it; recorded because the doc says otherwise. The quota verdict above is unchanged by any of this: nothing on the stream carries a window, a reset or a limit. What was added is a **spend count** (§7.16's vocabulary — a count with no denominator), accumulated per api_request event into `~/.telltale/usage/grok.json`, display held. §7.16a is the design record. A measurement of the completions headers would not move the verdict, because three separate rules already close this and none of them turns on whether the number exists: 1. it needs `auth.json`, and this adapter resolves no path outside the sessions tree (above); 2. it needs a network call, and §4a.5's adapter contract is explicit that implementations "must not write to vendor state, and must not make network calls or read credentials"; 3. §9.42 draws the probing line at **cost and side effect** — `doctor` may run `--version` precisely because it starts no turn and bills nothing. A quota probe against the completions endpoint would spend from the very pool it is trying to read. And the quantity would be the wrong one regardless. `x-ratelimit-remaining-requests`, `-remaining-tokens` and `x-ratelimit-reset-requests` are documented for **`api.x.ai`**, the metered developer API. This CLI rides `cli-chat-proxy.grok.com` on an OIDC session token against a **SuperGrok subscription**, whose binding limit is a shared weekly pool surfaced only in grok.com's own Settings → Usage. An RPS/TPM header is not the window the usage pane means by quota: §4a.3's window carries a *length* and a `ResetsAt`, and a per-second request cap has neither. Relabelling one as the other would be a duration claim with no source — the exact move §4a.3 already forbids — and would land a real number on screen answering a question nobody asked. **Two smaller findings worth writing down.** One session's `summary.json` carries `"sandbox_profile": "bogus-profile-xyz"` — the invalid profile §9.39 fed the CLI to prove `--sandbox` validates nothing. grok not only accepted it, it **persisted it**, which is why that key is not rendered as an Extra: it would put an unvalidated word on screen. And `summary.json` carries a git block (`git_root_dir`, `git_remotes`, `head_commit`, `head_branch`) on exactly the one session whose cwd was a git repo — a branch column is available for later, and is out of scope here. **The frame, generated by the build** (`internal/hud/testdata/golden/grok-row.txt`, at 120 columns; both rows are synthesized). The COST column is not narrow here, it is ABSENT — the vendor writes dollars and no row claims a session total. The first row's bar carries no estimate marker because the percentage was read rather than derived; the second is a headless run with no title and no `signals.json`, so its label falls back to the workspace and its CONTEXT is an em dash rather than a zero: ``` telltale │ 2 sessions │ grok 2 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL CONTEXT AGE ● GR │ Adapter Field Map Review C:\src\code grok-4.5 ▊─────────── 7% │ 20s ◐ GR │ example-app C:\src\code grok-4.5 — │ 6m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ``` **Adapter built, same day (`internal/adapter/grok`).** Name, model, workspace, context % and last_activity REPORTED; cost, quota, liveness and sub-agent count `CapNone`; the derived set deliberately EMPTY. `summary.json` and `signals.json` are read whole under a 64 KB cap each (largest observed: 918 and 1538 bytes) and a file past its cap degrades the fields it feeds rather than being slurped; `updates.jsonl` is tail-read like every other JSONL here. **Live verification, same day** (`go test ./internal/adapter/grok -tags=live -run TestLiveGrokStore`, which drives the HUD's own `Scan` and `Render` over the real store and is excluded from CI because CI has no store to read): **31 sessions discovered and read**, **31 of 31** sourcing both a model and a workspace, 18 titled, 24 carrying a context reading, 24 carrying a turn cost. Every session passed `Validate`; not one produced a cost, a quota window, a sub-agent count or a liveness hint. The frame showed the zero-vs-absent distinction doing real work: the 7 sessions with no `signals.json` rendered an em dash in CONTEXT while their neighbours rendered an unmarked bar, on the same screen. First `Discover` plus 31 `Read`s completed in 80 ms. What this does **not** cover, itemized: no session was sampled while a turn was streaming (the `active_sessions.json` probe ran against a headless turn, not against a HUD scan); no interactive TUI session was running during a scan, so nothing exercised the NTFS-deferred mtime the Q8 fold exists for; the summary-past-its-cap path has never fired on real data; the corpus is one machine, one day, one grok version; and `active_sessions.json` has never been observed non-empty **at all**, so "it does not track headless sessions" is the most this survey can say about what it would otherwise contain. #### `active_sessions.json` re-measured 2026-08-15 — the registry populates now, and liveness stays `CapNone` **Environment:** grok 1.0.4 (d846eb93d9), the same Windows 11 box, model `grok-4.6`, cwd `C:\Users\sanle`. The paragraph above is the reason this ran. The 2026-08-09 survey never saw the file hold anything, so it could not say what a populated registry would mean. The file holds records now, and a record with a **dead pid** was observed surviving overnight. That one fact makes the question answerable, so it was measured rather than argued. The measurement cost **3 grok model turns**; every other step below is a process start, which calls no model. **The record shape, from two sides.** The live file holds a JSON array of objects with four keys, and the binary at this pinned build names the same four: `struct ActiveSession` with `session_id`, `pid`, `cwd`, `opened_at`, written by `crates/codegen/xai-grok-active-sessions/src/lib.rs` through `active_sessions.lock`, `active_sessions.json.tmp` and `active_sessions.json`. A live record read verbatim: ```json [ { "session_id": "01a0084f-e3b1-7ff2-be5c-2afaed6a6581", "pid": 17096, "cwd": "C:\\Users\\sanle", "opened_at": "2026-08-16T02:04:08.999641600Z" } ] ``` `session_id` is the session directory's own UUID, so the record joins to §3.9a's store directly. `pid` is the grok.exe process, confirmed against `tasklist` at the moment of capture. `cwd` is the workspace in native absolute form, and it equals the `info.cwd` the same session writes. `opened_at` is UTC RFC3339 and stamps the OPEN, not the process start: a resumed session gets a new `opened_at` and the same `session_id`. The vendor also names the failure mode itself, verbatim from the binary: `active_sessions.json is corrupted, starting with empty list`. **The four questions, each measured.** | Question | Answer | Evidence | |---|---|---| | Does a session start gain an entry, and is the pid real? | **Only when a session OPENS, and the pid is real.** | A fresh interactive session registered 0.7 s and 1.3 s after the first prompt was submitted, on two runs, and never before it. `pid 17096 -> ALIVE as grok.exe` at capture. A `--resume` registered 1.6 s after launch with **no prompt at all**, which is how the rest of this block was measured for free. | | Does a CLEAN exit reap the entry? | **Yes, immediately.** | `/exit` in the TUI returned the file to the two bytes `[]` within 0.3 s of the command, with its mtime freshly stamped. The process removes its own record on the way out. | | Does a KILLED process leave the entry behind? | **Yes, and the dead pid persists.** | `taskkill /T /F` on the pid the registry itself named left the record byte-identical, mtime unmoved, for the whole sampling window. `pid 17096 -> DEAD` while the record still claimed it. Reproduced twice, and it confirms the overnight observation deliberately. | | What are the exact fields and their meanings? | The four above. | Live file plus the pinned binary's own struct, which agree. | **The lifecycle, stated as write points.** Four events touch the file, and only these four were observed to. A session OPEN appends its record. A clean exit removes that session's record. Any grok agent start **rewrites the file and drops records whose pid is dead**. A kill removes nothing, so the record outlives the process until the next start sweeps it. **The sweep is pid-aware, not a truncation, and that was measured separately.** One stale record and one live session were put on disk together, then a second grok process was started. The stale record disappeared and the live record survived unchanged, so the sweep reads each pid rather than clearing the list. This also explains the overnight persistence with no contradiction: nothing swept the record because no grok agent started in between. **What does NOT register, measured one shape at a time.** Each of these ran with the file sampled every 0.3 s throughout, and each left it at `[]` while the process was alive: `grok -p ` (a full headless turn, which is the 2026-08-09 result reproduced at 1.0.4), `grok agent stdio`, `grok agent leader`, and an interactive TUI sitting at its input box with a session directory already created and no prompt yet sent. Every one of those starts moved the file's mtime, so the vendor wrote the file and put nothing in it, exactly as the original survey found. `grok --version` does not touch the file at all. **Verdict: liveness stays `CapNone`, and §4a.4's ruling is not falsified.** A populated registry is not a liveness source, because both directions fail: - **Presence proves nothing.** A record can name a dead pid for an unbounded time. The bound is the next grok agent start, which is an event telltale cannot predict and must not cause. An adapter that read presence would report a session as live all night. - **Absence proves nothing.** Four live session shapes register nothing at all, and a headless turn is one of them. An adapter that read absence would report a running turn as gone. §4a.4 rules process-existence out for every vendor, and it named Claude's registry only as the first instance. A grok registry that populates is a different fact from a grok registry that is trustworthy, and only the second one would touch that ruling. It does not. **One usable shape falls out, and it is recorded rather than spent.** The registry supplies a `session_id` to `pid` binding, and a dead pid is checkable directly. So a stale `permission_requested` in `events.jsonl` (the needs-input seam §3.9a already recorded and declined) **can be NEGATED by a dead pid, and can never be ASSERTED by a live one**. The asymmetry is the whole content of the finding, and it survives its own caveats in one direction only. A recycled pid reads ALIVE, which withholds a negation rather than inventing one, so the failure is conservative. A session that never registered cannot be negated, which costs coverage and claims nothing false. Nothing is built on this. The measured-silence advisory that would consume it sits on the post-demo shelf by owner ruling, and this block exists so that work starts from a measurement instead of a memory. **What this does not cover.** Every registering session ran in one cwd on one box at one build. No two sessions were ever registered at once, so the multi-record ordering is unmeasured. `active_sessions.lock` was never opened and `active_sessions.json.tmp` was never observed mid-write, so the write is atomic by the vendor's own naming rather than by observation here. The adapter reads none of these bytes, and this measurement did not change that. #### The `usage` object re-measured 2026-08-29 — the token counts were always beside the cost, and this section's own sweep could not have seen them **Environment:** grok 1.0.5 (5115b46bc9) [stable], the same Windows 11 box, model `grok-4.6`. One billed headless turn (`grok --output-format streaming-json --single=…`) from a fresh empty workspace, with the session's `updates.jsonl` read back and cross-checked against the same turn's wire output. **This block corrects the section above rather than reporting vendor drift, and that is the point of it: the miss was telltale's instrument, not the vendor's format.** **What the cost row says, and what the record actually carries.** The row names `usage.costUsdTicks` and nothing else, because the sweep behind it was `"[a-z_]*cost[a-z_]*"` — lowercase, and anchored on the substring `cost`. A key spelled `inputTokens` matches neither half. The regex answered the question it was asked and was **structurally incapable** of the question a reader takes the row to answer. The `turn_completed` record's `usage` object carries **eleven keys** — nine scalar counts, a per-model breakdown of the same nine, and a turn count — read verbatim off the measured turn: ```json "usage":{"inputTokens":22772,"outputTokens":27,"totalTokens":22799, "cachedReadTokens":256,"cacheCreationTokens":0,"reasoningTokens":22, "modelCalls":1,"apiDurationMs":3481,"costUsdTicks":77047400, "modelUsage":{"grok-4.6-build":{ the same nine scalars, per model }}, "numTurns":1} ``` **It is not drift, and that was checked rather than asserted.** A session directory written 2026-08-09 at 09:35 local — the same day, the same build (grok 1.0.0) and the same model (`grok-4.5`) this section surveyed — carries the identical key set, `numTurns` included: `"inputTokens":22222, "outputTokens":37, "totalTokens":22259, "cachedReadTokens":128, "cacheCreationTokens":0, "reasoningTokens":29, "modelCalls":1, "apiDurationMs":2605, "costUsdTicks":444484000`. The counts were on disk before the adapter was written. The vendor changed nothing; the survey looked with an instrument that could match one key and reported one key. **The wire and the disk spell `input` differently, and the difference is exactly the cache read.** The same turn's `end` event on `--output-format streaming-json` reports `input_tokens: 22516`, `cache_read_input_tokens: 256`, `output_tokens: 27`, `reasoning_tokens: 22`, `total_tokens: 22799`. The disk reports `inputTokens: 22772`, which is 22516 + 256. Both envelopes agree on 22799 and reach it by different arithmetic: **the wire's `input_tokens` EXCLUDES the cache read; the disk's `inputTokens` INCLUDES it.** So `inputTokens + cachedReadTokens` on disk double-counts, and the two seams' "input" columns are not the same statistic. §7.16a's collector reads the OTLP wire, whose `input_tokens` follows the wire spelling — anything that ever holds both must convert rather than add. **`costUsdTicks` re-pins at 1e10 on the same record.** The wire printed `total_cost_usd: 0.00770474` beside `total_cost_usd_ticks: 77047400`, and the disk record carries `costUsdTicks: 77047400` for that turn. The 2026-08-09 unit measurement holds at 1.0.5, value for value, on a fourth run. **The counts are per-turn and they do not accumulate — measured on a two-turn session.** Its two `turn_completed` records read `inputTokens 21548 / totalTokens 21588` then `inputTokens 21958 / totalTokens 21999`. The second is not the first plus anything; each record describes one API call, and `inputTokens` grows between turns because the context is resent, not because a counter advanced. `numTurns` is `1` on **every** record, both turns included, so it counts the turn the record describes and is not a session ordinal. That is the same shape as `costUsdTicks`, measured again in a second unit. **What this changes, and what it does not.** - **The cost row's verdict stands unchanged, and it now covers nine counts instead of one.** A session total needs every record, `updates.jsonl` reached 818 KB in one observed session, and a tail-window sum is a lower bound. A lower bound in a TOKENS column is the same derived number in a different unit. - **The seam map is what moves.** §7.16a opens with grok's OTLP stream as the vendor's one designed-for-reporting surface, and that sentence is still true about *designed reporting*. It is no longer true that the stream is the only place telltale **could** read grok's per-turn counts. The disk is a second source, and a passive one: no collector running, no double opt-in, no port held. - **That second source creates a constraint, recorded here before anything is built on it.** §7.16a's replay guard exists because one number arriving twice must be counted once, and the same rule already chose the event stream over the redundant metric. A disk reader would be a **third** envelope for the same turn, and the guard cannot see it: the collector's guard lives in collector memory and keys on `(session.id, event.sequence)`, while a disk reader would key on a file offset. **A disk reader and the OTLP listener must never both count one turn into `~/.telltale/usage/grok.json`.** One envelope per turn is the rule. Which envelope is the design question, and this block does not answer it. **Nothing is built on this, deliberately.** The DISPLAY of grok spend is HELD by owner ruling (§7.16's amendment, applied in §7.16a), so a reader has no consumer to feed, and a reader design would first have to settle the one-envelope question above, the read-budget question the cost row raises, and the `input` spelling. This block exists so that work starts from a measurement instead of from a row that could not see it. #### Session-file drift, censused 2026-08-29 at grok 1.0.5 — `signals.json` is 64 keys, and there is a `subagents/` directory Read-only census over the whole sessions tree: **108 session directories in 27 workspaces**, against the 30 in 8 this section surveyed. Nothing below changes a verdict, and two entries say the *description* above is narrower than the store. **The inventory's invariant holds.** `summary.json` is present on 108 of 108, which is the one file every required field is sourced from. `chat_history.jsonl`, `events.jsonl`, `prompt_context.json` and `system_prompt.txt` are on 107, `updates.jsonl` on 105, `signals.json` on 95, `resources_state.json` on 73. **`signals.json` carries 64 keys, uniform across all 95 files that have one** — measured by parsing every one and comparing key sets, which returned a single set. The context row above names three of them (`contextWindowUsage`, `contextTokensUsed`, `contextWindowTokens`); all three are still present and still the adapter's source, so no code is wrong. The description is: a fixture synthesized to a three-key file models a file that does not exist. The other 61 are session statistics — `turnCount`, `toolCallCount`, `errorCount`, `compactionCount`, `totalTokensBeforeCompaction`, `sessionDurationSeconds`, `modelsUsed`, `primaryModelId`, plus latency and line-count blocks. **No cost key and no quota key is among the 64**, so the quota verdict survives one more sweep. **Fifteen entry names the inventory above does not list**, with the count of session directories carrying each: `title_refresh_idx` (19), `last_recap_main_turn` (11), `recap_requests` (11), `hunk_records.jsonl` (9), `subagents` (6), `web_fetch` (5), `compaction` (2), `compaction_checkpoints` (2), `compaction_requests` (2), `mcp` (2), `plan.json` (2), `plan_mode.json` (2), `outputs` (1), `plan.md` (1), `workflows` (1). The adapter reads none of them. They are recorded because a presence-census fixture built from this section's ten names would call every one of them unexpected. **One of the fifteen contradicts a verdict's REASON, and it is the cost regex's mistake a second time.** The field map rules sub-agent count ABSENT because `"subagent[A-Za-z_]*"` matched nothing on disk outside the system prompt's own tool description. That sweep read file *contents*, so it could not match a **directory name** — and `subagents/` is a directory. Six sessions carry one; the largest holds 18 child directories, each named with a UUID, each holding `meta.json` and `output.json`. `meta.json`'s key set, read as keys and not values, is `subagent_id`, `subagent_type`, `parent_session_id`, `child_session_id`, `child_cwd`, `status`, `turns`, `tool_calls`, `duration_ms`, `started_at`, `completed_at`, `effective_model_id`, `effective_context_source`, `description`, `prompt`. **This is not drift either**: the largest `subagents/` tree is dated 2026-08-09, the day of this survey. **The verdict is NOT changed here, and the reason is a measurement nobody has taken.** `CapNone` on sub-agent count refuses to assert "this session is running no sub-agents", and a directory that is present on 6 of 108 sessions does not yet establish that its absence means zero rather than not-yet-written — which is §4a.1's zero-versus-absent question, and answering it needs a live run that spawns one and watches the directory appear. Two further facts a future lane must start from: every child UUID observed is **also a top-level session directory** in the same workspace, so the HUD already lists sub-agents as independent rows and a count would need to decide whether they stay listed; and `meta.json`'s `description` and `prompt` are user content, so they fall under the same rule that keeps `prompt_context.json` closed. Recorded as the seam, not spent. ### 3.9b Pi seam — source-surveyed then live-verified 2026-08-16, `pi 0.84.1` **Environment and evidence class.** This section started in the §3.2 class (researched from source, NOT live-verified) and has since been measured against a live corpus on the Windows reference box. The heading and this paragraph said "not installed here, nothing live-verified" until 2026-08-16; that is no longer true, and the verdict table below carries what the bytes actually said. The source read came first, and it is still the reason the shape was known before any file was opened. That read was the writer's own code — `packages/coding-agent/src/core/session-manager.ts`, `src/config.ts`, and `packages/ai/src/types.ts` — in **`earendil-works/pi`**, default branch on 2026-08-16, newest release `v0.84.2` (2026-08-14). Both former names (`badlogic/pi-mono`, `earendil-works/pi-mono`) redirect there. The npm package `@mariozechner/pi-coding-agent` stopped at 0.73.1 in May 2026; current distribution is the `@earendil-works` scope and compiled Bun binaries. The §3.7 lesson predicted that the first live pass would falsify something here. It did: `responseModel` (see the verdict table) does not exist on any assistant message in the live corpus. That is what the prediction was for, and the falsified claim is corrected in place rather than deleted. **Why the survey ran when it did.** Pi had no HUD coverage at the time, and §9.1 already rejected it as a council *seat* under the harness-re-host class. This section is the adapter half: "we looked", so a future session does not start from "nobody looked". The coverage gap it names is closed; the section is kept as the evidence the adapter stands on, not as an open question. **Store inventory (from source, not from a listing).** ``` ~/.pi/agent/ sessions/ ----/ one directory per workspace: the cwd with leading separator dropped and / \ : each replaced by "-" _.jsonl one session; ts = ISO timestamp with : and . -> "-" auth.json provider API keys / OAuth tokens — NEVER read settings.json models.json themes/ tools/ prompts/ bin/ ``` Two relocation facts an adapter must pin: `PI_CODING_AGENT_DIR` and `PI_CODING_AGENT_SESSION_DIR` override the paths, and the package's `piConfig` (`name`, `configDir`) can rebrand the whole product — a rebranded build stores under a different dot-directory entirely. The adapter would pin the default `.pi` and say so. **Format.** JSONL, session `version` 3 (`CURRENT_SESSION_VERSION`). The first record is a header: `{type:"session", version, id, timestamp, cwd, parentSession?}`. Every other record carries `{type, id, parentId, timestamp}` — **entries form a tree, not a list**. Pi branches within a session, so conversation order is a `parentId` walk from the leaf, not file order. Entry types: `message`, `thinking_level_change`, `model_change`, `compaction`, `branch_summary`, `custom`, `custom_message`, `label`, `session_info`. One write-path quirk: a session is not flushed to disk until it holds an assistant message, so an abandoned prompt may never produce a file. **Field verdicts.** The prospect column this table shipped with is now a verdict column, because the live pass ran. The corpus is small and its size is part of every verdict: **4 sessions, 689 records, 197 assistant messages, `pi 0.84.1`.** A verdict of "present" means the field was found on every record that should carry it; a verdict of "absent" means a grep over that corpus returned zero matches, which is the §3.9a standard and not the same as "the source did not mention it". | Field | Verdict (live, `pi 0.84.1`) | Source in the format | |---|---|---| | name | **absent in this corpus** — no `session_info` entry appears in any of the 4 sessions, so the fallback is the operative path, not the exception | `session_info` entry's optional `name` (user-set). Absent is a state; the workspace-basename fallback (§3.7 precedent) applies. | | model | **present** — 5 `model_change` entries; `model` on 197/197 assistant messages. `responseModel` is on **0/197**: the source read claimed it and the corpus falsifies it | `model_change` entries carry `provider` + `modelId`; assistant messages also carry `model`. | | workspace | **present** — `cwd` on 4/4 headers | header `cwd`, verbatim. The directory slug is the fallback, as in §3.7. | | last_activity | **present** — `timestamp` on 685/685 non-header records | newest entry `timestamp` (ISO), folded with mtime per §6 Q8; assistant and toolResult messages also carry unix-ms timestamps. | | tokens | **present** — `usage` on 197/197 assistant messages, with all six documented keys | **every assistant message writes `usage`**: `input`, `output`, `cacheRead`, `cacheWrite`, optional `reasoning`, `totalTokens`. | | cost | **present per message, absent per session**, with the §3.9a ruling attached — `usage.cost` on 197/197, and no session total anywhere | the second vendor that writes money down, and it writes **dollars, per message**: `usage.cost {input, output, cacheRead, cacheWrite, total}`. No session total exists on disk — the TUI sums in memory (`usage-totals.ts`). A summed total is a derived number wearing a read one's clothes (§3.9a), and the tree sharpens it: all-entries and active-path totals differ. The honest carry is the last message's `cost.total` as a labeled Extra, grok-style. | | context % | **CapNone, now measured** — a grep for `contextWindow`/`maxTokens`/`context_window` over the corpus returns zero, and this box's `~/.pi/agent/models-store.json` is an empty object, so the denominator is in neither place | a numerator prospect exists (last assistant `usage` occupancy) but the denominator is nowhere in the session file — it lives in Pi's shipped model catalog, which is the §3.8 1048576 trap again. | | quota | **CapNone, measured absent** — the §3.9a grep now has its corpus: `quota`/`rateLimit`/`resetsAt`/`plan`/`subscription` return zero matches | Pi is bring-your-own-key, so structural absence is expected, and the corpus confirms it rather than assuming it. | | sub-agents | **absent in this corpus** — `parentSession` on 0/4 headers, so the link was never exercised here; the field's existence is still a source claim only | header `parentSession` is a child→parent link (the inverse of agy's structural nesting) — countable only by scanning siblings for parents. | | liveness | **CapNone, measured absent** — a grep for `pid`/`lock`/`heartbeat`/`isActive` returns zero | nothing in the session files, matching the source read. | **What this corpus does not cover.** Four sessions from one box, one operator, one day (2026-08-11), all on the default `.pi` directory. Five of the nine documented entry types (`compaction`, `branch_summary`, `custom`, `custom_message`, `label`) never appear, so their shapes remain source claims. No session was observed mid-write, and no two Pi sessions ran at once, so the tree's branching behavior is documented but not watched. A `session_info` entry and a `parentSession` header are the two things a wider corpus would most likely add. **The credential boundary is cleaner than Cursor's, and the live pass confirms it.** `auth.json` is a sibling of `sessions/`, not inside the session files — an adapter that reads only `sessions/**` never opens a credential-bearing file. The directory listing on the reference box matches: `~/.pi/agent/` holds `auth.json`, `settings.json`, `models-store.json`, `AGENTS.md`, `extensions/` and `sessions/` side by side, so the allowlist is a path prefix rather than Cursor's row-by-row filtering of one shared SQLite file. A grep for `apiKey`/`accessToken`/`refreshToken`/`authorization` over the session corpus returns zero matches. That lowers the read-allowlist burden; it does not remove the planted-marker test, because session *content* is still untrusted and a zero today is a measurement of this corpus, not a guarantee about the format. **The extension seam is the part no other vendor has.** Pi's product thesis is an in-process TypeScript extension system. That means a *Pi extension* could write telltale's relay files (§7.15/§7.16 shapes) directly, with vendor-computed numbers, no adapter parse at all — which would make Pi the first external writer of the relay contract rather than the sixth in-tree adapter. Which of the two paths to build is an owner decision deferred to post-launch demand; this survey only records that both are open and neither is blocked by the format. **Where this leaves the two Pi questions.** They are separate questions with separate answers, and this section is the reason they can be answered separately. The **HUD** question is settled: `internal/adapter/pi` is the in-tree observer, and the verdict table above is what it rests on — `Session.Cost` stays CapNone because `usage.cost.total` is per message, and it is carried as a labeled Extra instead. The **council seat** question is unchanged and still refused under §9.1's re-host class; a measured format was never an argument for a seat. The **relay-extension** path — a Pi extension writing §7.15/§7.16 files directly — remains open, is not what the adapter does, and stays gated on demand rather than on anything this survey found. ### 3.10 The canary set — what each adapter actually watches Every survey above pins an adapter to a private, unversioned on-disk format. §7 records how drift *renders*; this is the other half, and without it the next person re-verifying a vendor cannot know what was being watched. `grep -n canary docs/design.md` used to return nothing, which was the whole of the gap. A **canary** is a structural fact the survey established is present on *every* well-formed unit of that vendor's corpus. A read that examined units and found no canary is reading a corpus that has moved, and says so. `internal/adapter/drift` holds the mechanism; this is the inventory. | adapter | verified against | canary | fields it feeds | |---|---|---|---| | Claude Code | `Claude Code 2.1.233` | `sessionId` — on every JSONL record that feeds a field | name, model, workspace | | Codex CLI | `codex-cli 0.147.0` | `envelope type` — on every rollout record | model, workspace, quota, context % | | | | `session_meta record` — the FIRST record of every rollout | workspace | | Gemini CLI | `gemini-cli v0.53.1` | `metadata record` | name, subagents | | Antigravity | `agy 1.1.13` | `gen_metadata table` | model | | | | `trajectory_metadata_blob table` | workspace | | Cursor | `Cursor 3.14.7` | `composerHeaders timestamp columns` | last activity | | | | `meta.json updatedAtMs` — the CLI manifest's clock, on 43 of 43 manifests | last activity | | Grok CLI | `grok 1.0.4 (d846eb93d9)` | `summary.json info.id` — the identity envelope, on 30 of 30 sessions | name, model, workspace, last activity | | Pi | `pi 0.84.1` | `session header id` — first JSONL record is `type=session` with a non-empty `id` | name, model, workspace, last activity | The middle column quotes each canary by the **name the adapter gives it**, not a paraphrase, so the string in this table is the string in the code — which is what makes the guard tests below able to check it at all. **Cursor's row gained a second canary on 2026-08-29, and its `verified against` cell did not move.** The cell names the Cursor APPLICATION the SQLite store was surveyed inside, and `internal/adapter/pins` already records that this pin and an installed `cursor-agent` version cannot be compared. The second canary watches a different store written by a different program, so it carries its own pin — `cursor-agent 2026.08.11-e8db854`, in `chats.go`'s `chatsVerifiedAgainst` — rather than borrowing this cell. The `pins` table stays one row per vendor: `pins.For` answers per vendor id, and a second Cursor row there would make that answer depend on ordering. **Grok's row moved to 1.0.4 on 2026-08-14, and the row's two halves were re-checked to different depths.** The version came from re-measuring the seat after four patch bumps went unnoticed (§9.39's 2026-08-14 amendment). What was re-read on disk is one session directory that 1.0.4 itself wrote: `info.id` is there, so is every other key `internal/adapter/grok` names, and a rate/limit/quota sweep over it still matches nothing account-level. What was NOT re-run is the 30-session census the canary's own phrase quotes — that number is still the 2026-08-09 survey's, and it is left saying so rather than quietly re-attributed to a build it was never counted on. **One table rather than a paragraph in each of §3.1–3.9**, deliberately, and against the first instinct that a canary is a survey finding belonging beside its own survey. It is — but the question this answers is asked *across* adapters ("what is being watched, and where is the gap"), and five copies of the same claim in five subsections is five places for it to drift out of step with `internal/adapter`. The survey sections keep the evidence; this keeps the inventory. **What the columns are not.** `verified against` is CONTEXT, never a trigger — a version comparison would fire on every vendor release that did not move a byte, and a report nobody reads is worse than none. `fields it feeds` is what degrades when that canary goes missing, which is why a canary is the load-bearing subset of a schema fingerprint rather than the fingerprint: a vendor ADDING a column costs this program nothing, because every reader here addresses columns by name. **2026-08-16 — `telltale doctor` now carries the comparison this column would not make, and that changes nothing above.** The sentence before this one still holds on the read path: no adapter compares a version, nothing here triggers on one, and a canary is still structural. What changed is that the `verified against` column acquired a second reader with a different audience. §9.42's amendment has the full argument; in short, `agy` and `grok` self-update, so a pin in this table goes stale silently and CI — which installs no vendors — can never notice. The preflight already asks each seat its version, so it now says on the seat's own line when the installed build is not the pinned one, and names the § to re-measure. It is a **staleness note about this repository**, not a check of the operator's machine: no check fails, no tally moves, the exit code is unchanged, and a seat whose version could not be read gets no verdict in either direction. **This table is now machine-read.** `internal/adapter/pins` mirrors it, and every pin in that table is the adapter's own exported constant rather than a copy. `pins_doc_test.go` parses the rows above and compares them **cell by cell**, in both directions. That is the fix §3.8 named and left unowned: the six per-adapter guards match a pin anywhere in this file, so a dated paragraph quoting a new pin can turn them green over a stale cell — which is precisely how the Antigravity row survived a release reading `agy 1.1.9`. Editing a pin in this table without moving the adapter constant now fails the build, and so does adding a row no adapter claims. ## 4. Adapter contract (v1) One module per vendor implementing: - `discover()` — find live/recent sessions from vendor-native data on disk - `read(session)` — return the normalized session model (schema TBD, documented here) - `capabilities()` — which normalized fields this vendor can actually source The contract, the normalized schema, and a worked third-party example are documentation deliverables of v1, not afterthoughts. The Go form is §4a.5 and the worked example is §4a.7 — whose subject, Gemini CLI, became a real built-in adapter on 2026-08-02; the example keeps its original sketch precisely because live verification overturned part of it (see the §4a.7 postscript). ### JSONL framing rule (binding on every adapter) Both vendors' on-disk sources are JSONL (Claude transcripts, `~/.codex/sessions`), and both carry model-authored text. **A JSONL record is framed by the `\n` byte (0x0A) and nothing else.** U+2028 (LINE SEPARATOR) and U+2029 (PARAGRAPH SEPARATOR) are legal *unescaped* inside a JSON string value, so a reader that splits on "lines" in the Unicode sense tears one record in two and both halves fail to parse — a HUD row that silently loses sessions. Node's `readline` has exactly this bug; pi's `packages/coding-agent/src/modes/rpc/jsonl.ts` hand-rolls a `\n`-only splitter to avoid it, with the reasoning written down. Go is structurally safer: `bufio.ScanLines`, `bufio.Reader.ReadBytes('\n')` and `bytes`/`strings.Split` all match the 0x0A byte exactly, and the UTF-8 encodings of U+2028 (`E2 80 A8`) and U+2029 (`E2 80 A9`) contain no 0x0A byte. So the rule for adapters is: **split at the byte level, never adopt a dependency that does Unicode line-breaking.** The property is pinned by tests in `internal/claude/stdin_test.go` rather than assumed. Two adjacent traps to avoid when the adapters land: - `bufio.Scanner` caps a token at 64 KiB by default and then returns `bufio.ErrTooLong`. Transcript records routinely exceed that, and an ignored scanner error truncates the rest of the file — reading as "no more sessions". Use `bufio.Reader.ReadBytes('\n')`, or `Scanner` with an enlarged buffer, and **check `Err()`**. - A trailing partial line (the vendor is still writing) is not a record. Hold it until its `\n` arrives; a half-record must degrade to `—`, never to a parsed-looking value. **Audit record (2026-08-01):** the repo was swept for this hazard at the point the pi harness surfaced it. At that time the only parse site was `claude.Parse` (`internal/claude/stdin.go`), a streaming decode of a single JSON value from stdin with no line splitting anywhere — not exposed, nothing to fix. This section exists so the question is answered before the HUD adapters are written, not re-audited after. **Implementation (added with the adapters):** the rule now has one tested home, `internal/jsonl` — `Split`, `Scan`, `Head` and `Tail`. `Tail` additionally discards the first fragment after a backward seek, because a bounded tail read lands mid-record and parsing that fragment invents a record the vendor never wrote. Both adapters use it; no adapter reads bytes on its own. ## 4a. The normalized session model *(Answers §6 Q2. Implementation: `internal/model/session.go`, package `model`, stdlib only — the statusline path must never link a TUI framework.)* Adapters produce `model.Session`; the statusline and the HUD only read it. Nothing downstream of an adapter knows a vendor's field names, units, or file formats. ### 4a.1 Absence is two different things The honest-gauge rule (ADR-001) says a displayed value must come from vendor data. That forces a distinction most schemas skip, because the HUD renders the two cases differently: | | what it is | how it is encoded | how it renders | |---|---|---|---| | **absent now** | the adapter can source this field, but there is no value for this session right now (Claude's `rate_limits` on an API-key login) | nil pointer **+** capability declared | `—` in the cell | | **can't know** | the vendor exposes no such thing, ever | nil pointer **+** capability not declared | the column is dropped for that vendor | So presence lives in two places on purpose: **the value** (a nil pointer, per session) and **the capability** (a static declaration, per adapter). Neither alone is enough — a column of dashes for a vendor that could never fill it is itself a small lie about what was measured. Every optional field is a pointer. There is no "unset" sentinel number anywhere in the package: `0` always means the vendor said zero, and the zero `time.Time` is invalid input (`Validate` rejects it) rather than a stand-in for "no timestamp". ### 4a.2 Fields Required — an adapter that cannot produce these has no row to render: | Go field | Type | Notes | |---|---|---| | `Vendor` | `VendorID` | stable lowercase id, matches the adapter package name; appears in config keys and fixtures | | `ID` | `string` | opaque, unique within the vendor, stable for the session's life — the HUD matches rows across polls with `Vendor/ID` | | `ObservedAt` | `time.Time` | when **this snapshot was read**, not when the session last did anything. The HUD marks rows whose snapshot has aged out because polling failed; showing an old snapshot as current is precisely the failure the honest-gauge rule exists to prevent | Optional — each has a stable **field id** used by `Capabilities`, fixtures, and this doc: | Field id | Go field | Type | Meaning | |---|---|---|---| | `name` | `Name` | `*string` | human label. Model-authored text: may contain U+2028/U+2029, so renderers must not assume one line (§4) | | `model` | `Model` | `*Model` | `{ID, DisplayName}`; `Name()` falls back to the id, same rule as the statusline | | `workspace` | `WorkspaceDir` | `*string` | absolute native-format path. `WorkspaceName()` gives the basename for display | | `context_pct` | `ContextPercent` | `*Percent` | 0–100, as vendors report percentages. Convert once at the edge; nothing downstream rescales | | `cost` | `Cost` | `*USD` | USD only. A vendor reporting another currency declares `CapNone` rather than converting at an unsourced rate | | `quota` | `Quota` | `[]QuotaWindow` | labeled usage windows, see below | | `last_activity` | `LastActivity` | `*time.Time` | last observable activity. An **input** to liveness, not a claim about it | | `liveness` | `LivenessHint` | `*Liveness` | the adapter's own verdict; see 4a.4 for when you are allowed to set it | | `subagents` | `Subagents` | `*int` | count of the session's recently-written sub-agent transcripts. **Zero is a measurement** (we looked and found none) and must survive as one; nil means the count could not be taken | Three more per-snapshot annotations, all optional: - `Derived FieldSet` — fields whose value in *this* snapshot the adapter computed rather than read (summing transcript token counts into a context percentage, say). Must be a subset of the adapter's declared `Capabilities.Derived`, and every marked field must actually carry a value. The HUD renders these with an estimate marker; ADR-001 requires inferred values be visibly marked, not silently mixed in with reported ones. - `Degraded FieldSet` — fields the adapter tried to read and failed (a truncated JSONL record, an unparseable number). Degraded fields must be absent. **Degraded and plain-absent render identically as `—`**; the difference is diagnostic only, shown in the detail pane. If "we failed to read it" got its own gauge glyph it would start to read as data. - `Diagnostics []string` — operator-facing notes explaining degradation. Never rendered as values, and (public repo) they describe structure — `"record 41 truncated"` — never transcript content. `Extras []Extra` is the escape hatch for vendor-specific labeled strings, so an adapter with something extra to show does not stuff it into a field that means something else. Extras are display-only: no thresholds, no colors, no sorting, detail pane only. If an extra deserves a gauge, it deserves a `Field` — propose one. *(v1.1: the detail pane (§7.11) is that surface, and it is now the only place extras appear — `TestDetailPaneIsTheOnlyPlaceExtrasAppear` asserts both halves. Both adapters populate them: git branch, CLI version, Claude's context token count, Codex's plan and history mode.)* **Why `subagents` is a `Field` and not an Extra.** It fails the Extra test in both directions: it is a *number* with an absent-versus-zero distinction the Extra type cannot carry (an Extra is a string, and `""` would collapse "none running" into "could not count"), and it renders as a gauge-adjacent mark in the grid rather than as a labelled line in the pane. That is exactly the "if an extra deserves a gauge, it deserves a Field" case, taken rather than dodged. ### 4a.3 Quota windows Windows are a slice, not named fields, because the set is vendor-defined. Emit only the windows your vendor actually has, in display order, shortest first. Presence works at two levels and both are load-bearing: - a window the vendor does not have is **absent from the slice**; - a window that exists but has no usage figure yet is **present with a nil `UsedPercent`** and renders `—`. Never `0%`. Each window carries `ID` (stable snake_case key, e.g. `five_hour`), `Label` (short display string, ≤ 4 cells — the statusline is character-budgeted), `UsedPercent`, and `ResetsAt`. A nil `ResetsAt` hides the countdown rather than guessing one. A window whose *length* the vendor did not report gets a positional label (`1st`, `2nd`) from the adapter: calling it "5h" on a guess would be a duration claim with no source. ### 4a.4 Liveness: who decides **The HUD decides.** It classifies from `LastActivity` against one shared `LivenessThresholds`, so every vendor's rows are judged by the same rule and are comparable side by side. Defaults (`DefaultLivenessThresholds`): | age since last activity | class | |---|---| | ≤ 2 min | `live` | | ≤ 15 min | `idle` | | > 15 min | `stale` | | no timestamp, no hint | `unknown` → renders absent | 2 minutes because a working agent can go a long single turn without writing anything to disk, and a boundary shorter than the longest quiet stretch of real work flaps mid-task. 15 minutes because past that a session is nearly always one the user walked away from, so it sorts to the bottom instead of competing for attention. Both are defaults, not constants: the HUD may expose them and the eval harness pins renders against explicit values. A `LastActivity` in the future (file mtime vs. local clock skew) clamps to age zero rather than going negative — and the adapters degrade it to absent before it gets that far, so the clamp is a floor rather than a render path. `unknown` is a real state and renders as absent — **never as `stale`**. "Stale" is a claim; "we have no activity signal" is not. **The adapter's only input is `LivenessHint`, and it wins when set.** Set it *only* from a positive vendor signal the HUD cannot see: - ✅ a turn-started / turn-ended event from a hook or notify stream; - ✅ the vendor has recorded the session as ended → `LivenessStale`, even though `LastActivity` is seconds old (this is the strongest legitimate hint); - ❌ anything computed from the age of `LastActivity` — that is the HUD's job, and duplicating it per-adapter makes vendors incomparable at different boundaries; - ❌ "a process with that name is running" — that is evidence a process exists, not that the session is doing anything. A hint the model cannot check is the one place an adapter can lie undetected, which is why the bar for emitting one is a signal that actually separates working-now from process-exists. **Neither v1 adapter emits one** (§3.1, §3.2). ### 4a.5 The adapter interface ```go type Adapter interface { Vendor() VendorID Capabilities() Capabilities Discover(ctx context.Context) ([]SessionRef, error) Read(ctx context.Context, ref SessionRef) (*Session, error) } ``` - **`Capabilities()`** is static: which normalized fields this vendor can source, and how. It must not vary with what a particular session happens to contain — that is what nil pointers are for. Callers may cache it. ```go type Capabilities struct { Reported FieldSet // read from vendor output verbatim (modulo unit conversion) Derived FieldSet // computed by the adapter from something that isn't the value } ``` The two sets are disjoint; a field in neither is `CapNone` — "can't know". Declaring a capability is a promise about the *source*, not about any given session. - **`Discover()`** must stay cheap: directory listing and `stat`, no parsing. The HUD calls it every poll tick. It returns `SessionRef{Vendor, ID, Locator, LastActivity}` — `Locator` is vendor-private (a path, a pipe name), opaque to the HUD, never rendered (on a shared machine it can name another user's paths), and handed back to `Read` unchanged. `SessionRef.LastActivity` is a scheduling hint (typically an mtime) so the poll loop can skip unchanged sessions; the value the HUD *displays* comes from `Read`. - **`Read()`** parses one session. **Partial failure is not an error**: a field you cannot parse is left nil, added to `Degraded`, and explained in `Diagnostics` — the row still renders with `—` in that cell. Return an error only when there is no session to report at all. - Errors the HUD handles by showing *less*, not by showing a banner: `ErrVendorAbsent` (vendor not installed — the vendor disappears from the HUD entirely; a user without Codex should not stare at a Codex error forever) and `ErrSessionGone` (the session vanished between `Discover` and `Read` — the row drops silently). Any other `Read` error drops that one row and nothing else; the Codex adapter uses this for `ErrSubAgentThread`, a rollout that is a sub-agent's thread rather than a session and cannot be identified before parsing. - Implementations must be safe for concurrent use (the HUD polls vendors in parallel), must not write to vendor state, and must not make network calls or read credentials. - **JSONL adapters:** §4's framing rule is binding. Split on the `0x0A` byte only, use `bufio.Reader.ReadBytes('\n')` (or `Scanner` with an enlarged buffer) and **check `Err()`**, and treat a trailing partial line as not-yet-a-record. `internal/jsonl` is the shared implementation; use it rather than re-deriving it. ### 4a.6 The validation gate `(*Session).Validate(caps)` is the machine-checkable form of the honest-gauge rule. It rejects a session that: - carries a value for a field the adapter declared unsupported; - marks a field derived that was not declared derived, or marks a field derived that carries no value; - reports a field as both present and degraded; - has a percentage outside 0–100 (drop it and mark it degraded — a clamped value is invented data), a negative cost, or a zero `time.Time` used to mean "absent"; - has no `Vendor`, `ID`, or `ObservedAt`, or duplicate/unlabeled quota windows. The eval harness runs it over every fixture; run it in your adapter's own tests too. It is not on the render path. ### 4a.7 Worked example: adding a Gemini CLI adapter *(2026-08-02: this example became real — `internal/adapter/gemini`, seam verification in §3.7. The sketch below is kept AS WRITTEN, wrong guess included, because the gap between it and the shipped adapter is the section's whole lesson; see the postscript.)* **Step 0 — verify the seam before writing a line of code.** ADR-001's live-doc rule applies to third-party adapters too: check the vendor's *current* docs for what it actually writes to disk. This repo's own Codex plan was falsified on first check (no statusline hook), and the sketch below is deliberately a *shape*, not a claim — the file locations and field names are placeholders until you have verified them. A capability you cannot point at a documented source for is `CapNone`. **Step 1 — write the capability table first**, in your ADR or PR description. It is the honest inventory of what your vendor can actually answer, and it is the thing reviewers argue with. Illustrative shape: | field id | capability | source | |---|---|---| | `name` | reported | session file header | | `model` | reported | session file header | | `workspace` | reported | session file header | | `context_pct` | **derived** | token counts summed from records — not a vendor-reported percentage | | `cost` | none | vendor exposes no cost | | `quota` | none | vendor exposes no quota window | | `last_activity` | reported | timestamp of the last record | | `liveness` | none | no turn-start/turn-end signal to read | **Step 2 — implement.** ```go // Package gemini adapts Gemini CLI's on-disk session data to model.Session. // Source paths and field names verified against on . package gemini const Vendor = model.VendorID("gemini") type Adapter struct{ root string } // e.g. filepath.Join(home, ".gemini", "sessions") func (a *Adapter) Vendor() model.VendorID { return Vendor } // Capabilities: context_pct is DERIVED — we sum token counts ourselves, so the // HUD marks it as an estimate. Cost, quota and liveness are absent from the // vendor's data entirely and are declared nowhere: the HUD drops those columns // for gemini rather than printing dashes it can never fill. func (a *Adapter) Capabilities() model.Capabilities { return model.Capabilities{ Reported: model.NewFieldSet( model.FieldName, model.FieldModel, model.FieldWorkspace, model.FieldLastActivity, ), Derived: model.NewFieldSet(model.FieldContextPercent), } } // Discover stats the session directory only — no parsing on the poll path. func (a *Adapter) Discover(ctx context.Context) ([]model.SessionRef, error) { entries, err := os.ReadDir(a.root) if errors.Is(err, fs.ErrNotExist) { return nil, model.ErrVendorAbsent // not installed: hide the vendor } if err != nil { return nil, err } var refs []model.SessionRef for _, e := range entries { info, err := e.Info() if err != nil { continue // racing the vendor's writer is normal, not fatal } refs = append(refs, model.SessionRef{ Vendor: Vendor, ID: strings.TrimSuffix(e.Name(), ".jsonl"), Locator: filepath.Join(a.root, e.Name()), LastActivity: model.TimePtr(info.ModTime()), }) } return refs, nil } func (a *Adapter) Read(ctx context.Context, ref model.SessionRef) (*model.Session, error) { f, err := os.Open(ref.Locator) if errors.Is(err, fs.ErrNotExist) { return nil, model.ErrSessionGone } if err != nil { return nil, err } defer f.Close() s := &model.Session{Vendor: Vendor, ID: ref.ID, ObservedAt: time.Now()} // §4 framing rule, via the shared implementation: records are framed by // 0x0A and nothing else, there is no 64 KiB cap, and a trailing partial // line is not a record. err = jsonl.Scan(f, func(line []byte) error { var rec record if json.Unmarshal(line, &rec) != nil { // One bad record degrades the fields it fed, not the whole row. s.Degraded = s.Degraded.With(model.FieldContextPercent) s.Diagnostics = append(s.Diagnostics, "unparseable record skipped") return nil } apply(s, rec) return nil }) if err != nil { return nil, err } // Derived, and declared as such: this is a computed estimate, and the HUD // renders it with an estimate marker rather than as a vendor-reported figure. if pct, ok := estimateContext(s); ok && !s.Degraded.Has(model.FieldContextPercent) { s.ContextPercent = model.PercentPtr(pct) s.Derived = s.Derived.With(model.FieldContextPercent) } // Cost and quota are never set: not declared, so not knowable. // LivenessHint is never set: no positive signal, so the HUD classifies by age. return s, nil } ``` **Step 3 — fixtures and the gate.** Add synthesized fixtures per state — healthy, empty, degraded (a truncated final record), and whatever your vendor's equivalent of "logged in a way that hides quota" is — and assert `Validate` passes plus the exact render for each. Fixtures are **synthesized**: fake session ids, fake text, fake paths, realistic in shape only. This repo is public and real transcripts carry private material; no real session content enters `testdata/`, ever. **Step 4 — write down what you could not source**, in your adapter's package doc. A field declared `CapNone` with a one-line reason is a finished answer. A field quietly filled with a plausible number is the bug this whole schema exists to make hard. **Postscript — what Step 0 did to this very sketch.** The sketch above, written before anyone read the vendor's source, guessed `context_pct: derived` ("token counts summed from records"). The source read (§3.7) falsified it: Gemini's per-message token counts are real, but no context-window size reaches disk — the CLI's own percentage divides by a static table compiled into its binary, which is exactly the assumed denominator decisions/001 forbids. The shipped adapter declares `context_pct: CapNone` and carries the token reading as a display-only extra instead. The hypothetical also missed the two things only the source could reveal: message records are upserts (same id re-appended), and the writer deletes non-resumable sessions on exit. If you skip Step 0, those two become bugs; the wrong capability guess becomes a fabricated gauge. ## 5. Eval harness Fixture-driven, in-repo, CI-gating (`.github/workflows/ci.yml` runs `go vet ./...` and `go test ./...` on `windows-latest`, then smoke-tests the built binary against a statusline fixture). What it asserts today: | Layer | Package | What is pinned | |---|---|---| | Statusline renders | `internal/statusline` | every segment against five stdin fixtures, including the API-key login that must render no quota | | Statusline (agy) | `internal/statusline` + `internal/antigravity` | four agy fixtures + inline cases: full render exact, confirm? outranking the state word, ctx 0% as a reading beside hidden quota, bucket-without-reading hides, unknown state verbatim, reset_time fallback, U+2028 in a string value, the product routing marker | | Framing rule | `internal/jsonl` | 0x0A-only framing, a 300 KiB record surviving the `bufio.Scanner` cap, read errors surfacing, torn tails held back, a seek fragment discarded | | Schema gate | `internal/model` | `Validate` over every rejection case, liveness boundaries, presence semantics | | Claude adapter | `internal/adapter/claudecode` | discovery filters, the `` trap, the `input_tokens` trap, torn tail invisibility, torn-only session, future mtime, capability table | | Codex adapter | `internal/adapter/codex` | envelope + internally-tagged event parsing, derived context, quota window presence, null `rate_limits`, sub-agent rejection, capability table | | SQLite reader | `internal/sqlite` | record decoding across every storage class, the zero-width serial types, the overflow-page chain (25 KiB blob against a 4 KiB page), missing-table-is-absence, and the WAL overlay: a committed sidecar value winning, a corrupt frame ignored with a note, a bad header rejected whole, mismatched salts ignored, a torn tail never assembled | | SQLite reader (v1.2) | `internal/sqlite` | `Rows` streaming the same rows as `Table` and stopping when the callback says so; `Columns` splitting a CREATE statement across every quoting style, a parenthesized type, a comma inside a default and a table-level constraint, and yielding nothing rather than a guess on a statement it cannot read | | Antigravity adapter | `internal/adapter/antigravity` | discovery that ignores the stale summary index and the sidecars, WAL overlay changing the reported model, overflow-spanning generation blob, dedup on the response id rather than the constant conversation UUID, invariant violation dropping the tokens with a diagnostic, zero tokens as data, absent workspace as absence, missing transcript as a typed sentinel, unreadable database degrading rather than dropping the row, Q8 fold, future mtime, transcript content never reaching a field, capability table | | Cursor adapter | `internal/adapter/cursor` | discovery dropping the five row shapes that are not sessions (draft sentinel, `value.isDraft`, archived, sub-agent, empty-window), a store whose main file is one empty page reading entirely out of its sidecar, the vendor's own percentage reported unmarked, a computed one marked, neither-present as absence, `default` rendered literally, unpopulated zeros never becoming a cost or a token reading, missing workspace mapping as absence and an unparseable one as degradation, mixed epoch-ms/ISO-8601 timestamps, future-skew, all-timestamps-unreadable degrading, the stale `ItemTable` mirror losing to the table, an unrecognized schema erroring rather than reporting zero, capability table, and the allowlist: three planted markers (prompt text, credentials, plan entitlements) reaching no field, extra or diagnostic | | Gemini adapter | `internal/adapter/gemini` | fixed-depth discovery (legacy `.json` and nested sub-agent files excluded), upsert last-wins, `$set` summary/lastUpdated, registry workspace lookup (verbatim, corrupt-degrades, absent-is-absent), sub-agent nest counting, `kind:"subagent"` rejection, torn tail invisibility, future mtime, Q8 fold, capability table | | Claude adapter (v1.1) | `internal/adapter/claudecode` | the sub-agent count: recency boundary, future-mtime exclusion, non-transcript neighbours ignored, absent sidecar as a measured zero | | HUD renders | `internal/hud` | every golden frame byte-for-byte at 52/72/80/120 columns (count enforced by TestEveryGoldenIsClassified, not restated here — a literal drifted once already), the §7.4 gauge table, the estimate marker, threshold colours, frame width/height invariants | | HUD behaviour | `internal/hud` | vendor status words from adapter errors, key handling, one-scan-in-flight, spinner lifecycle | | HUD behaviour (v1.1) | `internal/hud` | esc unwinding one layer at a time, find mode swallowing the keyboard, selection carried by session key across a re-sort, the pane closing when its session ends | | Burn arithmetic | `internal/hud` | the minimum basis, the four refusals, least-squares slope against injected series, rollover detection vs. `resets_at` jitter, sample throttling and eviction | | Fixture legality | `internal/hud` | every session behind every golden passes `model.Validate` against its vendor's declared capabilities — a golden may not pin a render of a state the schema forbids | | Doc/code sync | `internal/hud` | every render pasted into `docs/design.md` §7.3/§7.11–§7.14 still matches its golden, and every golden is either embedded or explicitly exempted | | Picture/code sync | `internal/hud` | `README.md`'s hero picture is re-emitted from the `readme` golden and byte-equal to the committed file, with the characters read back out of the emitted markup and diffed against the render — plus no dollar sign anywhere, and the estimate marker surviving into the picture | | Picture/code sync | `internal/council` | the hero picture `README.md` and `docs/council.md` both show is re-emitted from the `activity` golden with its all-blank rows dropped, byte-equal and read back the same way, with exactly one seat wearing the focus mark | | Fast path (ADR-002) | `internal/statusline` | the statusline's transitive import graph reaches no `charm.land/bubbletea` or `charm.land/lipgloss` package — the framework cannot be initialized on a path it does not import | Rules that outrank convenience: - Fixtures are **synthesized**. No real session content enters `testdata/`, ever. - Golden renders use a pinned clock, an explicit terminal size and a plain style set, so they never depend on the CI terminal. - A failing render assertion fails the build. - No number appears in README/launch material unless this harness generated it. **Amendment, 2026-08-16: the fast path is gated structurally and reported numerically.** ADR-002's fast-path rule — the statusline never initializes Bubble Tea — was the oldest load-bearing claim in this document with nothing checking it. Any `import` added to `internal/theme` or `internal/model` for one convenient helper would have compiled, passed every golden and every smoke, and put the TUI framework's package init on a path that runs on every prompt. Two things now exist, and they are deliberately unequal. **The hard gate is structural.** `TestFastPathNeverReachesTUIFramework` (`internal/statusline/fastpath_test.go`) asks the toolchain — `go list -deps` — for the statusline's transitive imports and fails if any is under `charm.land/bubbletea` or `charm.land/lipgloss`. It runs inside `go test ./...`, so it gates on both CI jobs. It was verified against an injected violation before it was trusted: a bare `import _ "charm.land/lipgloss/v2"` in `internal/theme` turns it red, naming the package. A gate nobody has seen fail is a gate nobody has tested. **The timing is reported, and holds only a gross-regression ceiling.** A `ci.yml` step spawns the built binary 15 times against `full.json`, asserts the line rendered on every sample, prints min/median/max, and fails only if the median exceeds **400 ms**. That ceiling is not the budget and does not pretend to be. A shared runner's scheduling noise is larger than the quantity ADR-002 cares about, so pinning single-digit — or even double-digit — milliseconds there would buy a flaky build, not a guarantee. **What the gate deliberately does NOT assert**, stated so a later reader does not credit it with more than it does: - **Not the single-digit-millisecond budget.** Nothing in CI measures that. The measurement lives on the reference workstation (i7-7700K, Windows 11, 2026-08-16): parse+render 14 µs, end-to-end median 26.6 and 29.0 ms over two runs of the CI harness. The end-to-end figure is process start, not telltale's work. - **Not "the binary does not link the framework".** It does — see §7.5's 2026-08-16 correction. One binary carries the HUD, so the module is linked and the structural claim can only ever be about the statusline's own import graph. - **Not a small constant regression.** The 400 ms ceiling catches a change in the SHAPE of the path — a whole-corpus scan (§7.18 measured cold scans at 896–1200 ms), a network round trip, a TUI init and its terminal-capability queries. A path that got three times slower and stayed under the ceiling passes, and the printed median is the only thing that would show it. That is a human's job, on purpose. - **Not the other two gauges.** `internal/hud` and `internal/council` are TUI surfaces; the rule does not apply to them and the gate does not look at them. **Amendment, 2026-08-18: what the long-running modes cost after the first minute, measured.** Everything above measures one call, and the 2026-08-16 timing step measures 15 of them. Three modes do not stop after one call. `telltale events` and `telltale otel grok` are listeners that an operator starts once and leaves. `telltale council` holds a room and one long-lived child for each seat. A benchmark cannot answer the question these modes raise: does an idle process hold a constant amount of memory across a day? Nothing in this repository measured that. This amendment records the instrument and the first measurement. **The instrument is `tools/soak.ps1`, and it is not a gate.** It has three modes. Residency mode samples one process tree at an interval, and it records the working set, the private bytes, the handle count, the live CPU and the child count. Per-fire mode starts a short-lived process many times and records each run, because `telltale statusline` has no residency to sample. Summary mode renders a finished JSONL file again, so a later reader re-reads an arm without a second soak. The samples are the artifact, and a table is one view of them. Re-read a finished arm with `.\tools\soak.ps1 -Summarize "$env:TEMP\soak-events.jsonl"`. Three decisions in that script are load-bearing, and each one comes from a measured failure: - **`Win32_Process`, and never `Get-Counter`.** The counter set names of `Get-Counter` are localized, so `\Process(*)\Working Set` does not resolve on a Windows installation in another language. Its instances also carry process names, so three `telltale.exe` processes arrive as `telltale`, `telltale#1` and `telltale#2`, with no stable relation to a pid and no parent data. A tree soak needs the parent data. One CIM query also serves a whole sample at any tree size. - **The summary prints a median-based drift beside the least-squares slope.** One outlier moves a slope, and no outlier moves a median. Arm C below proves why that matters: its slope and its drift disagree, and the drift is the honest figure. - **An absent figure prints `--`, and never `0`.** `PeakWorkingSet64` reads 0 after a process exits, and PowerShell evaluates `$null / 1MB` as 0. Both defects produced a measured-looking zero in the first smoke arms. That is the zero-vs-absent collapse ADR-001 forbids, inside the instrument that exists to find drift. **Conditions.** The reference workstation: Intel i7-7700K (4 cores, 8 logical threads), Windows 11, Windows PowerShell 5.1.26100.9168, Go 1.26.5. The binary came from `d024f03`. Both listeners bound non-default loopback ports (14519 and 14318), and every arm ran with `USERPROFILE` redirected to a scratch directory. The redirect is verified in both directions: the listeners wrote their stores under the scratch home, and `~/.telltale` held no file newer than the previous day when the arms ended. No arm read or wrote a real store. Arms A and B ran concurrently, and arm C ran on the same machine while they sampled. **Arm A — `telltale events`, idle, 34.8 min, 210 samples at 10 s.** | metric | min | median | p95 | max | slope/hour | |---|---|---|---|---|---| | working set | 9.14 MiB | 9.23 MiB | 9.23 MiB | 9.25 MiB | +0.11 MiB | | private bytes | 45.32 MiB | 45.32 MiB | 45.35 MiB | 45.36 MiB | -0.02 MiB | | handles | 128 | 128 | 128 | 128 | 0.0 | | processes | 1 | 1 | 1 | 1 | 0.0 | Robust drift: working set 0.00 MiB, private bytes 0.00 MiB, handles 0.0, processes 0.0. CPU over 209 comparable intervals, 0 dropped: min, median and max all 0.000%, total 0.02 s. **Arm B — `telltale otel grok`, idle, 34.8 min, 210 samples at 10 s.** | metric | min | median | p95 | max | slope/hour | |---|---|---|---|---|---| | working set | 9.02 MiB | 9.11 MiB | 9.11 MiB | 9.13 MiB | +0.11 MiB | | private bytes | 45.15 MiB | 45.15 MiB | 45.18 MiB | 45.20 MiB | -0.02 MiB | | handles | 128 | 128 | 128 | 128 | 0.0 | | processes | 1 | 1 | 1 | 1 | 0.0 | Robust drift: working set 0.00 MiB, private bytes 0.00 MiB, handles 0.0, processes 0.0. CPU over 209 comparable intervals, 0 dropped: min, median and max all 0.000%, total 0.00 s. **Arm C — `telltale statusline`, 1000 fires.** Every fire exited 0, and every fire rendered the `Opus` marker that `ci.yml` asserts on its own 15 samples. | metric | min | median | p95 | max | slope/1k fires | |---|---|---|---|---|---| | wall ms | 20.0 | 20.9 | 26.0 | 2035.7 | -41.2 | | cpu ms | 0.0 | 15.6 | 31.3 | 93.8 | 0.0 | | peak working set | 9.32 MiB | 9.43 MiB | 9.61 MiB | 9.63 MiB | -0.01 MiB | | peak private | 44.90 MiB | 45.24 MiB | 45.44 MiB | 46.25 MiB | -0.01 MiB | Robust drift: wall -0.2 ms, cpu 0.0 ms, peak working set 0.00 MiB, peak private 0.00 MiB. p99 is 55.23 ms. 11 fires passed 50 ms, and 6 of those passed 100 ms. The cpu figure is quantized to the 15.625 ms scheduler tick, so only its distribution carries information. **What the three arms found.** 1. **Neither listener leaks.** The handle count and the process count held exactly constant across 210 samples of each arm. Both working-set slopes read +0.11 MiB/hour, and that figure is smaller than the 0.11 MiB total spread of the same arm. The robust drift is 0.00 MiB on every metric. The slope is therefore the noise of the arm, and it is not a trend. Each arm ran 34.8 min, so every hourly slope above is a 1.7x extrapolation. 2. **Idle CPU is under the measurement floor.** Every comparable interval read 0.000%, and no interval was dropped. The two listeners spent 0.02 s and 0.00 s of CPU across 35 minutes. An idle listener costs nothing this instrument can measure. 3. **The statusline cost does not drift, and its tail is the finding.** 6 fires of 1000 blocked for 2019-2036 ms, while their CPU stayed at the usual 15-31 ms. The process waited; it did not compute. The -41.2 ms per 1000 fires slope comes from those 6 samples alone, and the median-based drift reads -0.2 ms across the same run. This is the disagreement the drift figure exists to expose, and the drift is the honest one. **This arm does not name a cause.** Two residency arms polled `Win32_Process` on the same machine throughout, so the condition is recorded and the cause is not claimed. A statusline fires on every prompt, so a 2 s stall earns a later arm on an idle machine. **What this measurement does NOT assert**, stated so a later reader does not credit it with more than it did: - **Not a CI gate.** `tools/soak.ps1` is an operator instrument. No workflow runs it, and a leak introduced tomorrow turns nothing red. - **Not a multi-hour figure.** 35 minutes is the arm. A leak under this arm's noise floor stays invisible, and every hourly slope is an extrapolation. - **Not a loaded listener.** Both listeners sat idle, and no hook posted to either port. A listener under traffic is a different measurement. - **Not the council room.** That arm was owed when this list was written; the dated payment below records its run (2026-08-18). - **Not macOS.** The instrument uses PowerShell and `Win32_Process`, so it is Windows-only. **Owed: the council arm, which an operator must run.** The room cannot be soaked headlessly, for two reasons that belong to the mode rather than to the script. The room is a TUI, so it needs a real terminal. Its seats are live vendor CLIs that read their credentials from the real home directory, so this arm cannot use the redirected `USERPROFILE` that isolated the other three arms. The room is also the arm that matters most, because it is the only mode whose process count can move: each persistent seat is a long-lived child, and residency mode counts the whole tree. Terminal 1 opens the room and leaves it idle. `--read` seats every vendor and forbids every write, which is what an unattended 35-minute arm needs: ```powershell cd C:\Users\sanle\code\telltale .\telltale.exe council --read ``` Terminal 2 resolves the room's pid, then samples the tree. `-Name telltale` is wrong here, because it fails whenever more than one `telltale.exe` runs: ```powershell cd C:\Users\sanle\code\telltale $room = @(Get-CimInstance Win32_Process -Filter "Name='telltale.exe'" | Where-Object { $_.CommandLine -match '\bcouncil\b' }) $room.Count # must print 1 before the next command runs powershell.exe -NoProfile -ExecutionPolicy Bypass -File .\tools\soak.ps1 ` -Id $room[0].ProcessId -Label council-idle ` -Out "$env:TEMP\soak-council.jsonl" -IntervalSeconds 10 -DurationMinutes 35 ``` The arm prints its own table when it ends. Read the same file again at any later time with: ```powershell powershell.exe -NoProfile -ExecutionPolicy Bypass -File .\tools\soak.ps1 -Summarize "$env:TEMP\soak-council.jsonl" ``` **Paid, 2026-08-18. The owner ran the arm exactly as written above.** A `--read` room with all five seats reattached sat idle for 34.8 minutes, 210 samples at 10 s, on the same box as the three headless arms (Intel i7-7700K, Windows 11). One honest difference from those arms is stated up front: the room ran the operator's own clone build (the binary a CI-gate run leaves in the repo root), so the revision is not pinned, and the operator used the machine for other work during the window — the sampler is tree-scoped, so only the room's own process tree was counted, but the CPU percentages share a loaded box. | metric | min | median | p95 | max | |---|---|---|---|---| | working set | 21.55 MiB | 587.95 MiB | 727.81 MiB | 1790.64 MiB | | private bytes | 56.58 MiB | 648.42 MiB | 795.28 MiB | 2217.25 MiB | | handles | 155 | 949 | 1500 | 6978 | | processes | 1 | 5 | 9 | 34 | Robust drift (median of the second half minus the first): working set +2.44 MiB, private bytes −1.38 MiB, handles −3, processes 0. CPU over 205 comparable intervals: median 0.429%, max 1.779%, 91.16 s total. Three readings, in the order they matter: 1. **The room is not the residency; the seats are.** The tree's first sample is 21.55 MiB and one process — the room alone, before reattach. Steady state is ~590 MiB, five processes and ~950 handles: one TUI plus four persistent vendor children. An all-day room costs what its vendors cost, and the gauge binary itself is the smallest process in its own tree. 2. **No leak at steady state.** The robust drift is +2.44 MiB of working set and exactly zero processes across the halves. The large negative least-squares slopes (−173 MiB/hour) are the reattach transient, not a decline: the peaks — 34 processes, 1.79 GiB, 6978 handles — all belong to the opening minutes when five conversations reattached at once, and the tree settled from there. The same outlier-versus-median disagreement the statusline arm showed, with the sign flipped. 3. **"Idle" seats are not quiescent.** 91 CPU-seconds over an idle 35 minutes, and bursts to nine processes mid-window: the vendor CLIs run their own background machinery inside an idle room. That is vendor behavior observed from outside, not attributed to any one vendor — a per-seat attribution would need one arm per seat, which nobody has run. What this arm still is not: one run, one build, one box, 35 minutes — and a loaded box for the CPU column. The reattach transient it measured is also the demo's opening beat, which is worth knowing before a stage. ## 6. Open design questions 1. ~~Language/stack~~ — **ANSWERED, ADR-002:** Go + Bubble Tea/Lipgloss, one binary, two modes. Windows-first hardened; macOS/Linux deferred post-v1. 2. ~~Normalized session schema~~ — **ANSWERED, §4a:** `model.Session` + `Capabilities`, pointers for absent-now vs. undeclared capability for can't-know, liveness classified by the HUD from `LastActivity` with an adapter override only on positive vendor evidence, and `Validate` as the machine-checked honest-gauge gate. 3. ~~HUD refresh model~~ — **ANSWERED, and the measurement it was waiting for has been taken.** A 1 s `tea.Tick` poll, not a file watcher. The survey's inputs argued for it: 837 sessions across 33 directories, transcripts up to 7.7 MB, and a projects tree that provably mutates mid-sweep. `Discover` is stat-only and `Read` is head+tail bounded, which is what made polling affordable; a watcher over a mutating tree on Windows is a larger correctness surface for a smaller win. The open part was the cold cache, and `BenchmarkScan` closed it over a synthesized 1,400-session corpus: the warm scan went 798 ms → **82 ms** with the `(size, mtime)` cache, and on the live corpus 1.84–3.37 s → **181–204 ms**. **The cold scan did not move** (896 ms → 994 ms, unchanged within noise), and that is the answer rather than a remaining gap: the first frame must genuinely read everything, which is what the spinner exists for. The poll was never the cost; re-reading unchanged files was. 4. ~~Exact Claude/Codex on-disk data sources~~ — **ANSWERED, §3.1–3.3**, with Claude verified live and Codex's first live pass run 2026-08-01 (§3.4; short remainder itemized there). 5. ~~Distribution naming (`telltale-hud` on any registry; winget/scoop manifests) — at packaging time. Go binary means npm is optional, not required.~~ — **ANSWERED 2026-08-08, §8 "Packaging decisions":** the bare name `telltale` was free on scoop and winget, so the `telltale-hud` fallback goes unused; winget takes the publisher-qualified `sanlee-ys.telltale`; npm stays skipped rather than renamed. 6. ~~HUD UI design section~~ — **ANSWERED, §7:** layout grid, colour/threshold tokens shared with the statusline, motion rules, degraded-state renders. Written before HUD build per ADR-002; every render in §7.3 and every row in §7.7 is a golden/fixture. 7. **Cross-vendor context percentage — ANSWERED PROVISIONALLY, wants a ruling.** Claude and Codex context percentages are different statistics (§3.3), and Claude's is not derivable from disk at all. Three options were on the table: 1. show raw token counts only, cross-vendor, no percentage anywhere; 2. show a percentage only where a vendor ships a denominator, accepting a ragged column; 3. show each vendor's own formula, labelled as vendor-native and non-comparable. **v1 ships (2).** Codex's `context_pct` is `CapDerived` and renders with an estimate marker; Claude's is `CapNone` and renders absent; and when no visible row can fill the column it is dropped entirely. (3) is rejected outright: two numbers in one column that are not the same statistic is exactly the lie the honest gauge forbids. (1) is the more conservative answer and remains the better one if the ragged column reads badly in daily use — it needs a `context_tokens` field in the schema, which is an additive change, and both adapters already carry the token count as an extra so nothing has to be re-derived. **Decide after two weeks of dogfood, not before.** 8. ~~LastActivity source on Windows~~ — **RULED 2026-08-01: option (iii), `max(mtime, newest record timestamp)`, implemented in both adapters.** NTFS defers mtime while the writer holds the file (~100 s observed on a hot rollout; ~20 min on a closing one, seen twice), so mtime alone under-reports on exactly the rows the HUD exists to watch. The newest record `timestamp` (RFC3339, vendor-written on records in both formats — verified live; some Claude housekeeping records omit it) comes out of the tail window the adapters already read, so the fold is zero extra I/O. Rules: each signal independently passes the future-skew guard or is excluded (a wrong vendor clock cannot fresh-wash a row); the fresher valid signal wins; the field degrades only when BOTH are unreadable. Still `CapReported` — the max of two vendor-written stamps invents nothing. Pinned by `TestLastActivityUsesNewestRecordTimestampOverStaleMtime` in each adapter; no HUD golden changed (fixtures inject `LastActivity` directly). ## 7. HUD UI design Written before the HUD was built, per ADR-002. This section is binding: every render below is a golden-test target in `internal/hud/testdata/golden/`, and the degraded states in §7.7 are eval fixtures like the statusline's. Library facts were verified against live docs on 2026-08-01 and then against the compiled API: `charm.land/bubbletea/v2` **v2.0.8** and `charm.land/lipgloss/v2` **v2.0.5**. Both are v2: `AdaptiveColor` and the global `Renderer` are gone, `Style` is a plain value type, and `Model.View()` returns a `tea.View` struct rather than a string. > **The renders in §7.3 are generated by the build, not drawn by hand.** They are pasted > from `internal/hud/testdata/golden/*.txt`, which `go test ./internal/hud -update` > regenerates. If a render here and the code disagree, the code is right and this section > is stale — fix it in the same change. ### 7.1 Principles Five rules, in priority order. Where polish and a rule conflict, the rule wins. 1. **The honest gauge extends to pixels.** Absent data renders as absent. Specifically: an *empty gauge track* means zero, so an absent gauge draws **no track at all** — the field is blank and the number beside it is `—`. `0%` and "no data" must never produce the same row of glyphs. This is the load-bearing render assertion of the whole HUD. 2. **Colour is always redundant.** Every distinction is carried by a glyph or a number first; colour only reinforces it. This makes `NO_COLOR` degradation correct by construction rather than by a second code path, and makes the HUD readable to colour-blind users without a mode. 3. **telltale may animate its own work; it must never animate the vendor's.** See §7.6. 4. **Still by default.** This is a glanceable monitor. In steady state the only cell permitted to change each second is the `AGE` of a session younger than one minute. If more than that moves, it is a bug, not a flourish. 5. **Same product as the statusline.** Identical thresholds, identical palette, identical `│` separator, identical `↻` countdown. The two surfaces share numbers through `internal/theme` (§7.5), not by coincidence. A sixth rule that is a consequence of #1 and worth stating on its own: **account-level quota appears once per vendor, in the header, never per row.** `rate_limits` is a property of the account, not the session; repeating it on every row would assert per-session quota, which is false. If no source can honestly supply it, that vendor's block is absent — not zeroed — so the block count is itself a measurement: the header shows exactly as many vendors as telltale can speak for. Where the readings come from and how the line fits them is §7.15. ### 7.2 Anatomy ``` header identity, session counts, account quota 1 line (2 below 100 cols) rule ───────────────────────────────────────── 1 line col header SESSION MODEL CONTEXT COST AGE 1 line rows one per session, sorted, scrollable n lines rule ───────────────────────────────────────── 1 line footer key hints (left) · state notices (right) 1 line ``` Chrome is 5 lines at wide, 6 below 100 cols (the quota block wraps to its own line). The column-header row is dropped when the body is not the grid — over the help overlay or the empty state it would label columns that are not on screen. **Column grid.** Widths are fixed; the `SESSION` column is the only flexible one and absorbs all slack, which right-anchors the numeric block at every terminal width. Offsets below are 1-based for the wide tier at 120 columns. | Cols | Field | Width | Align | Notes | |---|---|---|---|---| | 1 | pad / **selection** | 1 | | blank, or `▸` on the selected row (§7.11) | | 2 | state dot | 1 | | `●` live / `◐` idle / `○` stale / blank unknown | | 4–5 | vendor | 2 | left | `CC` / `CX` | | 7 | separator | 1 | | dim `│` | | 9–67 | **session** | **W−61** | left | flexes; `…` truncation | | 70–82 | model | 13 | left | normalized display name | | 85–96 | context gauge | 12 | | see §7.4 | | 98–103 | context % | 6 | right | 6, not 5: a derived value carries a `~` marker | | 106–112 | cost | 7 | right | | | 114 | separator | 1 | | dim `│` | | 116–119 | age | 4 | right | | | 120 | pad | 1 | | | Only two `│` separators per row, deliberately. They cut the row into three zones — **identity** (dot, vendor), **measurement** (name, model, gauges, cost), **time** (age). A pipe between every column reads as a spreadsheet; two pipes read as structure. **Session label content.** The session's own `name` if the vendor has one, else the workspace basename, else the vendor session id. Then the sub-agent chip if the session is fanning out (§7.13). Then, only if ≥14 cells remain free, two spaces and the parent directory (left-elided with `…`). The parent path disambiguates same-named projects under different roots and stops the wide tier from opening a dead gulf between the name and the model. It drops out automatically as the terminal narrows. The chip's width is reserved **before** the name is truncated, and the name loses the character. A chip that vanished on a long project name would make the same session look like a different kind of session at a different terminal width — a lie by omission — and the name is the field that can afford to lose a character, because the parent path and the detail pane both still carry the identity. > **Deviation from the original spec, deliberate:** the `⌥worktree` mark is not rendered. > `worktree.name` exists only on the statusline's stdin payload, which the HUD does not > consume, so **no adapter can source it** — a cell no adapter can fill is a cell that > should not be in the grid. The glyph and its ASCII form stay in `internal/hud/glyphs.go` > for the day a vendor writes it to disk. **Responsive tiers.** Breakpoints are on width only; the shedding order is fixed, so the layout at any width is a pure function of the width. | Tier | Width | Model | Gauge | Cost | `SESSION` width | At 120 / 80 / 72 | |---|---|---|---|---|---|---| | wide | ≥ 100 | 13 | 12 cells | shown | W − 61 | 59 | | compact | 80–99 | 13 | 8 cells | **dropped** | W − 48 | 32 | | narrow | 60–79 | 13 | **dropped** | dropped | W − 38 | 34 | | floor | < 60 | — | — | — | — | one-line notice | Cost sheds before the gauge, and the gauge before the model, because the gauge is a redundant encoding of a number that stays on screen, and the model is identity — the answer to "which of my agents is this?" — which nothing else supplies. `MODEL` is never narrowed below 13: `gpt-5.1-codex` is exactly 13 columns, and truncating a model name to `gpt-5.1-c…` destroys the one field a user scans for. Height tiers: **H ≥ 9** full chrome; **6 ≤ H < 9** drops both rules and the column-header row (header + rows + footer only); **H < 6** shows the floor notice. Row overflow is not paginated — the footer gains `+3 more`, and the viewport **follows the selection**: `↑`/`↓` move the cursor and the visible window slides to keep it on screen. That arithmetic lives in `Render` rather than in `Update`, because `Update` does not know how tall the chrome is this frame and a second copy of the height maths is a second thing to get wrong. Floor renders, exactly: ``` telltale needs 60 columns (have 52) ``` ``` telltale needs 6 rows (have 4) ``` **Column auto-hide.** A column that would render `—` for *every* visible row is dropped entirely and its width returned to `SESSION`; a full column of dashes is noise, not information. This applies to `CONTEXT` and `COST` only (never to `MODEL` or `AGE`), is computed per frame from the visible rows, and is therefore deterministic. The help overlay lists any column hidden this way and why. ### 7.3 Target renders These are the golden-test targets, pasted from the generated files. All data is synthesized. **A — wide, healthy (120 cols).** The reference render. Rendered with a *synthetic* vendor that declares every capability, so the whole grid is exercised; render I below shows what the real v1 capability mix produces. ``` telltale │ 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL CONTEXT COST AGE ● CC │ telltale C:\src\code Opus 5 █████████▎── 84.2% $2.41 │ 12s ● CC │ acme-api C:\src\work Sonnet 4.5 ████▌─────── 41% $0.18 │ 48s ◐ CX │ notes-api C:\src\code gpt-5.1-codex — — │ 4m ○ CC │ learning-notes C:\src\code Haiku 4.5 ██████████▏─ 92.6% $11.07 │ 22m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ``` Row 3 is the honest-gauge case in its normal habitat: a session whose adapter can source a model and an age but not context or cost. Blank gauge field, `—` in both numeric columns. Nothing about it looks like zero. **B — compact (80 cols).** Cost gone, gauge halved, quota wrapped to its own line. ``` telltale │ 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h ────────────────────────────────────────────────────────────────────────────── SESSION MODEL CONTEXT AGE ● CC │ telltale C:\src\code Opus 5 █████▉── 84.2% │ 12s ● CC │ acme-api C:\src\work Sonnet 4.5 ██▉───── 41% │ 48s ◐ CX │ notes-api C:\src\code gpt-5.1-codex — │ 4m ○ CC │ learning-notes C:\src\code Haiku 4.5 ██████▌─ 92.6% │ 22m ────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage ? keys ``` **C — narrow (72 cols).** Gauge gone; the number it encoded stays. Vendor names shorten. ``` telltale │ 4 sessions │ cc 3 cx 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h ────────────────────────────────────────────────────────────────────── SESSION MODEL CTX AGE ● CC │ telltale C:\src\code Opus 5 84.2% │ 12s ● CC │ acme-api C:\src\work Sonnet 4.5 41% │ 48s ◐ CX │ notes-api C:\src\code gpt-5.1-codex — │ 4m ○ CC │ learning-notes C:\src\code Haiku 4.5 92.6% │ 22m ────────────────────────────────────────────────────────────────────── q quit / find ? keys ``` **D — degraded rows (120 cols).** Four distinct failure shapes in one frame. Rows are sorted by activity, so they do not appear in the order they are described. ``` telltale │ 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL CONTEXT COST AGE ● CC │ telltale C:\src\code Opus 5 ──────────── 0% $0.04 │ 3s ● CC │ a-really-long-project-name-that-overflows-the-label-column… Opus 5 ███████████─ 99.9% $340.50 │ 9s ◐ CX │ 4f2a9c81-1d3e-4a77-9b02-000000000000 — — │ 7m CC │ acme-api C:\src\work Sonnet 4.5 — $1.02 │ — ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ``` Row 1 is at exactly 0% and draws a full track. Row 2 is the label-overflow case, truncated at `…` with the grid intact. Row 3 is a session discovered by filename whose only record was torn, so nothing parsed — the label falls back to the session id, the model cell is blank and every sourced field is `—`. Row 4 has a record timestamp in the future (clock skew): its `AGE` is `—`, never a negative or a zero, and because its liveness is `unknown` **its state dot is blank, not `○`** — unknown is not a claim of staleness. **J — zero versus absent (120 cols).** The single assertion the build exists to protect. ``` telltale │ 2 sessions │ claude 2 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL CONTEXT COST AGE ● CC │ at-zero C:\src\code Opus 5 ──────────── 0% $0.00 │ 5s ● CC │ no-source C:\src\code Opus 5 — — │ 6s ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ``` `0%` is a full track of `────────────`; absent is whitespace. If these two rows ever render the same, the build fails. **E — stale scan (120 cols).** The scan has been failing for 47 seconds. Values are the last ones actually measured; the whole row area renders `Muted` (invisible in a plain golden, asserted separately), and the footer's right slot carries the notice. The header is never used for notices — it holds identity and quota only, which keeps it from overflowing at any width. ``` telltale │ 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL CONTEXT COST AGE ● CC │ telltale C:\src\code Opus 5 █████████▎── 84.2% $2.41 │ 12s ● CC │ acme-api C:\src\work Sonnet 4.5 ████▌─────── 41% $0.18 │ 48s ◐ CX │ notes-api C:\src\code gpt-5.1-codex — — │ 4m ○ CC │ learning-notes C:\src\code Haiku 4.5 ██████████▏─ 92.6% $11.07 │ 22m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── ⚠ last scan 47s ago Access is denied. ``` **F — filter and sort active (120 cols).** The header count reads `3 of 4` so it cannot contradict the per-vendor totals beside it. Non-default filter/sort is stated in the footer, because a monitor that silently hides rows is a liar. ``` telltale │ 3 of 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL CONTEXT COST AGE ○ CC │ learning-notes C:\src\code Haiku 4.5 ██████████▏─ 92.6% $11.07 │ 22m ● CC │ telltale C:\src\code Opus 5 █████████▎── 84.2% $2.41 │ 12s ● CC │ acme-api C:\src\work Sonnet 4.5 ████▌─────── 41% $0.18 │ 48s ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys filter claude sort context ``` **G — empty (120 cols).** Distinguishes "watching, found nothing" from "vendor not installed". Two different facts; two different words; never a fake row and never an error dialog. ``` telltale │ 0 sessions ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── no active sessions agy not detected %USERPROFILE%\.gemini\antigravity-cli claude watching %USERPROFILE%\.claude\projects codex not detected %USERPROFILE%\.codex cursor not detected %APPDATA%\Cursor\User gemini not detected %USERPROFILE%\.gemini\tmp ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ``` The vendor status word is one of exactly four: `watching` (directory exists and is readable), `not detected` (directory absent), `unreadable` (the vendor's data is there and the adapter cannot read it — an OS refusal, or a store whose schema the adapter does not recognize (§3.9); rendered `SevWarn` with the reason appended), `drifted` (the store opened and read, and at least one session's read could not find the structure the adapter was verified against — `internal/adapter/drift`; also `SevWarn`). On the dev machine today the Codex line reads `not detected`, since `~/.codex` is absent (§3.2). The third word: ``` telltale │ 0 sessions ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── no active sessions agy not detected %USERPROFILE%\.gemini\antigravity-cli claude unreadable %USERPROFILE%\.claude\projects Access is denied. codex not detected %USERPROFILE%\.codex cursor not detected %APPDATA%\Cursor\User gemini not detected %USERPROFILE%\.gemini\tmp ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ``` `drifted` is the word the other three cannot say. Those three are all answers from `Discover`, known before a single session is read. Drift is knowable only *after* the read: it is the finding that a store which opened, listed and parsed no longer carries the structure the adapter's readings hang off. Calling that `unreadable` would borrow the word for "the OS refused" to describe a store the OS handed over intact — the same collapse `zero` and `absent` are kept apart to avoid (§4a.1). It renders with its **scope** appended: how many of the vendor's sessions reported drift, out of how many this scan read. One of forty-one is a vendor mid-rollout; forty-one of forty-one is a format that moved under the whole store, and the word alone cannot tell those apart. The scope deliberately does **not** reuse the header's `n of m sessions` sentence — the header counts visible-of-total across every vendor, this counts drifted-of-read for one vendor, and the two land on the same screen. In the borrowed grammar the vendor line would read as a claim about how many sessions are *showing*, which the header directly contradicts; naming what the numerator counts is what keeps them apart. The scope is the only part of the line that gives way when the width runs out, the same way the grid sheds `COST`; the word never does. This state needs sessions to exist and every one of them to be hidden — below, by the 8-hour idle cutoff — because a vendor cannot drift without having produced the sessions that revealed it. The ordinary case is render M. The fourth word: ``` telltale │ 0 of 2 sessions │ codex 2 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── no active sessions agy not detected %USERPROFILE%\.gemini\antigravity-cli claude not detected %USERPROFILE%\.claude\projects codex drifted %USERPROFILE%\.codex 1 drifted of 2 read cursor not detected %APPDATA%\Cursor\User gemini not detected %USERPROFILE%\.gemini\tmp ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ⚠ codex drifted ``` **H — help overlay (120 cols).** Replaces the row area rather than floating over it; a floating panel on a monitor obscures the thing being monitored. ``` telltale │ 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit (also ctrl+c) ↑/↓ move the selection (also j / k) enter open the detail pane for the selected session u what each vendor has left, and what it spent w this week: the fleet's slow windows only / find: narrow rows by name or path esc close the pane, or cancel the find, or quit v vendor: all > claude > codex > gemini > agy > cursor s sort: activity > context > cost a show all (include sessions idle > 8h) r rescan now ? close this help ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── ? close ``` **I — the v1 capability mix (120 cols).** What the real adapters actually render today: Claude sources neither context nor cost from disk, Codex sources a **derived** context percentage (marked `~`) and real quota windows. Nothing sources cost, so the `COST` column auto-hides and its width returns to `SESSION`. The `AG` row carries no `Name` at all — the only free text on agy's disk is prompt content (§3.8), so this field is `CapNone` and the HUD falls back to the workspace basename, the same fallback a Gemini row takes when it has no summary of its own (ruled 2026-08-12). Its `MODEL` cell truncates because the vendor's display string is 23 characters against a 13-column cell — both are what the HUD really shows. The `CU` row is the one to read next to the `CX` row: both carry a context bar, and only one of them carries a `~`, because Cursor persists its own `contextUsagePercent` and telltale reads it rather than computing one (§3.9). ``` telltale │ 5 sessions │ claude 1 codex 1 gemini 1 agy 1 cursor 1 codex 5h ██████▎─ 88.4% ↻ 3h02m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL CONTEXT AGE ● CC │ telltale C:\src\code Opus 5 — │ 12s ● CU │ multi-vendor orchestration C:\src\code composer-2.5 ████▏─────── 37% │ 1m ● CX │ example-app C:\src\code gpt-5.1-codex ███████▋──── ~69.8% │ 1m ● AG │ example-app C:\src\code Gemini 3.6 F… — │ 2m ◐ GE │ glossary tooltips ⑂~2 c:\src\code gemini-3-pro — │ 3m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ``` This render is the one to look at when judging §6 Q7. The ragged CONTEXT column is the cost of option (2); it is honest, and whether it is *legible* is a dogfood question. **K — every column hidden (120 cols).** No visible row reports context or cost. ``` telltale │ 3 sessions │ claude 2 codex 1 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL AGE ● CC │ telltale C:\src\code Opus 5 │ 12s ◐ CX │ notes-api C:\src\code gpt-5.1-codex │ 4m ○ CC │ learning-notes C:\src\code Haiku 4.5 │ 22m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ``` **L — ASCII glyph mode (120 cols).** `--ascii`, `TELLTALE_ASCII=1`, or a non-terminal output target. Absent renders `n/a`; the gauge loses its eighth-cell partials, which is a real precision loss in the bar and acceptable only because the number beside it carries the precision. ``` telltale | 4 sessions | claude 3 codex 1 claude 5h ###----- 42% ~ 2h13m 7d #------- 18% ~ 5d02h ---------------------------------------------------------------------------------------------------------------------- SESSION MODEL CONTEXT COST AGE * CC | telltale C:\src\code Opus 5 #########--- 84.2% $2.41 | 12s * CC | acme-api C:\src\work Sonnet 4.5 #####------- 41% $0.18 | 48s o CX | notes-api C:\src\code gpt-5.1-codex n/a n/a | 4m . CC | learning-notes C:\src\code Haiku 4.5 ##########-- 92.6% $11.07 | 22m ---------------------------------------------------------------------------------------------------------------------- q quit / find enter detail u usage w week v vendor s sort a all ? keys ``` **M — shape drift (120 cols).** A store that reads fine and no longer matches. Every row here is render A's except the Codex one, whose read found no `session_meta` record — so everything that record feeds is absent, and the row renders *exactly* as it would if Codex simply had nothing to say. That is the failure: nothing in the grid can tell those two apart, and the footer notice is the only thing on screen that knows. It is also why the fourth vendor word needs a second home. The vendor line renders in the empty state only, and a vendor cannot drift without having produced sessions — so the screen drift actually happens on is this one, where the vendor line is not present at all. `driftNotice` therefore renders under **every** body: grid, empty state, help overlay and detail pane alike. A warning that came and went with whichever pane was open would be one a reader could not trust to be there. ``` telltale │ 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL CONTEXT COST AGE ● CC │ telltale C:\src\code Opus 5 █████████▎── 84.2% $2.41 │ 12s ● CC │ acme-api C:\src\work Sonnet 4.5 ████▌─────── 41% $0.18 │ 48s ◐ CX │ 00000000-bbbb-4ccc-8ddd-000000000001 — — │ 4m ○ CC │ learning-notes C:\src\code Haiku 4.5 ██████████▏─ 92.6% $11.07 │ 22m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ⚠ codex drifted ``` ### 7.4 The gauge Glyphs: fill `█` (U+2588) with eighth-block partials `▏▎▍▌▋▊▉` (U+258F–U+2589), track `─` (U+2500). Full-height fill over a mid-height rule reads as a level above a baseline and keeps the row visually quiet; a shaded `░` track reads as texture and fights the text. Three rules, all of which exist to stop the gauge from lying: 1. **The last cell is reserved below 100%.** Fill is computed over `cells−1`, so a value of 99.9% always leaves one visible track cell and only an exact 100% fills the bar. A 92.6% bar that renders as solid is a gauge claiming "full" when it is not. 2. **Any nonzero value draws at least one eighth.** 0.4% must not be pixel-identical to 0%. 3. **Absent draws nothing.** Not an empty track — nothing. (Principle #1.) Verified scale at 12 cells (`TestGaugeScale` pins every row): ``` 0% ──────────── 0% 84.2% █████████▎── 84.2% 0.4% ▏─────────── 0.4% 92.6% ██████████▏─ 92.6% 5% ▌─────────── 5% 99.9% ███████████─ 99.9% 25% ██▊───────── 25% 100% ████████████ 100% 50% █████▌────── 50% absent — ``` **Number formatting**, shared with the statusline via `internal/theme`: - Percent: floored to one decimal, never rounded up (a usage gauge must not overstate); whole numbers drop the decimal; `100%` has no decimal. Guarantees a 5-column field, and the cell is 6 so a derived value can carry its `~`. - Cost: `$0.00` under $1000, `$1234` at or above. - Age: `12s` / `47m` / `2h` / `3d`, capped at 4 columns. Sub-hour precision is where a monitor's value is; `2h13m` precision belongs to the quota countdown, not to row age. - Countdown: `↻2h13m`, `↻47m`, `↻5d02h`. The days branch matters: without it a seven-day window renders `↻120h00m`. The statusline still uses its own `pct` and `shortDur`; see the divergence note in §2. ### 7.5 Colour and threshold tokens Thresholds are the statusline's, unchanged and now literally shared: **green < 60, yellow ≥ 60, red ≥ 85** from `theme.WarnPct` / `theme.CritPct`. The palette is deliberately the terminal's own 4-bit ANSI palette rather than hex truecolor. Reason: telltale then inherits whatever theme the user already chose, looks native in Windows Terminal's default scheme and in a light-background scheme without a second palette, and matches the statusline byte-for-byte in intent. Total palette: four hues, one attribute, and the default foreground. | Token | Meaning | ANSI | Statusline (raw) | HUD (lipgloss v2) | |---|---|---|---|---| | `Text` | primary values | default | *(unstyled)* | `NewStyle()` | | `Muted` | chrome, labels, rules, de-emphasis | — | `\x1b[2m` | `NewStyle().Faint(true)` | | `Identity` | model name, vendor tag | 6 | `\x1b[36m` | `Foreground(Color("6"))` | | `SevOK` | value < 60; healthy notices | 2 | `\x1b[32m` | `Foreground(Color("2"))` | | `SevWarn` | value ≥ 60; warning notices | 3 | `\x1b[33m` | `Foreground(Color("3"))` | | `SevCrit` | value ≥ 85; error notices | 1 | `\x1b[31m` | `Foreground(Color("1"))` | | `Track` | unfilled gauge cells | 7 / 8 | n/a | `Foreground(lightDark(Color("7"), Color("8")))` | Semantic aliases, so intent is greppable rather than inferred: `Absent() = Muted`, `Rule() = Muted`. Hue owns exactly one meaning: **cyan is identity, the green/yellow/red ramp is severity, faint is de-emphasis.** Nothing else gets a colour. In particular the state dot encodes liveness by *glyph and intensity* (`●` Text / `◐` Text / `○` Muted), never by hue — green already means "under 60%", and one hue meaning two things is how a colour system rots. **Shared code, without dragging Lipgloss onto the fast path.** ADR-002 requires the statusline to stay stdlib-only and never initialize Bubble Tea. So `internal/theme` holds only numbers and names — `WarnPct`, `CritPct`, the ANSI indices, and the shared format helpers — and no `Style` type at all. `internal/statusline` maps those indices to escape codes as it does today; `internal/hud/style.go` maps them to `lipgloss.Style` values. One source of truth for the thresholds, zero coupling of the statusline's **import graph** to the TUI stack. **"Import graph", not "binary" — corrected 2026-08-16.** This paragraph used to claim zero coupling of the statusline *binary*, and the shipped artifact refutes it: `go version -m telltale.exe` lists `charm.land/bubbletea/v2 v2.0.8` and `charm.land/lipgloss/v2 v2.0.5`, because §1 ships ONE binary and `telltale hud` is in it. §9.8 had already measured the consequence (the 14 MB binary costs ~54 ms per gated call, against 36.2 ms for a small one) without this sentence being brought into line. What ADR-002 actually buys, and all it buys, is that the code `telltale statusline` reaches never touches the framework: no renderer is constructed, no program is started, neither module's package init runs. That is now gated rather than asserted — see §5's 2026-08-16 amendment. **Light and dark backgrounds.** Lipgloss v2 removed `AdaptiveColor` and the global renderer, so adaptation is explicit: `Init()` lifts `tea.RequestBackgroundColor()` into a `Cmd`, `Update` handles `tea.BackgroundColorMsg` and calls `msg.IsDark()`, and the style set is rebuilt with `lipgloss.LightDark(isDark)`. Only `Track` consumes it (light gray on light backgrounds, dark gray on dark). Terminals that never answer the OSC query leave the default: assume dark. Because exactly one token depends on it and no layout does, golden layout tests are unaffected by which branch is taken — and `TestBackgroundColorRebuildsTheStyleSetWithoutMovingTheLayout` enforces that. **NO_COLOR.** `colorprofile` caps the profile at `Ascii` when `NO_COLOR` is set; Bubble Tea v2 downsamples internally, so no telltale code path is involved. Under `Ascii` every `Foreground` disappears while `Faint` survives — chrome still recedes. Nothing is lost, because by principle #2 colour was never the sole carrier of any distinction. No `--no-color` flag of our own: one mechanism, the standard one. **ASCII glyph mode** is a *separate* switch from colour — `--ascii`, or `TELLTALE_ASCII=1`. For legacy consoles and non-UTF-8 code pages: | Unicode | ASCII | | Unicode | ASCII | |---|---|---|---|---| | `●` `◐` `○` | `*` `o` `.` | | `─` (light rule / track) | `-` | | `━` (heavy rule) | `=` | | | | | `█` + eighths | `#` (no partials) | | `│` | `\|` | | `—` (absent) | `n/a` | | `…` | `>` | | `↻` | `~` | | `⌥name` | `(name)` | | `⚠` | `!` | | spinner | `-\|/` rotation | ### 7.6 Motion **The rule: telltale may animate its own work; it must never animate the vendor's.** Everything follows from it. A spinner on a session row would assert "this agent is working right now" — a claim telltale cannot source, since the adapters read files on disk and know a last-write timestamp, not liveness. That is a narrated animation, which is the honest-gauge violation in motion form. A tweened gauge is worse: every intermediate frame displays a value no vendor ever reported. **Animates — one thing.** A braille spinner `⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏` at 10 fps in the header, only while the *first* scan is in flight, and only once that scan exceeds 250 ms. It reports telltale's own I/O, which telltale is entitled to describe. After the first successful scan it never appears again; later slow scans surface as staleness (§7.7), not motion. `TestSpinnerStopsForeverAfterTheFirstScan` pins that. **Never animates.** Numbers. Gauge fills. Colour pulses or flashes. Row insertion and removal. Sort transitions. Per-row spinners or activity indicators of any kind. **Cadence.** Poll every 1000 ms via `tea.Tick`. Rows re-sort only when the underlying values change, and never as a tie-break wobble — the sort comparator falls back to session id so equal keys hold a stable order frame to frame. Steady-state churn budget is principle #4: the `AGE` cell of sessions younger than 60 s, and nothing else. The footer's scan-freshness notice appears *only when abnormal* (> 3 s), precisely so the healthy screen has no ticking element in it. The 1 s cadence is affordable because the poll is stat-first: `Discover` lists and stats only, and `Read` touches a bounded head and tail rather than a whole transcript. Any further tail-read optimization must honour §4 — a backward seek can land mid-record, so the first partial record after a seek is discarded, not parsed. `internal/jsonl.Tail` already does this. **Bubble Tea v2 implications** (verified against v2.0.8): - `Model` is `Init() Cmd`, `Update(Msg) (Model, Cmd)`, `View() View`. `View` is a struct, so alt-screen and cursor are view state, not program options: return a `tea.View` with `AltScreen: true`, `Cursor: nil`, `WindowTitle: "telltale"` (`--no-title` suppresses the title). There is no `WithAltScreen` in v2. - `tea.RequestBackgroundColor()` is a `Msg`, not a `Cmd`; it has to be lifted into a `func() tea.Msg` before `tea.Batch` will take it. - **Scanning never happens inside `Update`.** The tick dispatches a `tea.Cmd`; the scan returns a `scanResultMsg`. At most one scan is in flight and ticks arriving during a scan are dropped. This is a Windows correctness requirement, not tidiness: a `stat` against a disconnected network path blocks, and a blocked `Update` freezes input including `q`. - **`View()` is pure and never calls `time.Now()`.** The model carries `now`, stamped when the tick arrives — the same discipline as `statusline.Options.Now`. This is what makes the renders in §7.3 testable at all. - Do not raise the renderer FPS. Frames are byte-identical between data changes, so the framerate cap is not the thing limiting redraws — the data is. - Render nothing until the first `tea.WindowSizeMsg` arrives. One blank frame beats one frame of wrong layout. ### 7.7 Degraded and empty states Each row below is an eval fixture. Fixtures are synthesized — fake session ids, fake paths, fake text — never copied from real transcripts. | Fixture | Condition | Where it is asserted | Assertion | |---|---|---|---| | `zero-vs-absent` | one session at 0%, one with no context source | golden J + `TestAbsentGaugeIsNotAnEmptyTrack` | The two rows differ. The build fails if a gauge cannot tell "no data" from "zero". | | `degraded` row 3 | adapter sources model + age only | golden D | Blank gauge field, `—` in `CONTEXT` and `COST`; no `0`, no stale carry-over. | | torn tail | file ends mid-record, no trailing `\n` | `TestTornTailChangesNothing` (both adapters) | The partial line is held, never parsed (§4). A torn tail changes nothing on screen. | | torn-only | the file's *only* record is torn | `TestTornOnlyRecordStillListsWithEverythingAbsent`, golden D row 3 | Session still listed (discovered by filename), label falls back to session id, every sourced field `—`, age from file mtime. | | clock skew | record timestamp in the future | `TestFutureMtimeDegradesRatherThanClampingToZero`, golden D row 4 | `AGE` = `—`. Never a negative duration, never `0s`. Liveness is `unknown`, so the dot is blank. | | `stale-scan-47s` | scan older than 3 s | golden E + `TestStaleScanDimsTheRowArea` | Values are retained (they were true at the displayed `AGE`) but nothing renders at full intensity. | | `stale-scan-90s` | scan failing over 60 s | golden | Notice escalates to `SevCrit` and the header quota goes `Muted` too. Quota is as stale as everything else and must not look fresh. | | `empty-watching` | vendor dirs readable, no sessions | golden G | `watching` + the path actually checked, home-redacted. | | vendor missing | `~/.codex` absent | golden G, `not detected` + `TestVendorAbsentBecomesNotDetected` | No fake row, no error state; the other vendor still renders. | | `empty-unreadable` | dir exists, OS refuses | golden + `TestUnreadableVendorKeepsTheOSMessage` | Third word, distinct from the other two, with the OS message in `SevWarn`. | | `quota-absent` | API-key login, no `rate_limits` | golden + `TestNullRateLimitsYieldNoQuotaWindows` | Header quota block absent. Mirrors the statusline's load-bearing test — never `5h 0%`. | | `degraded` row 2 | label longer than the column | golden D | Truncation at `…`, grid intact. | | `column-hidden` | `CONTEXT` and `COST` absent for every visible row | golden K | Columns dropped, width returned to `SESSION`; help overlay names them. | | `floor-width` / `-height` | 52 cols / 4 rows | goldens | One line, no partial grid. | | gauge scale | the §7.4 table | `TestGaugeScale` | Exact glyph string per value: the reserve-last-cell and min-eighth rules. | | separator injection | a session name containing U+2028/U+2029 | `TestSessionNameSeparatorsCannotTearTheGrid`, `TestDetailPaneSanitizesModelAuthoredText`, `TestFindQueryCannotTearTheFooter` | The character never reaches the frame — grid, pane or footer — and no line exceeds the terminal width. | Added in v1.1: | Fixture | Condition | Where it is asserted | Assertion | |---|---|---|---| | `detail-pane` | pane over a Claude row, real capability table | golden + `TestDetailPaneSeparatesCantKnowFromAbsentNow` | Fields Claude declares `CapNone` get **no line**; they are named once on `not sourced`. | | `detail-degraded` | pane over a session whose records did not parse | golden + `TestDetailPaneShowsDegradedFieldsAndDiagnostics` | Degraded field names and every diagnostic are on screen; a declared-but-empty quota is `—`, never `0%`. | | clean session | no degraded fields, no diagnostics | `TestDetailPaneStatesTheAbsenceOfProblems` | The honesty block says `—` rather than going blank; a blank block is indistinguishable from a pane that forgot to render it. | | measured zero fan-out | `Subagents = 0` | `TestDetailPaneStatesAMeasuredZeroFanOut`, `TestSubagentChipOnlyAppearsForANonzeroCount` | Grid draws **no chip**; the pane says `~0 recent`. | | uncountable fan-out | sidecar unreadable | `TestDetailPaneRendersAnUncountableFanOutAsAbsent` | `—`, never `0`. | | selection vanishes | the selected session ends mid-poll | `TestASelectedSessionThatVanishesClosesThePane`, `TestDetailPaneSaysSoWhenItsSessionIsGone` | The pane closes rather than retargeting; an out-of-range cursor says "no longer listed". | | re-sort under the cursor | a bottom row becomes the newest | `TestSelectionFollowsTheSessionNotTheIndex` | The selection follows the **session key**, not the index. | | `row-grammar` | selection mark + fan-out chips | golden + `TestSelectionIsAGlyphNotAHighlight` | Selection is a glyph in the pad column, not reverse video. | | chip vs. truncation | a 73-character session name | `TestSubagentChipSurvivesLabelTruncation` | The chip survives at every width; the name gives way. | | `burn-forecast` | 7 samples over 18 min on one window, a near-flat second window | golden + `TestForecastArithmeticIsPinned` | Exact projected time and basis; the slow window renders **nothing**. | | below basis | < 3 samples, or a span < 5 min | `TestForecastRefusesToProjectBelowTheMinimumBasis`, `TestNoForecastRendersWithoutABasis` | Nothing renders. Not a placeholder, not a dash — the header cell simply ends. | | window rollover | usage drops, or `resets_at` jumps a window forward | `TestUsageDropClearsTheSamples`, `TestResetsAtJumpClearsTheSamplesButJitterDoesNot` | The buffer clears; three seconds of `resets_at` jitter does not clear it. | | `find-active` | find mode with a query typed | golden | The footer becomes the query line and says how to leave. | | `find-applied` | query applied, mode left | golden + `TestAnAppliedQueryAlwaysAnnouncesItself` | Header reads `2 of 4`; footer keeps naming the query. | | query hides everything | a query matching no row | `TestAnEmptyResultNamesTheQuery` | The empty state names the query rather than saying "no active sessions". | | over-long query | 156 characters at 60–120 cols | `TestALongQueryIsTruncatedNotDropped` | Truncated with `…`, never pushed off the footer — a query that vanished while still filtering is the silent row-hiding the footer exists to prevent. | | query with a trailing space | `"acme "` | `TestTheDisplayedQueryIsTheQueryBeingMatched` | The string on screen is the string being matched; the display is not trimmed. | Added with shape-drift reporting: | Fixture | Condition | Where it is asserted | Assertion | |---|---|---|---| | `shape-drift` | a read reports drift; every row still renders | golden M + `TestDriftIsVisibleOnTheGridNotOnlyInTheDetailPane` | The grid is unchanged and the footer carries `⚠ drifted`. A healthy frame never mentions drift, and the notice survives `--ascii` and `NO_COLOR` as a word. | | `empty-drifted` | sessions exist, all past the idle cutoff | golden + `TestTheDriftScopeCannotBeReadAsTheHeaderCount` | The fourth vendor word, with its scope in the slot `unreadable` gives to the OS message — and in a grammar the header's own count cannot be mistaken for. | | partial drift | 1 of 41 sessions reports drift | `TestOneDriftedSessionDriftsTheVendor`, `TestTheVendorLineStatesHowMuchOfTheStoreDrifted` | **Any** drifted session drifts the vendor, and the counts travel with the word. Every row still renders. | | drift under a failed `Discover` | vendor absent, or the OS refuses | `TestTheDiscoverTierStillWinsOverDrift` | `not detected` and `unreadable` are untouched: the roll-up only runs where `Discover` succeeded, so the ordering is structural rather than a comparison. | | reworded drift note | `drift.Watch` changes its wording | `TestDriftIsRecognizedFromTheNoteTheAdapterLayerActuallyWrites` | The HUD reads drift off `Diagnostics` text, which the compiler cannot check. The test folds a real `drift.Watch` so a rewording fails the build instead of silencing the vendor line. | | notice pile-up | 60 cols with a 24-char query, a filter, a sort, a stale scan **and** drift | `TestADriftedFrameStillFitsEveryTier`, `TestTheFooterGivesUpItsCheapestNoticesFirst` | Every line fits the terminal. Whole notices are dropped, cheapest first, and `…` says so. | Freshness escalation, stated once: **≤ 3 s** normal; **> 3 s** row area `Muted` + footer notice in `SevWarn`; **> 60 s** notice in `SevCrit` and the header quota goes `Muted` too. Retained values are not "presented as fresh" in any of these, because the age of the measurement is on screen next to them — that is the condition the honest-gauge rule actually imposes. Notice priority, stated once: the footer cannot always hold every notice — `joinEnds` has no truncation path, so a block that does not fit runs off the end of the terminal. The block is therefore fitted first, by dropping **whole** notices cheapest-first and prefixing what survives with `…`. The order is `sort`, `+N more`, `filter`, `find`, the stale-scan warning, drift. That is the rule `joinEnds` already applies between the key hints and the notice block, asked one level down: what survives is what the reader cannot find out anywhere else on this screen. `sort` hides nothing at all; `+N more` sits above a row area the reader can see is full; a filter and a query hide rows silently, but the header's `N of M sessions` still declares *that* rows are hidden, so only the cause is lost; a stale scan re-announces itself every tick and clears the moment a scan succeeds; drift does neither, and is the last to go. A single notice wider than the whole line is truncated rather than dropped — an ellipsis on a warning still says a warning is there, and a footer that dropped its last one would quietly claim nothing is wrong. #### The zero-config first frame — measured 2026-08-15, then narrowed The empty states above are all about a machine telltale already lives on. This subsection is about the frame before that: what a stranger sees on the first run, with nothing configured and no vendor store anywhere. **It was measured before anything was built.** A clean profile was made by pointing `HOME`, `USERPROFILE`, `APPDATA` and `LOCALAPPDATA` at an empty directory and running the real binary. The isolation was verified rather than assumed — `telltale snapshot` reported `vendors_not_detected: 6` and `sessions: 0`, so no real session leaked into any reading below. | Mode | What a stranger actually saw | Verdict | |---|---|---| | `telltale` (bare) | all 203 lines of `usageText`, on **stderr**, exit **2** | **fail** — true and useless. Eight modes, no start-here, and a failure code for typing the binary's own name | | `telltale hud` | `no active sessions`, then six vendors as `not detected` with the path checked for each | **partial** — every word true and complete, and silent on the reader's next question | | `telltale doctor` (vendors present) | five seats, binary path, version, and `not checked` said out loud for auth and network | **pass** on truth, no next step | | `telltale doctor` (bare `PATH`) | five `FAILED` rows under `0 checks passed, 5 failed` | **pass** on truth, and the frame most likely to read as telltale being broken | | `telltale council` | the room opens; unseatable seats fold out and `collapsedNotice` names each one and why | **pass** — not the zero-config entry point | The frame the HUD draws on that profile, generated by the build like every render in §7.3 — this is `empty-nothing-detected`, and the last line is what this subsection added: ``` telltale │ 0 sessions ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── no active sessions agy not detected %USERPROFILE%\.gemini\antigravity-cli claude not detected %USERPROFILE%\.claude\projects codex not detected %USERPROFILE%\.codex cursor not detected %APPDATA%\Cursor\User gemini not detected %USERPROFILE%\.gemini\tmp grok not detected %USERPROFILE%\.grok\sessions pi not detected %USERPROFILE%\.pi\agent\sessions telltale doctor checks the vendor binaries; this screen reads their stores ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ``` **The HUD's gap was not the empty state; it was one question the empty state cannot answer.** The HUD reads STORES. A vendor CLI that is installed and has never been run has no store, so `not detected` six times is the same picture on a bare machine and on a machine holding five unopened vendors. `telltale doctor` resolves the BINARIES, which is exactly the missing measurement — so the frame now names it. Three changes, and the narrowing is the design: 1. **A bare `telltale` prints a short first frame on stdout and exits 0.** A no-argument run is a first run, not a usage error. An unknown SUBCOMMAND keeps stderr and exit 2, because that one is an error and the manual is its correction. `telltale help` is new and reaches `usageText` — before this, every route to the manual was an error path, so the frame could not point at it without inventing a command. 2. **The HUD's empty state carries one pointer line, in one case only:** every vendor `not detected`, nothing watching, nothing drifted, nothing unreadable. A watching store needs no remedy; a drifted one needs the adapter re-verified; a store the OS refused wants a permission fixed, and doctor would report that vendor's binary `ok` and lead the reader away from the answer. `empty-watching`, `empty-drifted` and `empty-unreadable` are byte-identical after this change, which is the check that the narrowing held. 3. **`doctor` closes with a `What runs next` paragraph**, branching on the seat count the summary above already printed and claiming nothing more. The no-seat branch says the report is working rather than telltale failing, and still leaves a command that runs. **Nothing invented, at either end (ADR-001).** `firstFrameText` prints before `main` has stat'd a store or resolved a binary, so it asserts nothing about the machine at all — it names modes and points at the one that measures, and `TestTheFirstFrameClaimsNothingAboutThisMachine` keeps it that way. `nextStep` stays pure over its `Report` and never touches auth or network, which are `not checked` on every seat, always. | Fixture | Condition | Where it is asserted | Assertion | |---|---|---|---| | `empty-nothing-detected` | six vendors, every one `not detected` | golden + `TestTheZeroConfigEmptyStateNamesDoctor` | The frame names `telltale doctor` — the measurement this screen cannot make. | | the other three empty states | watching / drifted / unreadable | `TestOnlyTheNothingDetectedFrameNamesDoctor` | None of them names doctor. A pointer on every frame is a pointer meaning nothing. | | no vendor looked for | an empty vendor slice | `TestAnEmptyVendorListIsNotNothingDetected` | Not the same fact as every vendor missing (§4a.1, one level up from the gauge). | | narrow terminal | the pointer at `MinWidth`–120 cols | `TestTheDoctorPointerIsShedRatherThanTearingTheFrame` | Shed whole, never wrapped and never past the frame — council's `collapsedNotice` rule, for its reason. | | first frame | `firstFrameText` | `TestTheFirstFrameIsShortAndNamesTheModeThatMeasures` | Under 30 lines, and it names every mode it sends a reader to. | | doctor next step | no seat ready, and one ready | `TestTheReportSaysWhatToRunNext` | Both branches name a command that runs on the machine just described. | ### 7.8 Keyboard Minimal, and every key earns its place. | Key | Action | |---|---| | `q`, `ctrl+c` | quit | | `esc` | close the detail pane → close help → clear the find query → **then** quit | | `↑`/`↓`, `j`/`k` | move the selection (scrolls the help overlay while it is open) | | `enter` | open the detail pane for the selected session; close it if open | | `/` | find: type-to-filter on name or path | | `v` | vendor filter cycle: all → claude → codex → gemini → agy → cursor → all | | `s` | sort cycle: activity → context → cost → activity | | `a` | toggle show-all (default hides sessions idle > 8 h) | | `r` | rescan now | | `?` | toggle help | In **find mode** the keyboard belongs to the query: only `esc` (clear and leave), `enter` (keep and leave), `backspace` and `ctrl+c` are commands, and everything else is text. That is why the mode takes over the whole footer — a mode that silently changes what `q` means without saying so is how a read-only monitor surprises someone. `--vendor all|claude|codex|gemini|agy|cursor` sets the starting filter; the cycle takes over from there. `antigravity` is accepted as a synonym for `agy` and `composer` for `cursor`; the short forms are the ids the footer and the header counts print. Cycles, not multi-select menus: with six vendors and three sorts, a cycle is one keystroke and no mode. Non-default filter, sort or query is always visible in the footer. The help overlay writes the cycle with `>` rather than `->` for one reason worth recording: the fourth vendor pushed that line past the 60-column floor, and a golden test at that width is what caught it. The **sixth** vendor exhausted that trick, and the cycle now wraps onto a continuation line indented under the first hop — shortening the vendor names instead would have made the overlay teach a name the footer does not print. > **Reversed in v1.1, deliberately.** v1 said: *"There is no selection cursor — the > default sort puts the interesting sessions on top, and a cursor invites drill-down, > which is a different product."* The roadmap (§8) then decided drill-down **is** the > product: the schema already carried `Diagnostics`, `Degraded` and every `Extra` with no > surface to show them on, and that machinery is the thing this project is actually > about. The original objection is answered rather than ignored — the cursor starts at > **no selection** and the mark appears the first time the user asks for it, so the > steady-state monitor frame is byte-identical to v1's. Anything that changes *which rows are visible or in what order* (`v`, `s`, `a`, a new query) **drops the selection** and closes the pane. The cursor is an index into the visible rows, so a different row set makes the old index point at a different session. Between polls the selection is carried by **session key**, not by index, because the activity sort re-orders rows as sessions write — holding the index would silently move the selection, and with the pane open would relabel one session's diagnostics with another's. Show-all deliberately does **not** hide a session with no activity timestamp: "we have no signal" is not evidence that a session is old. Deliberately absent: mouse support, fuzzy/regex/embedding search (the query is displayed literally, and a syntax that can mean something other than what it looks like is a filter that hides rows without saying so), and configuration UI. And one invariant that outranks all future feature requests: **the HUD is strictly read-only. No keybinding may ever mutate vendor state or send anything to a running agent.** The HUD is a telltale. That invariant is scoped to the observation surfaces — `hud` and `statusline` — and it does not weaken. `telltale council` (ADR-008, §9) is a separate subcommand that *does* dispatch to vendor CLIs; it is a dispatch room, not a gauge, it is entered deliberately, and it says so on screen. Nothing in §7 may reach for it. ### 7.9 Golden tests - `Render` is pure over `State` — `(sessions, vendors, now, width, height, filter, sort, showAll, help, scroll, scanning, thresholds)`. Tests construct the state directly and compare against `internal/hud/testdata/golden/*.txt` — no terminal, no program loop. `go test ./internal/hud -update` regenerates them. - Two families. **Layout goldens** render with `PlainStyles()`, a style set in which every `Render` is the identity, at widths 120 / 80 / 72 / 52 and at the height floor — so they never depend on the CI terminal's colour profile. **Style assertions** render with `NewStyles(true)` and check one escape code per severity band, mirroring `TestThresholdColors` in `internal/statusline/render_test.go`. - Width is measured with `lipgloss.Width`, never `len()` — the label column carries arbitrary project names. `TestNoLineExceedsTheTerminalWidth` sweeps seven widths with and without the help overlay. ### 7.10 Known limitations - Every glyph in the visual language — `● ◐ ○ ─ │ █ ▏▎▍▌▋▊▉ … — ↻ ⚠ ▸ ⑂ ·` — is East-Asian-**Ambiguous** width. Windows Terminal, the reference environment, renders ambiguous as narrow, which is what the grid assumes. A terminal configured to render ambiguous glyphs double-width will shear the layout; `--ascii` is the escape hatch. Stated here rather than discovered later. - `⑂` (U+2482 OCR FORK) is the least-common glyph in the set and the most likely to miss from a font. It appears **only** on a session that is fanning out, so a font gap costs a tofu box on a minority of rows rather than a broken grid — and `--ascii` renders it `Y`. It was chosen over a second `│` or a bracket because both already mean something here (`│` separates zones; `]` is the ASCII selection mark). - The detail pane does not scroll. A pane taller than the row area is clipped, and on a terminal shorter than about 16 rows a long extras list can run off the bottom. The arrows are spent on moving between sessions, which is the more valuable binding while the pane is open; a scrollable pane needs a second axis and is deferred. - The `█` fill and `─` track differ in glyph height by design. Verified legible in Cascadia Mono; other fonts may render the step more harshly. - Fill resolution is one eighth of a cell (1.04% at 12 cells). The number beside the bar carries the precision; the bar carries the glance. - ~~The account quota block is sourced from one session (§7.1). A second quota-bearing vendor needs a per-vendor block.~~ **Closed in two steps.** §7.15 (2026-08-07) gave the *header* a block per vendor, from the statusline relay alongside the transcript reading. §7.17 (2026-08-09) gave the per-vendor block its own surface: `u` opens a body with one block per vendor, the gauge at 20 cells instead of the header's 8, and — the part the header has no room for at all — a stated reason wherever a vendor has nothing to say. What remains is not this limitation but §7.17's own: an aged-out relay reading and one that never arrived render alike. - ~~The 1 s poll has not been measured on a cold cache over an 837-session tree (§6 Q3).~~ **Measured** — see §6 Q3 and the `BenchmarkScan` table. The cold scan is the half that did NOT improve, and that is the ruling rather than the residue: the first frame has to read everything, and the spinner is what covers it. - The burn forecast's sampling history lives in the process and dies with it. Restarting the HUD restarts the basis at zero, and for the first five minutes of every run there is no forecast at all. Persisting samples would mean writing to disk, which "telltale never writes" forbids, so this limitation is load-bearing rather than an oversight. ### 7.11 The detail pane **The problem it solves.** v1 carried `Diagnostics`, the `Degraded` field set and every `Extra` from adapter to renderer and displayed **none of them**. The grid can only draw one kind of nothing: a dropped column and an em dash both read as "no value here", and §4a.1 insists there are two different facts underneath. The pane is where the difference gets said in words. It is the honesty machinery becoming product rather than plumbing. `enter` opens it on the selected row; `enter` or `esc` closes it. It **replaces** the row area rather than floating over it, for the same reason the help overlay does — a panel covering the thing being monitored is a monitor you have to move to read. **Layout.** Line one is literally the selected row's identity zone (dot, vendor, `│`, label), so the pane opens where the row was. Everything below hangs off the `SESSION` column at offset 8: a 12-column muted label, two spaces, then the value. Field order mirrors the row's three zones — identity, measurement, time — then extras, then the honesty block, separated by one blank line because it is a different kind of statement. ``` telltale │ 4 sessions │ claude 3 codex 1 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── ● CC │ telltale ⑂~2 C:\src\code session 00000000-aaaa-4bbb-8ccc-000000000001 workspace C:\src\code\telltale model Opus 5 subagents ~2 recent activity live · 12s ago branch main cli 2.1.219 ctx tokens 215k degraded — diagnostics — not sourced context_pct, cost, quota ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── esc close ↑/↓ session ``` Read the last three lines together, because they are the whole point: - **`not sourced`** is "can't know" — the fields this vendor declared `CapNone`. They get no line of their own at all, exactly as the grid drops a column no visible row can fill. This line is the answer to "why is this row's CONTEXT cell empty?", and it is the first surface in the product that answers it. - **`degraded`** is "we tried and failed", named field by field. §4a.2 requires degraded and plain-absent to render identically in the grid — otherwise "we failed to read it" starts to look like data — and this is the one place that difference is legible. - **`diagnostics`** is why. One line per note, structure only, never transcript content. A clean session prints `—` on both rather than going blank: a blank honesty block is indistinguishable from a pane that forgot to render one. Degraded and absent under real failure: ``` telltale │ 4 sessions │ claude 3 codex 1 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── ◐ CX │ 4f2a9c81-1d3e-4a77-9b02-000000000000 session 4f2a9c81-1d3e-4a77-9b02-000000000000 workspace — model — context — quota — activity idle · 7m ago degraded workspace, context_pct diagnostics 2 unparseable records skipped no turn_context record in the read window not sourced name, cost, subagents ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── esc close ↑/↓ session ``` Every `—` above is a field Codex **can** source and has no value for right now; every name on `not sourced` is one it never could. Same glyph in the grid, two different facts, and this is where they separate. **`activity` is the one line that never renders `—`.** It reports the liveness *class*, which the HUD can always produce, with the age only as the evidence behind it. A session with no timestamp reads `unknown`, not `—`: the em dash would say "no value" where the truthful statement is "no basis for a claim" (§4a.4). **Selection.** `▸` in the row's leading pad column — the column that was already blank, so selection costs the grid no width. A glyph rather than reverse video, because §7.1 rule 2 says every distinction is carried by a glyph or a number first and a highlight-only cursor disappears under `NO_COLOR`. The mark is jammed against the state dot (`▸●`) on purpose: it reads as a pointer at the row's state, and the alternative is a dedicated column on every row forever to serve a mark that is off most of the time. ``` telltale │ 4 sessions │ claude 3 codex 1 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL AGE ▸● CC │ telltale ⑂~2 C:\src\code Opus 5 │ 12s ● CC │ acme-api C:\src\work Sonnet 4.5 │ 48s ◐ CX │ 4f2a9c81-1d3e-4a77-9b02-000000000000 │ 7m ○ CC │ learning-notes ⑂~5 C:\src\code Haiku 4.5 │ 22m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ``` That frame is also §7.13's: row 2 measured **zero** sub-agents and therefore draws no chip, and the CONTEXT and COST columns are auto-hidden because these are real Claude rows. ### 7.12 The burn-rate forecast **What makes this ours.** The incumbents in this lane project a burn line against a plan budget nobody publishes. That is the exact fabrication decisions/001 exists to forbid, and §8's "deliberately rejected" list names it. telltale instead samples the vendor's own `used_percentage` **over its own runtime**, reports the slope it measured, marks it derived, and states the sampling window beside it. The number is telltale's measurement of telltale's own observations, which is the one kind of computed figure this product is entitled to show. Rendered in the header beside the window it describes, never per row (the §7.1 corollary: quota is a property of the account): ``` telltale │ 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m ~13:27 · 18m basis 7d █▎────── 18% ↻ 5d02h ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL CONTEXT COST AGE ● CC │ telltale C:\src\code Opus 5 █████████▎── 84.2% $2.41 │ 12s ● CC │ acme-api C:\src\work Sonnet 4.5 ████▌─────── 41% $0.18 │ 48s ◐ CX │ notes-api C:\src\code gpt-5.1-codex — — │ 4m ○ CC │ learning-notes C:\src\code Haiku 4.5 ██████████▏─ 92.6% $11.07 │ 22m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ``` Both windows in that frame have the same 7 samples over the same 18 minutes. The 5h window is moving fast enough to project; the 7d window renders **nothing at all** — not a dash, not a placeholder, the cell simply ends. That contrast is the feature. **The four refusals.** A forecast renders only when all of these hold, and each one exists to prevent a specific lie: | Condition | The lie it prevents | |---|---| | ≥ **3 samples** spanning ≥ **5 minutes** | Two samples fit a line through themselves and cannot disagree; the third is the first one that can. Five minutes because a stepped percentage sampled over ninety seconds measures the step, not the rate. | | **positive slope** | "You will never run out" is not a time. A flat or falling window renders nothing rather than an infinity. | | exhaustion **before the window resets**, when `resets_at` is known | Projecting past the reset describes a window that will not exist. | | exhaustion **within 24 h** | The render is a wall clock with no date on it. `~04:12` sixteen hours out is misleading, not informative. | **The arithmetic**, pinned by `TestForecastArithmeticIsPinned`: least-squares slope over the retained samples, projected from the **last observed value**. Least squares rather than a first-to-last difference because vendor usage percentages move in steps and a two-point slope is dominated by whichever endpoints straddle a step. Anchored to the last observed reading rather than to the fitted line so the projection starts from the number printed next to it — a forecast that quietly starts from 44% while the cell says 42% is a small lie in the place this product is least allowed one. **Sampling.** One sample per *completed* scan (a failed scan contributes nothing rather than a repeat of the last reading, which would flatten the slope with data we did not measure), throttled to one every 15 s, bounded to 30 minutes and 128 entries. A window with a nil `UsedPercent` this scan is a **gap, not a reset** — the history stands. **Rollover clears the buffer**, on either of two signals: usage dropping (monotonic within a window, so a drop is a rollover), or `resets_at` jumping forward by more than a minute (a rollover moves it a whole window; jitter does not). Fitting a line across a rollover reports a negative rate or a wild one, and every sample before it describes a window that no longer exists. **Amendment to §7.1 rule 4** ("still by default"), stated rather than quietly taken: the forecast cell may change when a new sample lands, which is at most once every 15 s and only when the measurement itself moved. That is a measurement changing, not an animation — the §7.6 rule is about telltale never animating the *vendor's* state, and this is telltale reporting its own arithmetic on a new reading. ### 7.13 The sub-agent chip `⑂~2` after the session label on any row whose adapter counted recently-written transcripts in that session's `subagents/` sidecar. Sourced by a stat pass (§3.1), Claude only. **Why the `~`.** The count is exact — telltale listed the directory. What is *inferred* is the 15-minute recency boundary that turns "written lately" into "a fan-out is running now", and ADR-001 requires the inferred part be visible. So the chip carries the same estimate marker the CONTEXT column does, and it means the same thing: this number was computed by telltale, not reported by the vendor. **Zero draws nothing.** The absence of a chip is not a claim, and a `⑂0` on every Claude row would be noise asserting a fact nobody asked for — the same reasoning as an absent gauge drawing no track. The measured zero is not discarded, though: the detail pane says `~0 recent`, where there is room to distinguish "we counted none" from "we could not count". A sidecar the OS refuses renders `—` there, never `0`. Styling: the chip renders in `Text`, not `Muted`. `Muted` is this palette's "chrome or absent" (§7.5, `Absent() = Muted`), and rendering real measured data in it would put a sourced number in the same visual class as a missing one. ### 7.14 Type-to-filter `/` opens the query; typing narrows rows by case-insensitive substring; `enter` keeps the query and hands the keyboard back; `esc` clears it. ``` telltale │ 2 of 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL CONTEXT COST AGE ● CC │ acme-api C:\src\work Sonnet 4.5 ████▌─────── 41% $0.18 │ 48s ◐ CX │ notes-api C:\src\code gpt-5.1-codex — — │ 4m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── /api_ esc clear enter apply ``` and once applied, with the mode left: ``` telltale │ 2 of 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL CONTEXT COST AGE ● CC │ acme-api C:\src\work Sonnet 4.5 ████▌─────── 41% $0.18 │ 48s ◐ CX │ notes-api C:\src\code gpt-5.1-codex — — │ 4m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys find "api" ``` Four rules, all of them the same rule — **a monitor that hides rows must say so**: 1. The header count reads `2 of 4`, so the headline can never contradict the per-vendor totals beside it. (Same mechanism as the vendor filter; they compose.) 2. An applied query keeps announcing itself in the footer after the mode is gone. A filter the user has forgotten about hides rows just as silently as one they cannot see. 3. If nothing matches, the empty state says `no sessions matching "zzz"` rather than `no active sessions` — naming the thing that emptied the list. 4. The match is a literal substring, displayed literally. No globs, no regex, no fuzzy or embedding search: a syntax that can silently mean something other than what it looks like is a filter that hides rows without saying so. **What it matches:** the vendor's session name, the workspace path, and the session id. The id is in the set because a torn-record row is *labelled* by its id (§7.3 render D row 3), and matching only the "title" would make the one piece of text on that line unable to find it. `/` is a mode, and it is the product's only one — which is why it takes over the whole footer instead of quietly changing what an unmodified key does. ### 7.15 The quota relay — every vendor the header can honestly speak for Added 2026-08-07, San's ruling. The header's quota block was one vendor (Codex) because Codex is the only vendor whose quota exists on disk where a passive reader can see it: Claude's `rate_limits` arrive **only** on its statusline stdin payload (§3.1 — the live corpus was grepped, nothing quota-shaped reaches the transcripts), and agy's named buckets exist only in its statusline payload the same way (§3.8). Cursor's store holds plan-entitlement constants that must never render as usage (§3.9), and Gemini has nothing — so those two vendors have no quota **anywhere**, relay or not. **The mechanism: the statusline relays what it just rendered.** After the line is on stdout, `telltale statusline` writes the payload's quota windows to `~/.telltale/quota/.json` (`internal/quotacache`), and the HUD's scan reads every surviving entry alongside the vendor stores. This is a deliberate, scoped amendment to "the gauges never write" (§1, CLAUDE.md): - **numbers only, never content** — vendor id, timestamp, window ids/labels, percentages, reset instants. The same keys-not-content standard as council's `room.json`, pinned by a test that walks the serialized form field by field. - **atomic and best-effort** — temp + rename in the same directory, error ignored after the render is delivered; the cache can never cost a statusline frame or a torn read. - **self-expiring** — the reader drops a window whose reset has passed (its percentage is not stale, it is *false*), and whole entries past 24h or stamped from the future beyond clock-jitter tolerance. - **age travels with the reading** — past 5 minutes a relayed block carries `· 2h ago` at every dress level, the §7.12 basis rule applied to time: shedding the age would re-present a stale number as fresh. Past `quotaAgeWarn` it stops being muted chrome and escalates to `· ⚠ stale 19h ago`; §7.17 as amended argues the threshold and owns both surfaces' wording. **One block per vendor, transcript outranks relay.** A vendor sourced from its own store (Codex) is re-measured every scan; its relay entry, if one ever exists, is as old as the last statusline render. The scan-fresh reading wins and the vendor renders once. Only the transcript-sourced block may carry a burn forecast — window ids collide across vendors (Claude and Codex both have a `seven_day`), and re-reading an unchanged cache file is not a new observation, so a forecast on a relayed block would be one vendor's slope pinned to another's account. **The line fits by shedding decoration, never fact.** Dress levels, tried in order until one fits: full (names, gauges, countdowns, forecasts) → drop forecasts → names to two-letter tags → drop gauges (the percentage beside each bar says the same thing) → drop countdowns. Vendor, window label, reading, and a stale reading's age survive every level. If even the barest level overflows, whole trailing blocks are dropped and an ellipsis says so — the footer's dropping-is-never-silent rule. The generated render (`quota-fleet` golden): ``` telltale │ 1 session │ codex 1 ag gemini-weekly 38% ↻ 3h00m │ cc 5h 42% ↻ 2h13m 7d 6% ↻ 5d00h · 2h ago │ cx 7d 79% ↻ 22h48m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL AGE ◐ CX │ notes-api C:\src\code gpt-5.1-codex │ 4m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ``` At 120 columns three vendors and four windows already shed to tags-without-gauges; the full dress needs ~145. That is the honest trade as measured, not a bug: the bars went first precisely so every fact could stay. What the relay does **not** change: the statusline's own display (it renders from stdin as before, the write happens after), the HUD's read-only posture toward *vendor* files, and the absence rule — a vendor whose statusline never fires simply never appears, and one that stops firing ages out. ### 7.16 The token relay — what Cursor cost, from a seam with no network call Added 2026-08-08. §3.9 declared Cursor's `cost` and `quota` **ABSENT**, and re-verifying it that day made the verdict harder rather than softer: `usageData` was `{}` in 19 of 19 blobs, `tokenCount` was zero in 1,622 of 1,622 message rows, 78 `turn_ended` records and 51 transcripts carried status and no numbers, and the only account figures anywhere on disk were Statsig experiment values stamped `is_user_in_experiment:false`. Nothing about consumption reaches the store as a byproduct of a turn. So the number was fetched rather than found: **Cursor Hooks**, the vendor's own documented and versioned contract (cursor.com/docs/hooks), whose `afterAgentResponse` step hands a command hook the turn's token counts on stdin. **Why the hook and not print mode — the derived-`inputTokens` trap.** `cursor-agent -p --output-format json` prints a `usage` block, and reaching for it would have been the obvious move. It is the wrong one: its `inputTokens` is **not** the raw count. The CLI publishes `max(raw − cacheRead − cacheWrite, 0)`, measured printing **24,076 where the un-derived input was 48,012**. Rendering that under the label "input tokens" would be telltale repeating a vendor's arithmetic as if it were a reading — the ADR-001 violation this project exists to refuse, and a *quieter* one than usual, because the number looks perfectly plausible. The hook payload carries the vendor's own `tokenUsage` fields untouched, which is the entire reason it wins. Source-read at **cursor-agent 2026.08.04-aaa8809** (`8674.index.js`, `./src/after-agent-hooks.ts`), where the payload is assembled as `{conversation_id, generation_id, model, text, input_tokens: tokenUsage?.inputTokens, output_tokens: …, cache_read_tokens: …, cache_write_tokens: …}` and then enriched by the executor (`190.index.js`) with `hook_event_name`, `cursor_version`, `workspace_roots`, `session_id`, `transcript_path` and **`user_email`** before it reaches stdin. Both Windows transports (`argv_heredoc`, the default, and `windows_temp_file`) deliver that JSON on the command's **stdin**, under PowerShell. **The payload is the reason the allowlist is a struct.** This is the first telltale seam where the numbers arrive in the same object as the model's full reply *and* the user's email address. `internal/cursorhook` decodes into a four-field struct of integer pointers; `encoding/json` discards everything with no destination, so no content field can reach the cache unless someone adds a field on purpose — the technique `internal/adapter/cursor` already uses against a store that keeps OAuth tokens beside session state (decisions/007), pointed at a payload that keeps PII beside numbers. A test plants markers in every content-bearing field of a real payload shape and asserts none of them survives, at the parser AND again on the serialized cache file. **Three fields were left out on purpose, and they are not content.** `model` and `generation_id` are per-turn facts and the entry is a TOTAL — naming one turn's model beside a sum invites reading the sum as that model's. `conversation_id` names a cursor-agent **CLI** conversation, and the HUD's Cursor rows come from the **IDE's** Composer store (§3.9); the CLI keeps a separate one. Storing it would dangle a join that does not exist, which is also why this reading is not rendered on a session row. **Amended 2026-08-29: the `conversation_id` ruling HOLDS and its reason no longer does.** The HUD draws CLI rows now, out of `~/.cursor/chats///meta.json` (§3.9's 2026-08-29 addendum), so the join has something to join to for the first time. It is still not built and the field is still not stored, because whether this `conversation_id` IS that directory's session uuid was never measured. A key stored on the assumption that two ids match is how a relay begins attributing one session's tokens to another, and §7.16's whole argument is that the tokens must not be attributed to a row on a guess. Measure the two ids against each other first; nothing about the held display changes either way. #### The accumulation ruling: a total, and never without its window A hook fires once per agent response, so the file is either the last turn's numbers or a running total. It is a **running total**, and the price of that choice is that the window is not optional: - a single turn's counts answer a question nobody asks — the turn you just watched finish — and go stale the instant the next one starts. A *counter* is the thing a token figure wants to be. - but a sum over an unbounded window is a different and much weaker claim than a reading. So the entry carries `since` and `turns`, both travel to the screen, and the renderer may never print the sum without them. "48k" is a number pretending to be a state; "in 48k · out 1.2k · 14 turns over 12m" is a measurement with its scope attached, the same §7.12 basis rule the burn forecast and the relayed quota block already follow. - the window's boundaries are mechanical rather than chosen. Accumulation continues onto any entry a *reader* would still accept, and opens a fresh window otherwise — first turn ever, first turn after a day of silence, first turn after a corrupted or clock-skewed file. `internal/usagecache.readEntry` is shared by `Add` and `ReadAll` precisely so those two can never disagree; without that, a sum could silently span a week-long gap and still call itself a total. - **a partial turn is refused, not part-counted.** A payload missing any of the four counts is not accumulated at all. Summing the three that arrived and treating the fourth as zero would leave the total wrong by an amount nothing on screen could name, while it kept looking like a total; refusing makes the counter go quiet, and a visible absence is the failure mode §7.7 prefers every time. Every count in the file is therefore a sum of complete readings, and `turns` says exactly how many. Everything else is §7.15's mechanism copied deliberately, function for function: one file per vendor under `~/.telltale/usage/`, atomic temp+rename in the same directory, best-effort, self-expiring at 24h and on future-skew, and the reading's age travelling with it past five minutes. `internal/usagecache` is a **sibling package** rather than a second store inside `internal/quotacache` because the two share their mechanism exactly and their schema not at all — quota is windows with percentages, resets and a "reset has passed, so this window no longer exists" rule that means nothing to a counter — and folding them together would put one keys-not-content test in charge of two unrelated formats. #### What it rendered, and what stayed absent *This subsection describes the display as built on 2026-08-08. It was retired on 2026-08-09 — see the amendment below — and is kept in the past tense because the rules it worked out still bind the one spend line that remains (§7.17).* **Tokens spent are not quota, and the render may never blur that.** There is no denominator anywhere in this reading, so no percentage, no gauge, no countdown, no bar — any of them would invent a ceiling out of nothing, the same class of error as filling a `CapNone` field with a plausible guess. The spend block therefore got **its own header line, never shared with quota at any width**, and carried a verb: `cursor spent`. The verb is a word, not a glyph or a colour, so `--ascii` and `NO_COLOR` lose none of the claim. It rendered as: ``` telltale │ 2 sessions │ codex 1 cursor 1 codex 7d █████▌── 79% ↻ 22h48m cursor spent in 48k · out 1.2k · cache read 1.9M · cache write 62k · 14 turns over 10m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL CONTEXT AGE ``` Codex's quota was on the quota line; Cursor's was nowhere, and Cursor's spend was on a line of its own. The shed cascade followed §7.15's grammar — cache pair first, then the turn count, then full names down to the two-letter tags — and vendor, verb, `in`, `out` and the window survived every level, with the ellipsis saying so whenever anything went. `theme.Tokens` floors at every step for the same reason `theme.Percent` does: this is what a machine *spent*, and rounding 47,950 up to "48.0k" invents fifty tokens nobody was billed for. #### The display is retired; the relay is not (owner's ruling, 2026-08-09) **Ruled by the owner: the Cursor spend line comes off every surface. The seam, the hook, the cache and the HUD's read of it all stay.** The reason was not that the number was wrong — it was measured end to end and it was right. It was that *it bought no decision*. A running token count for a vendor with no ceiling anywhere answers a question nobody was asking, and it was answering it from a header line the header does not have to spare: §7.15's whole design is a shed cascade fighting for one or two rows, and this was permanently occupying a third. That is a product judgement, not an honesty one, and it is worth naming which because the two have different consequences. Nothing here was retracted. §7.16's measurements, its vocabulary rules and its accumulation ruling all still stand and all still bind — the fleet usage view's remaining spend line is held to them (§7.17), and the amendment below is the first place they were applied to a sum of a different shape. What changed, exactly: - **removed:** the header's spend line, and the usage view's Cursor spend row. Nothing renders a `usagecache.Total` anywhere. - **kept, deliberately untouched:** `telltale hook cursor`, `internal/cursorhook`, `internal/usagecache` and every test either owns; `~/.cursor/hooks.json` on this machine; and the HUD's own read — `Snapshot.Spend` is still filled by every scan and is read by nothing. `internal/hud/state.go` says so on the field, because a reader who finds an unused field will otherwise correctly conclude it is dead and delete it. Reinstating the display is a call site, not a re-plumb, and the accumulating file means the day it comes back it has history in it rather than starting from this minute. - **pinned:** `TestTheCursorSpendDisplayIsRetiredEverywhere` renders the same fixture that produced the old block — relayed total still in the snapshot — at five widths, in both glyph sets, with the usage view open and closed, and fails on the verb or on either of the counts appearing anywhere in the frame. `TestTheRetiredDisplayStillHasItsRelayUnderneath` is the other half, and it is the one that catches "retired" being implemented as "deleted". `TestCursorIsStillGivenNoQuotaBlock` survives from the old pair: losing its spend line is exactly the moment a renderer would be tempted to find Cursor a home on the quota one. The same fixture now renders (`cursor-without-spend` golden) — a two-line header where there were three, Cursor's row still on the grid, and its quota still visibly nowhere: ``` telltale │ 2 sessions │ codex 1 cursor 1 codex 7d █████▌── 79% ↻ 22h48m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL CONTEXT AGE ● CU │ Multi-vendor orchestration C:\src\code composer-2.5 ████▏─────── 37% │ 1m ◐ CX │ notes-api C:\src\code gpt-5.1-codex — │ 4m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ``` #### The boundary amendment This is the **third** bounded write on a gauge path, and §1 and `CLAUDE.md` name it alongside council's `room.json` and the statusline's quota relay. It meets the same bar: under `~/.telltale/`, numbers and keys only, atomic, best-effort, and pinned by a test that walks the serialized form field by field. It is also the first one written by neither gauge — `telltale hook cursor` is its own mode, because a hook's stdout is parsed by the vendor as a hook result, so this is the one path in the binary where printing nothing is the contract and every exit is clean. #### Live verification, 2026-08-08 — and the half of it that did not hold Measured, on Windows 11 at cursor-agent 2026.08.04-aaa8809: - the relay path end to end, driven by the real binary from a real turn's real counts (`input 23,941 / output 33 / cache 0 / 0`, from a print-mode `result` whose zero cache figures make its derived `inputTokens` equal to the raw one). Two turns relayed; `~/.telltale/usage/cursor.json` read back `{"vendor":"cursor","turns":2,"input_tokens":47882,"output_tokens":66,…}` — accumulated, with no `text`, no `user_email`, no `conversation_id` and no `model` anywhere in it — and the HUD rendered `cursor spent in 47.8k · out 66 · cache read 0 · cache write 0 · 2 turns over 6s` from that file. - **the vendor never invoked the hook, on any surface reachable from a script.** Marker hooks were installed at `~/.cursor/hooks.json` on both `beforeSubmitPrompt` and `afterAgentResponse` and neither fired for: `-p --output-format json`, a non-TTY run without `-p`, or an **ACP** turn driven over JSON-RPC. The ACP result is corroborated by source rather than resting on one capture — the ACP chunk (`8096.index.js`) contains **zero** references to `hookExecutor` — and the only call site of the `afterAgentResponse` helper sits in the React/ink agent-app module, beside the protobuf agent-server dispatcher the IDE talks to. The config itself was not the problem: `type` defaults to `"command"` in the bundle's own validator, and an invalid config produced no `[hooks]` warning on those paths either, which is consistent with the subsystem not being initialized there at all. So the seam is real and its payload is verified by source read at a pinned version, and the **invocation is verified for no surface yet**. What is unverified, itemized: that a true-TTY interactive `cursor-agent` session fires the hook; that the Cursor IDE's agent server does; and therefore that any figure ever appears without being fed in by hand. The hook is left wired at `~/.cursor/hooks.json` so the first session that does fire it captures a total. Worth stating plainly because it bears on the roadmap: **council's Cursor seat runs on ACP (§9.36), and ACP does not carry hooks** — so the seat that most wanted per-turn cost is, at this version, the one surface that provably cannot supply it. **AMENDED 2026-08-15: the last sentence above is too wide, and the ACP path DOES carry hooks.** The paragraph is kept whole because its two measurements still stand — neither `afterAgentResponse` nor `beforeSubmitPrompt` fired, and the ACP chunk really does hold no `hookExecutor` reference. What is wrong is the generalization from those two events to the subsystem. **The hook subsystem is live on the ACP path, per event.** Measured at the SAME build the paragraph above was measured at, `cursor-agent` **2026.08.04-aaa8809**, on Windows 11, driven from a PowerShell parent (`PARITY.md`'s launch-parent trap), with the handshake `internal/council/vendors/cursoracp.go` builds: `initialize`, `session/new`, `session/prompt`, and any `session/request_permission` answered `allow-once`. The observable is the filesystem. | arm | `~/.cursor/hooks.json` | marker created | trials | |---|---|---|---| | baseline | telltale's `afterAgentResponse` only, unchanged | **yes** | 1/1 | | deny | plus one `beforeShellExecution` returning `permission: "deny"` | **no** | 2/2 | **The hook's own breadcrumb is the proof, and the vendor stamps its build into it.** The payload arrived carrying `"hook_event_name":"beforeShellExecution"`, `"command":"mkdir cursor-deny-marker"` and `"cursor_version":"2026.08.04-aaa8809"`. So a `~/.cursor/hooks.json` hook runs on a turn driven entirely over the ACP wire, and its denial holds: the directory was never created, on either trial. **Three records disagreed and each was right about its own event.** §7.16 measured `afterAgentResponse` and `beforeSubmitPrompt` and found neither firing. `cursoracp.go` and `PARITY.md` recorded a hook blocking an ACP tool call, which is `beforeShellExecution`. The research read that found a hook executor in chunk 4414 was reading a chunk the ACP path reaches, not the ACP chunk itself: at this build `8096.index.js` holds **zero** `hookExecutor` references and `4414.index.js` holds one, while `beforeShellExecution` appears in `190.index.js`, `3384.index.js` and `index.js`. A per-event answer reconciles all three without withdrawing any of them. **The token relay's own conclusion is unchanged, and it was re-checked rather than assumed.** Telltale's `afterAgentResponse` entry stayed wired through all three ACP turns above. `~/.telltale/usage/cursor.json` still reads `"turns":2` with a `written_at` of 2026-08-08, so that event still does not fire on ACP. The seat that most wanted per-turn cost still cannot supply it; what changes is only the reason, which is the event rather than the protocol. **A trap for the next hook written against this vendor.** The payload cursor writes to the hook's stdin begins with a UTF-8 **BOM**, so a plain JSON parse of it fails. `telltale hook cursor` is not exposed to this today, because it is the one event that does not fire here. #### Known limitations - **Accumulation is a read-modify-write and takes no lock.** Two hook processes finishing in the same instant can lose one turn. Accepted rather than locked: the loss is bounded and self-consistent (the `turns` count drops with the counts it names, so the total never disagrees with its own window), and the alternative is a lock file on a path a vendor's turn is waiting on — a gauge that can hang a turn is strictly worse than one that undercounts. - **`~/.cursor/hooks.json` is un-versioned machine state.** It names an absolute path to the binary, so it does not travel; versioning it belongs in the dotfiles repo, not here. - The counter says nothing about *which* conversation or model spent the tokens, by the design ruling above. If a future seam makes a CLI conversation joinable to a HUD row, that is a new claim and needs its own section. #### The statusline seam, measured per surface (2026-08-16) A third Cursor seam, measured on the same build the two above were measured on — **cursor-agent 2026.08.04-aaa8809**, Windows 11. `~/.cursor/cli-config.json` accepts a top-level `statusLine` object using the stdin-JSON/stdout-text contract telltale already implements twice. §2.2 is the design record and the renderer shipped; this is the measurement. **The §7.16 lesson was applied before the fact.** "A config that validates is not a subsystem that initializes" is what the deny-composition block above cost to learn, so no surface was reasoned about — each was driven with a marker command that tees its stdin to a file and exits 0. | surface | how it was driven | marker fired | |---|---|---| | interactive TUI, real console | hidden-window `cmd /c cursor-agent.cmd`, known workspace | **yes** (2 invocations inside 0.4s) | | ACP | `cursor-agent acp`, `initialize` + `session/new` — the handshake `cursoracp.go` builds | **no** | | print mode | `cursor-agent -p`, one real turn | **no** | | interactive argv, piped stdio | no TTY | **no** — the TUI never mounts at all | | interactive, workspace never opened before | same as row 1, in a fresh git worktree | **no**, 2 trials totalling 55s | The source read agrees and explains the shape of it: the invocation is a React hook (`./src/hooks/use-status-line.ts`) called from the ink chat component in `./src/ui.tsx`, and `statusLine` appears in exactly one bundle chunk (`8674.index.js`). No headless surface mounts that component. The last row is recorded as measured and its mechanism was not chased — the fresh worktree had no `~/.cursor/projects/` entry and the workspace that fired did. **So this seam repeats §7.16's conclusion rather than relieving it: council's Cursor seat runs on ACP (§9.36), and this seam is interactive-only.** Two seams now, both real, both invisible to the one surface that wanted them. The reason is different each time — an unfired event there, an unmounted component here — and neither generalizes to the other, which is the mistake the 2026-08-15 amendment above exists to correct. **No vendor marker, so routing is a flag.** Every captured payload was checked. There is no `product`, no `hook_event_name`, and no other vendor name; `version` carries the CLI build string, a value rather than a marker. The payload is Claude-shaped by the vendor's own statement of intent. §2.1's affirmative-marker scheme therefore cannot be extended, and `telltale statusline --vendor cursor` is the routing (§2.2 argues why a structural guess was refused). **The BOM is NOT on this seam — and the check is now in the code anyway.** The trap note above records that cursor's HOOK payload begins with a UTF-8 BOM, which breaks a plain JSON parse. The statusline payload does not: both captures began `7B 22 73` (`{"s`). `cursorstatus.Parse` strips one regardless, and `TestCursorBOMIsStripped` pins it. One vendor writing two seams with two encodings is the case where three bytes of caution are cheaper than rediscovering the difference. **The ACP `usage` field is parsed and deliberately goes nowhere.** The response schema in `8096.index.js` declares session/prompt's result as `{stopReason, usage?}`, with usage as `{inputTokens, outputTokens, totalTokens, cachedReadTokens?, cachedWriteTokens?, thoughtTokens?}`. The thirteen live arms of §9.36 saw no `usage` on any turn, and that finding stands — the schema says optional and this build never sent one. `acpTurnEnded` now parses it into `acpUsage` and **nothing reads it, relays it or renders it.** The reason to stop there is on the record twice already: print mode's `result.usage` published a CLI- computed `inputTokens` under that exact name, and the statusline seam does the same thing today with `total_input_tokens`. Whether ACP's `inputTokens` is raw or derived is an open question a live capture at a pinned version must close, and displaying it first would be the violation rather than the discovery of one. **Hand-back — the one measurement this session could not take.** Every captured payload came from a session that had made no API call, so `context_window` arrived with all six keys `null`. The populated shape is therefore documented and synthesized, never observed: `used_percentage` has been read from the bundle as vendor-sourced, but its live magnitude and precision have not been seen. Closing it needs a human at a real TTY, which is the one thing an agent cannot drive here. Roughly a minute: 1. Add to `~/.cursor/cli-config.json`: `"statusLine":{"type":"command","command":"telltale statusline --vendor cursor"}` 2. Run `cursor-agent` in a workspace it has opened before, and send any one prompt. 3. Read the line above the input. `ctx N%` appearing after the reply is the confirmation; the segment staying hidden means `used_percentage` is null for longer than assumed. **Paid, 2026-08-17, at cursor-agent `2026.08.11-e8db854`.** The operator ran the three steps above. The environment was an interactive session in Windows Terminal, from a PowerShell parent, in this repository's own workspace. After the first reply the line read: ``` Composer 2.5 Fast │ ctx 12.7% │ telltale ``` The segment drew, so `used_percentage` populates on a real session. Three things are now observed rather than assumed. The magnitude is live: `12.7` is a reading, not a null and not a zero. The precision is **one decimal**, which is the precision the synthesized fixture already assumed — so the fixture shape stands unchanged. The field populates **after the first reply**, which answers step 3's alternative outcome; nothing waited longer than one turn. **The build is newer than this section's own pin, and that is stated rather than smoothed.** Every other measurement in this section came from `2026.08.04-aaa8809`, including the bundle arithmetic that ruled two fields derived. This capture ran at `2026.08.11-e8db854`, the build §9.46 pinned. So read the populated shape as confirmed at the newer build, and the derived-field ruling as measured at the older one. The two do not substitute for each other. **One limit, named.** This capture confirms that the field carries a number. It does not re-measure `remaining_percentage` or `total_input_tokens`, because §2.2 keeps both out of the structs and the line renders neither. A capture cannot observe a field the parser deletes. ### 7.16a The OTLP collector — what grok spent, pushed rather than hooked (2026-08-10) §3.9a closed grok's quota question three times over and ended on a seam: the vendor's one designed-for-reporting surface is an external OpenTelemetry stream that carries **spend and no quota at all**. This section is that seam, spent. `telltale otel grok` is a loopback-only OTLP/HTTP listener; grok's own exporter pushes to it, and each `grok_code.api_request` event's four token counts are folded into `~/.telltale/usage/grok.json` — the same cache, accumulation ruling and refusal gates as §7.16, fed by a push instead of a hook. The measured export shapes, capture environment and grok version are pinned in §3.9a's export addendum; `internal/grokotel`'s package doc carries the same facts beside the code. **Why a listening socket does not breach §4a.5.** The adapter contract's "no network calls" protects two things: a gauge that can stall on a wire, and a gauge that can *reach out* — toward an endpoint, with credentials, spending from the pool it reads. The collector does neither, and the direction of the arrow is the whole argument: telltale opens a socket on 127.0.0.1 and **the push is grok's**, exactly as §3.9a recorded when it named the seam. The gauges still make no network calls and read no credentials; they read the FILE this mode writes, exactly as they read the hook relay's. It is its own mode for the same reason `telltale hook cursor` is (§7.16's boundary amendment): its I/O contract — a foreground server holding a port — belongs to neither gauge. The bind refuses any non-loopback address at startup, mechanically: a collector reachable off-box would be an open door wearing a gauge's name. **One source, chosen over a redundant second — and there turned out to be a third.** The rule below is scoped to the two envelopes on the wire, which is all that was known here in 2026-08-10. A 2026-08-29 re-measure found the same per-turn counts on grok's DISK as well (§3.9a's `usage` re-measure), so the rule now has a wider job; the amendment at the end of this section states it. The stream carries the same counts twice — per-request on `api_request` events, aggregated on the `token.usage` metric — and §3.9a's capture measured them value-for-value equal. The collector reads the EVENTS and acknowledges `/v1/metrics` without reading it: one record is one claim, an event carries all four counts atomically (so §7.16's complete-or-refused gate maps onto it unchanged), and reading both envelopes would be two chances to count one number. The 200 on the unread path matters — an unacknowledged export is retried, and making the exporter loop on a signal nobody reads would spend grok's batches on nothing. **The window unit is the api request, and the entry says so.** grok's counts arrive per API call, not per turn (`turn_completed` carries no counts — measured), so the cache entry's window count is `requests`, a new sibling of `turns` in the §7.16 schema. The same amendment gave the schema `reasoning_tokens` and made two fields *optional with their absence meaning something*: a cursor entry carries `cache_write_tokens` and no `reasoning_tokens`, a grok entry the reverse, because each vendor's file may only claim the counts its vendor keeps — §4a.1's zero-versus-absent rule, applied to the serialized form. `TestTheGrokShapedEntryCarriesItsOwnKeysOnly` and the cursor keys test pin both shapes field by field. **What the wire carries and what survives.** Every record arrives with `session.id`, `user.id`, `team.id`, `model` and timing beside the counts; with a content gate open it would carry prompt text. Four counts survive. The extraction is an allowlist the same way `internal/cursorhook`'s struct is — an attribute key with no case in the parser falls through unread — and `TestNothingFromTheWireReachesDisk` plants content markers on a real record shape (plus a gate-open `user_prompt` event) and asserts nothing but the numbers reaches the file. `session.id` and `event.sequence` are read into collector *memory* for one purpose: the exporter retries unacknowledged batches, and a total that counts a retried batch twice is overstated by an amount nothing can name. A replayed (session, sequence) pair is refused; the guard is never written to disk. **The display is held, and it is the owner's own ruling applied.** §7.16's amendment retired the cursor spend line because a running count for a vendor with no ceiling anywhere buys no decision — and grok is *more* ceiling-less than Cursor, not less: §3.9a swept its disk twice, probed the free network half and read the vendor's own monitoring schema, and no quota exists anywhere. §7.17's Declined already refused grok a spend line sourced from disk ticks. So this relay ships exactly as the cursor one now stands: write, cache and the HUD's read of it wired (`Snapshot.Spend` carries the entry; nothing renders it), display a call site away, and the accumulating file means the day a display is ever ruled in it has history rather than starting from that minute. `TestTheGrokSpendRelayRendersNowhere` pins the hold at every width, in both glyph sets, with the usage view open and closed. **Wiring it on a machine** (the enable is machine-local config, deliberately not in this repo): ```toml # ~/.grok/config.toml — grok's double opt-in, pointed at the default local endpoint [telemetry] otel_enabled = true otel_logs_exporter = "otlp" ``` then leave the collector running while grok runs: ``` telltale otel grok ``` It listens on 127.0.0.1:4318 (OTLP's http default, so the zero-flag pairing finds itself; `--addr` moves it, loopback only) and prints one line per counted request. The content gates (`otel_log_user_prompts`, `otel_log_tool_details`) stay off; the collector keeps nothing they would add, and the planted-marker test is the proof, but a content-free wire is strictly better than a filtered one. Verified end to end on 2026-08-10: a config-driven `grok -p "hi"` (grok 1.0.0 (3cd0d0cbce), no env overrides, default batch intervals) against the running collector produced `{"vendor":"grok","requests":1,"input_tokens":23767,"output_tokens":96, "cache_read_tokens":1408,"reasoning_tokens":81,…}` — real numbers, keys only. **Amended 2026-08-16 — the collision on 4318, measured and then named.** 4318 is OTLP/HTTP's registered port. That is why this mode defaults to it, and it is also why every other local OTLP receiver takes it: Jaeger, the OpenTelemetry Collector and the vendor agents all default there. So the most likely startup on a working machine is the one that cannot bind, and nobody had measured what that looked like. **Measured 2026-08-16**, Windows 11, `main` at `4e0cf6b`, with a throwaway listener holding 127.0.0.1:4318: `telltale otel grok` printed one line on stderr and exited 1. ``` telltale otel: listen tcp 127.0.0.1:4318: bind: Only one usage of each socket address (protocol/network address/port) is normally permitted. ``` So the failure was already loud and already correctly coded. Nothing pretended to collect, and nothing hung. What the line did not carry is what to do next, and there are three parts to that: the likely holder is another OTLP collector, `--addr` moves this side (the flag already existed and was already documented), and moving this side ALONE counts nothing, because grok's exporter goes on posting to 4318. A collector listening on a port nobody pushes to reads exactly like "grok spent nothing", which is the failure §7.7 rates worst. The message now states all three and keeps the bind error verbatim underneath it. On a port the operator chose it names no likely holder, because telltale cannot know who took 4444, and it says that instead of guessing. `TestAHeldPortSaysWhoLikelyHasItAndHowToMove` and `TestTheDefaultPortCollisionNamesTheOtherCollectors` pin both branches; `TestAMovedPortBindsAndCountsARequest` pins that the way out works, and `TestAMovedPortIsStillLoopbackOnly` keeps the loopback bind absolute across the flag. The redirect the message prescribes is `OTEL_EXPORTER_OTLP_ENDPOINT=http://` in grok's own environment, beside the `[telemetry]` pair above. **That redirect is NOT measured, and the message says so on the line that prescribes it.** The capture pinned the exporter as OTel-OTLP-Exporter-Rust/0.32.0 posting to the default endpoint with no variable set; nothing here re-ran an export with the variable moved. A named knob marked unverified beats no knob at all, and marking it is §4a.1's estimate rule applied to a sentence instead of a number. One Windows detail earns its line, because the portable-looking version of it is wrong. `errors.Is(err, syscall.EADDRINUSE)` is FALSE on Windows for a real collision: the bind returns errno 10048 (`WSAEADDRINUSE`), while Windows builds define `syscall.EADDRINUSE` as one of Go's synthetic `APPLICATION_ERROR` constants, 536870914 on this box (measured, go 1.26). Windows is the primary target (ADR-002), so a one-arm check would have detected the collision on the two platforms CI does not run and missed it on the one it does. The detection carries both arms and cites the measurement beside them. #### Known limitations - **The collector must be running to hear the push.** grok's exporter retries briefly and then drops a batch; spend accrued while the collector is down is not counted later. The counter goes quiet rather than drifting — the §7.7-preferred failure — but "quiet" here can also look like "nothing spent", and only the window's `since` says how long the file has been accumulating. - **The replay guard is memory-only.** A batch retried across a collector restart is counted twice; bounded by one batch, and accepted for the same reason §7.16 accepted its write race — a guard file would be a second store keyed on session ids. - **A record without `session.id` or `event.sequence` is counted unguarded** rather than refused: both ids exist on every measured record, and if a later grok drops them the honest failure is a counter exposed to duplicate retries, not one that silently stops. - **The schema is the vendor's alpha (`grok_code.schema.version = v1`)** and the collector does not read the version attribute. A rename lands as quiet non-counting — visible as a counter that stops moving, and §3.9a's capture is the shape to re-measure against. - **The endpoint redirect is prescribed, not measured** (2026-08-16 amendment). Moving the collector with `--addr` is measured; moving grok's exporter to meet it rests on the exporter library's documented variable, not on a re-run capture. The startup message and this section both say so, so nobody quotes it as verified later. - The capture behind every claim here is one machine, one day, one grok version, one signed-in account. The §3.4 discipline applies: re-measure before extending any claim. **Amended 2026-08-16 — who may push here.** The listener above took any loopback POST, and "loopback" was carrying more weight than it could hold: measured the same day, a web page on another origin planted a forged `api_request` in `usage/grok.json` from a real headless Chrome, with no local code running at all. §7.24 is the measurement and the fix. Two things changed here: `/v1/logs` and `/v1/metrics` now refuse a request carrying `Origin`, and both require `Content-Type: application/x-protobuf` — the media type this section's own capture pinned on grok's exporter, so a correctly configured grok is unaffected. A local *program* is still trusted completely and deliberately, because it can write the cache file directly; §7.24 states that boundary rather than pretending a token would move it. **Amended 2026-08-29 — this listener is no longer the only way to read what grok spent, and the one-envelope rule is what that costs.** §3.9a's `usage` re-measure at grok 1.0.5 found `inputTokens`, `outputTokens`, `totalTokens`, `cachedReadTokens`, `cacheCreationTokens`, `reasoningTokens`, `modelCalls` and `apiDurationMs` sitting beside `costUsdTicks` on every `turn_completed` record on disk — present since 1.0.0, and missed only because §3.9a's cost sweep was spelled `"[a-z_]*cost[a-z_]*"` and could not match them. **Nothing in this section is falsified by that.** The opening sentence claims the OTLP stream is the vendor's one *designed-for-reporting* surface and that is still true; a session-update log persisted verbatim is not a reporting surface. What is no longer true is the unstated corollary a reader would draw, that the push is the only way telltale **could** obtain grok's per-turn counts. It is not. The disk is a second source and a passive one: nothing to run, no double opt-in, no port to hold, and no batch lost while a collector is down — which is this section's first Known limitation, answered by a seam it did not know it had. Three constraints bind any future disk reader, and they are recorded now precisely because nothing is being built here. 1. **One envelope per turn, and the replay guard cannot enforce it across seams.** The guard above refuses a repeated `(session.id, event.sequence)` pair, in collector memory, over OTLP records. A disk reader keys on a file and an offset and would be invisible to it, so a machine running both would fold one turn into `~/.telltale/usage/grok.json` twice and the file would be overstated by an amount nothing could name. That is the failure §7.7 rates worst, arriving through the front door. **A disk reader and this listener must never both count the same turn.** Which of the two is the source is a design question, and this amendment does not answer it — it only forbids the answer "both". 2. **`input` does not mean the same thing on the two seams.** Measured on one turn: the wire reports `input_tokens: 22516` beside `cache_read_input_tokens: 256`, and the disk reports `inputTokens: 22772` for that same turn, which is the sum. The wire excludes the cache read and the disk includes it. This cache's `input_tokens` field currently holds the wire sense, because this collector is its only writer. A disk reader writing `inputTokens` into the same field would silently change what the column means, and history already in the file would not convert. 3. **The read budget is unmeasured and it is the deciding question.** `updates.jsonl` reached 818 KB in one observed session and is append-only, so whether a complete read is bounded — not whether the fields exist — is what decides a disk reader. §3.9a's cost row ruled a tail-window sum a lower bound, and a lower bound accumulated into a total is the derived-number refusal in a second unit. **The display stays held** by the owner's ruling above, so none of this has a consumer today, and the reader is deliberately NOT built in the lane that measured it. Recorded as the seam, not spent — the same way §3.9a recorded this listener's own seam before it was spent. ### 7.16b The Claude statusline's token block — measured, modelled, relayed nowhere (2026-08-16) The token relay has two writers (§7.16, §7.16a) and an obvious-looking third. Claude Code's statusline payload grew token fields, and the Claude statusline is **already a relay writer** — it writes the quota it just rendered (§7.15). A writer fed from a payload the gauge is holding anyway would cost one function and no new seam. **It was measured first, and the measurement refused it.** This section is the record of a seam that was NOT spent, which is worth as much as one that was: the fields are modelled and parsed, nothing renders them, and `~/.telltale/usage/claude.json` is never written. #### What was measured, and how Pinned at Claude Code **2.1.233** — `GIT_SHA f8d57569aaf350fe25dc4dfa10cad59db8ea4d45`, `BUILD_TIME 2026-08-14T17:21:48Z`, `DD_SOURCEMAP_GROUP win32`, the build installed at `~/.local/share/claude/versions/2.1.233` and the one `claude --version` reports. **The method was a source read of the shipped bundle, not a live capture, and that is a weaker instrument in one specific way.** A capture shim was installed at the dispatcher (`~/.claude/statusline.sh`, backed up and restored byte-identically) and it produced **zero real payloads** in a fifteen-minute window: the harness session that installed it renders no statusline, and an idle session does not re-fire one. Rather than leave a modified dispatcher on the owner's machine, the shim came out and the bundle was read instead. `CLAUDE.md` admits both instruments ("a live run, a source read at a pinned version") and §7.16's own Cursor payload rests on exactly this one. What a source read buys here that a capture could not is the **arithmetic** — a capture shows numbers, and the numbers are not the problem. The whole block comes from one function: ```js function TAw(e,t){let r=wMo(e,t);return{ total_input_tokens: e ? e.input_tokens + e.cache_creation_input_tokens + e.cache_read_input_tokens : 0, total_output_tokens: e?.output_tokens ?? 0, context_window_size: t, current_usage: e, used_percentage: r.used, remaining_percentage: r.remaining}} ``` and `e` is `TDr(messages)`, which walks the message list **backwards and returns on the first usage it finds** — the single most recent assistant message: ```js function TDr(e){for(let t=e.length-1;t>=0;t--){let r=e[t],n=r?QUe(r):void 0; if(n)return{input_tokens:n.input_tokens,output_tokens:n.output_tokens, cache_creation_input_tokens:n.cache_creation_input_tokens??0, cache_read_input_tokens:n.cache_read_input_tokens??0}}return null} ``` The vendor's own doc comments in the same bundle say it in words, and they agree with the code: `current_usage` is "Token usage from last API call (null if no messages yet)", `total_input_tokens` is "Input tokens currently in the context window (incl. cache reads/writes)", `total_output_tokens` is "Output tokens from the most recent API response". #### What `current_usage` actually is — and why the write half dies **It is one API call, and the two totals are a LEVEL rather than a counter.** `total_input_tokens` is not a session spend; it is what is sitting in the context window right now, which is why it includes cache reads. It is also **derived** — the vendor sums three fields of the same single call — so it is `§7.16`'s derived-`inputTokens` trap wearing a different label, and a plausible-looking number is the quiet kind of wrong. `internal/usagecache` accumulates **per-turn counts taken as-is**. Nothing in this payload is one. Three routes to a spend figure exist and all three are refused: - **sum `current_usage` across fires.** Every fire reports the same last call, and the statusline is debounced at 300ms and re-renders on state changes, so this counts one call an unbounded number of times. - **difference successive renders.** This is deriving a number and presenting it as a reading — the ADR-001 refusal this project exists for — and `usagecache.Delta`'s contract says the vendor must report the count, not telltale reconstruct it. - **treat `total_input_tokens` as a total.** It is an occupancy level. It goes **down** after a `/compact`, and a "total" that decreases is not one. So the write half is dead on the measurement, not on taste. **The one thing that would revive it is the vendor reporting a per-turn count** — and the bundle shows it already computes something close (`U5d` sums input + cache_creation + output across messages, de-duplicated by id) and does **not** put it in the statusline payload. If a future version does, this ruling is re-openable against that field and nothing else. #### What was built - **`internal/claude`**: `context_window` gains `total_input_tokens`, `total_output_tokens` and a `current_usage` object (four counts); `StatuslineInput` gains `prompt_id`. All are pointers or omitempty, because a pre-2.1.233 CLI sends no such key and "this CLI does not report it" must stay distinguishable from the measured zero `TAw` emits before the first message (§4a.1). `CurrentUsage` is the allowlist for its block, the `internal/cursorhook` technique reused. - **nothing else.** No renderer, no relay, no `usagecache` converter, no HUD field. - **pinned three ways**: `TestTheTokenCountsParseAtTheMeasuredVersion` asserts the fields land *and* that the fixture still models `TAw`'s arithmetic — the moment those numbers stop agreeing, the fixture has begun claiming a cumulative total nobody measured; `TestAnOlderPayloadLeavesTheTokenCountsAbsent` is zero-vs-absent on a schema that grew; `TestTheTokenCountsAreParsedAndNeverRendered` guards the render, and CI asserts on the **built binary** that no count reaches stdout and that `usage/claude.json` does not exist. #### Two claims from the research brief that the measurement corrected - **`prompt_id` is real, and it is not on the documentation page.** It arrives via the vendor's shared session-basics helper (`py`), as `pr.requestJournal.promptId()`, not via the statusline's own assembly. `internal/claude`'s package doc now says which of its fields are documented and which are measured-only. - **`aborted` and `api_retry` appear in NEITHER payload.** Every `aborted` in the bundle is `AbortSignal`/undici machinery or a telemetry reason string. `api_retry` is a **stream event type** — "Emitted when an API request fails with a retryable error and will be retried after a delay" — alongside `assistant`, `result` and `compact_boundary`. Neither is a statusline field at this version. Nothing shelved was reopened to establish this; it is this session's own read. #### Known limitations - **A live capture now backs the field list — closed 2026-08-16, the same day.** The dispatcher wore a tee for one interactive session and came out again; the host was the same pinned 2.1.233. Seven fires landed, all inside one prompt (one `session_id`, one `prompt_id`). Every fire satisfied `TAw`'s arithmetic: `total_input_tokens` equalled the three `current_usage` input fields summed (e.g. 2 + 1122 + 66355 = 67479), and `total_output_tokens` equalled `current_usage.output_tokens`. Two refusals above were also witnessed rather than argued: two adjacent fires repeated one call byte-for-byte on the token block (a summer counts that call twice), and `total_output_tokens` fell 169 → 3 between fires inside the one prompt — a "total" that decreases, with no `/compact` involved. The host sent 16 top-level keys, a strict subset of what the bundle can assemble: `vim`, `pr`, `remote`, `agent_type`, `workspace.repo` and `worktree` never appeared (the session's cwd was not a repo), and `permission_mode` stayed absent — §2's exclusion holds on the wire, not only at the call site. The capture also found one field the source read missed: `effort` (`{"level": ...}`), present on every fire; it joins the deliberately-unmodelled list in the next limitation. Still open: one session on one machine, `prompt_id`'s rotation across prompts unobserved, and the occupancy level's decrease after `/compact` unwitnessed (the output half's decrease above is the same property on the cheaper field). - **The payload has grown fields this struct still ignores** — measured present at 2.1.233 and deliberately unmodelled, because nothing needs them: `exceeds_200k_tokens`, `fast_mode`, `output_style`, `thinking`, `vim`, `pr`, `remote`, `agent_type`, `workspace.added_dirs`, `workspace.repo`, an expanded `worktree`, and three more `cost` fields (`total_api_duration_ms`, `total_lines_added`, `total_lines_removed`). The live capture added `effort` to this list — the source read missed it entirely. Unknown fields are ignored by design (§2), which is why that growth was a non-event. - **§2's "permission mode is not in the payload" survives, for a changed reason.** `py` CAN emit `permission_mode`, but the statusline calls it with two arguments, so the field is `undefined` and never serialized. The exclusion is now a property of the call site rather than of the schema, and a future version that passes the argument would ship it. ### 7.16c The cache hit ratio — the vendor computes it, so telltale may render it (2026-09-04) §7.16b is the record of a seam that measured out to nothing. This is its mirror image on the same payload: a number arrived that telltale is allowed to display, and the reason it is allowed is the whole content of this section. **The question was whether any surface telltale already reads passively carries prompt-cache counts.** Two do, and the difference between them decided where the work went. #### What was measured, and how - **The transcript on disk carries the raw counts and not the ratio.** A read-only pass over the owner's corpus on 2026-09-04 — 300 transcripts sampled from 1,508, 65,589 records — found `message.usage.cache_read_input_tokens` and `message.usage.cache_creation_input_tokens` on 32,416 records, written by six CLI builds from 2.1.209 to 2.1.258, plus the same two keys nested under `message.usage.iterations[]` and under `toolUseResult.usage`. No `prompt_cache` and no `hit_ratio` key occurs at any depth. `internal/adapter/claudecode` has parsed both counts since v1 and folds them into `contextIn()`. - **The statusline payload carries the ratio itself, computed by the vendor.** Source read of the shipped executable at **2.1.260** (the same field names are present in the 2.1.251 build on the same machine, and 2.1.251 is the floor the vendor documents). One function assembles the block and returns an empty object — so the whole key is absent — while `requests` is 0: ```js function WZt(w=Date.now()){let x=JUt(void 0,w); if(x.requests===0||x.lastRequest===null)return{}; return{prompt_cache:{warm:x.warm,caching_observed:x.cachingObserved,ttl:…, requests:x.requests,misses:x.misses,expected_rebuilds:x.expectedRebuilds, hit_ratio:x.hitRatio,cache_write_tokens:x.cacheWriteTokens,…}}} ``` and the ratio comes from the accumulator's `summary()`: ```js let n=this.cacheReadTokens+this.cacheCreationTokens+this.inputTokens; … hitRatio: n>0 ? this.cacheReadTokens/n : null ``` #### The ruling **A ratio derived from the transcript is refused; the reported ratio is rendered.** They are the same formula and they are not the same claim. Dividing the adapter's own counts would be arithmetic telltale invented, which is the ADR-001 §4a.1 refusal, and it would additionally be wrong in scope: the adapter reads a head+tail window of a file that reaches 7.7 MB, so a "session" ratio built there is a ratio over whatever the windows happened to cover, silently. The vendor's figure is taken over every main-conversation request of the session. telltale reads that quotient and multiplies by 100 — the unit conversion §2.1 already permits for Antigravity's `remaining_fraction` — and computes nothing else. So the adapter gains nothing and stays exactly as it was. That is the answer to "add a cache ratio to the Claude adapter": the honest place for it was the other Claude seam. #### What was built - **`internal/claude`**: `StatuslineInput` gains `PromptCache`, modelling four of the block's fourteen fields. `HitRatio` renders. `CachingObserved` gates it. `Warm` and `Requests` are parsed and rendered by nothing, each for a reason stated on the field. The struct is the allowlist for the block, the `internal/cursorhook` technique reused. - **`internal/statusline`**: one segment, `cache 91%`, placed directly after `ctx` because the two answer one question together — how full the window is, and how much of what fills it came from cache. - **No threshold colour, and this is a rule rather than a preference.** Every other percentage on that line is a consumption, so `pct()` paints high values red. A hit ratio inverts that, and nothing here has measured the ratio at which a cache becomes bad, so inventing an inverted scale would be inventing a judgment. The value renders unpainted and the word `cache` carries the distinction, which is what §9's accessibility rule asks of every distinction anyway. - **Three absences hide the segment and none renders as zero**: no block (a CLI older than 2.1.251, or a session before its first API response), a null `hit_ratio` (the vendor's own `n>0` guard), and `caching_observed` false (the provider or gateway reports no cache tokens at all). The last is the sharp one: it is an unread field, not a 0% hit rate. A ratio of 0 **with** caching observed is a reading and renders `cache 0%`. - **Pinned on both halves**: the parse tests cover the measured shape, both flavours of an absent block, and the `caching_observed: false` case; the render tests pin the line, prove the segment reads the reported ratio rather than the per-call counts sitting beside it in the same fixture (they would yield 75%, not 91%), and assert the value carries no ANSI. CI asserts on the **built binary** that the 2.1.233 fixture grows no cache segment and that the 2.1.260 fixture renders `cache 91%` and nothing else from the block. #### Known limitations - **This is a source read with no live capture behind it**, which is weaker than §7.16b ended up being. The formula, the field names and the absence semantics come from the 2.1.260 executable and the vendor's documentation page agreeing with each other; no interactive session was teed to watch a real `prompt_cache` block arrive on the wire, and no cold-cache or `caching_observed: false` payload has been observed. A capture would close all three at once and is the obvious next measurement. - **The docs page's field table is not the whole block.** The same source read found two fields it does not list: `last_miss_cause` (an object carrying `causes`, `tools_added`, `tools_removed`, `system_char_delta`) and `miss_causes`. That is §7.16b's `prompt_id` finding repeating, and it is recorded rather than modelled — absence of need is a result. - **Eight documented siblings are deliberately unmodelled**: `ttl`, `expires_at`, `misses`, `expected_rebuilds`, `cache_write_tokens`, `miss_recache_tokens`, `last_miss_at`, `recache_tokens_if_cold`. `warm` is the one to reopen first — a cold prefix makes a healthy ratio unusable, and a cold mark on the segment is a second claim that needs its own ruling. - **The statistics cover the main conversation only.** The vendor excludes subagent requests, so a fan-out's cache behaviour is not in this number and the segment does not imply it is. - **§7.16b's relay refusal is untouched.** Nothing here writes `usage/claude.json`; a ratio is not a count, and the two CI relay assertions still stand beside the new ones. ### 7.17 `u`: the fleet usage view — two claims, and never one Added 2026-08-09; amended the same day by the reading pass below (the models census, the title's rule weight, and the age escalation the 19-hour incident forced). §7.15 gave the header a block per vendor and §7.16 gave it a spend line, and between them they filled the one or two lines the header has. §7.10's last open limitation said the quiet part: *"the account quota block is sourced from one session. A second quota-bearing vendor needs a per-vendor block."* It has one now — but not in the header, because the header is answering a different question. **Glance and read are different jobs.** The header answers *am I about to run out?* in the time it takes to look up from an editor, and its whole design is a shed cascade that spends decoration to keep facts on one line (§7.15). It cannot also answer *what can telltale actually say about each of my five vendors, and where it says nothing, why?* — that answer is a paragraph per vendor, and a paragraph per vendor is a body, not a header. So `u` opens a third body over the row area, on the detail pane's precedent (§7.11): it replaces the grid rather than floating over it, because a panel covering the thing being monitored is a monitor you have to move to read. **The header was left unchanged, and that was wrong for this one body** — see *the header stops repeating the page* below, which reverses it. Over the GRID the two surfaces render the same readings from the same assembly and the duplication is the point, one for glancing at and one for reading. Over this body there is nothing left to glance at: the reader is already looking at the read surface. #### The organizing insight, and it came from the measurement "Usage" is **two different claims**, and the view may never blur them. | | **Quota** | **Spend** | |---|---|---| | what it is | a reading against a limit the vendor published | a count of tokens with no denominator anywhere | | sources today | Codex (its own store, scan-fresh); Claude and agy (statusline relay) | agy (summed from the conversations this scan read, §3.8) | | may render | gauge, percentage, reset countdown, severity hue | a verb, the counts, and the accumulation window | | may never render | — | a gauge, a percentage, a countdown, a bar, or the sum without its window | The right-hand column is not a style preference. There is no ceiling anywhere in a token count — no vendor here publishes an account limit a passive reader can see (§3.8, §3.9, §3.9a) — so a bar or a percentage would **invent** one, which is the same class of error as filling a `CapNone` field with a plausible guess. `TestUsageSpendBorrowsNoneOfQuotasVocabulary` pins it, and since the header's spend line was retired (§7.16's amendment) that test is the *whole* of the guarantee rather than the second half of a pair. It matters most here anyway: this surface puts the two measurements four lines apart under one vendor name instead of on separate header rows, and **proximity is what makes this the riskier render of the two**. #### The layout One block per vendor, in **fixed fleet order** — claude, codex, gemini, agy, cursor, grok — never sorted by usage. Position is the navigation: a vendor moving must mean a vendor was added or removed, not that another vendor's percentage crossed it. `fleetOrder` is now one variable, walked by both the header's per-vendor counts and this view, so a vendor sits in the same place on both surfaces. Each block is a heading that states the **quota seam** — where the reading came from, or why there is none — and then one line per fact under it, hanging off the detail pane's label column so the two bodies read as one product. The generated render (`usage-fleet` golden, and the shape it took after the 2026-08-09 amendment below): ``` telltale │ 7 sessions │ codex 1 gemini 1 agy 3 cursor 1 grok 1 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── fleet usage quota is a reading against a limit; spend is a count with none ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ claude quota relayed by the statusline · 2h ago 5h ████████──────────── 42% ↻ 2h13m 7d █▏────────────────── 6% ↻ 5d00h codex quota read from its own store, this scan models gpt-5.1-codex 7d ███████████████───── 79% ↻ 22h48m gemini no quota reaches disk anywhere telltale can read models gemini-3-pro agy quota relayed by the statusline models Gemini 3.6 Flash (High) gemini-weekly ███████▎──────────── 38% ↻ 3h00m spent uncached in 1.2M · out 13.1k · summed across 2 sessions on disk, this scan cursor no quota anywhere · its store holds experiment values, not usage models composer-2.5 grok no quota anywhere · no window, no ordinal, no reset time on its disk models grok-4.5 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── esc close ↑/↓ scroll ``` Five things in that frame are doing work: - **The gauge is 20 cells, not the header's 8.** That is the concrete reason this view exists rather than a wider header: one window per line buys the bar room to be *read* rather than merely glanced at. Fill resolution goes from 1.6% to 0.66%. It sheds to 12 cells at the compact tier and disappears below 80 columns, on §7.2's own breakpoints rather than on new ones — and it is allowed to shed only because the number beside it stays. - **The heading names the source, and §7.15 makes that load-bearing.** A transcript-sourced block is re-measured every scan; a relayed one is exactly as old as the last statusline render, and only the first may ever carry a burn forecast. A reader deciding how much to trust a percentage needs to know which they are looking at. The relayed reading's **age survives every dress level** — the phrase around it gives way instead — because shedding the age would re-present a stale number as fresh. - **`agy` carries both kinds of claim, and they arrive from two different seams.** Its percentage is relayed by the statusline; its token counts are summed out of the conversations this scan read. The heading always speaks about quota and never about spend, deliberately: the spend line explains itself in its own vocabulary (a verb and a window), while an absent or relayed reading explains nothing at all unless something says it out loud. - **`gemini` is one line, and it is a line rather than a row of dashes.** - **`cursor` and `grok` have a sentence each and no numbers, and the sentences differ** because the measurements behind them do. - **The header above it is identity only.** It carried the quota strip when this view shipped; the amendment below took it off, because every figure on it is restated underneath with room. #### Three kinds of nothing, and the one that had to collapse §4a.1's rule is that the kinds of absence stay distinct. On this surface there are three, and they are the whole reason the absence line carries a reason rather than an em dash: | | Vendors | What it renders | Why that wording | |---|---|---|---| | **structurally absent** | gemini, cursor, grok | `no quota reaches disk anywhere telltale can read` / `no quota anywhere · its store holds experiment values, not usage` / `no quota anywhere · no window, no ordinal, no reset time on its disk` | There is no seam to fire. Cursor's verdict was re-measured 2026-08-08 and came back harder: the only account figures on its disk are Statsig experiment values stamped `is_user_in_experiment:false`, never consumption (§7.16). grok's was measured the same way on 2026-08-09: a rate/limit/quota sweep of the whole store matched tool-configuration keys and nothing else (§3.9a). Naming an action here would send someone to enable a thing that does not exist. **Each vendor gets its own sentence rather than sharing one** — the verdicts are the same shape and different measurements, and lending grok Cursor's wording would claim something about grok's disk that nobody looked for there. | | **seam exists, never seen** | claude, agy | `no quota relayed yet · the telltale statusline writes it` | This is the one absence a user can act on, so it **names the statusline**. The reading turns up as soon as the gauge runs in that vendor. An absence with an action behind it that does not say the action is just a shrug. | | **aged out** | any relayed vendor | *renders as never-seen* | `quotacache`'s reader drops a window whose reset has passed and any entry over 24h old before the HUD ever sees it (§7.15's self-expiry). Telling the two apart would mean **holding numbers §7.15 calls not stale but FALSE** so this view could display them. Losing one distinction is the cheaper of those two trades — and it is recorded as a limitation below rather than left to be discovered. | Codex is in none of the three: its quota comes from its own store, so an absence there is a statement about what this scan read (`no quota in the sessions read this scan`) rather than about a relay that never fired. Borrowing either of the other two sentences for it would name the wrong seam. **A vendor with no sessions, no quota and no total does not appear at all.** The view is a report on what is running, not a checklist of every adapter that was compiled in — and an absence line is for a vendor that is *here* and silent, not for one that is not here. When nothing anywhere has anything to say, the body is a sentence rather than five blocks of dashes, because a table of nothing is a table asserting it measured five things (`usage-empty` golden): ``` telltale │ 0 sessions ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── fleet usage quota is a reading against a limit; spend is a count with none ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ no vendor on this machine has reported a quota reading or a token count ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── esc close ↑/↓ scroll ``` #### The reading pass (amended 2026-08-09): it worked, and it did not read like a page The view above shipped the same day it was designed and it was *correct* — every claim in it is measured, every absence names its seam. Asked whether it was the best the product could do, the answer was no, and for three separable reasons: it did not say **which models did the work** (half of the original ask, dropped in the first cut), it had **no visual nesting** (a body title at the same weight and the same column as its own entries), and an **old reading looked exactly like a fresh one** apart from four muted characters. The third one had already cost something real, which is why it is the longest item here. **The census: which models actually did the work.** Each vendor block now carries a `models` row naming the model display names this scan saw under that vendor. ``` telltale │ 6 sessions │ claude 4 codex 1 gemini 1 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── fleet usage quota is a reading against a limit; spend is a count with none ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ claude no quota relayed yet · the telltale statusline writes it models Haiku 4.5, Opus 5, Sonnet 4.5 codex quota read from its own store, this scan models gpt-5.1-codex 7d ███████████████───── 79% ↻ 22h48m gemini no quota reaches disk anywhere telltale can read ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── esc close ↑/↓ scroll ``` It is a **third kind of line**, and it belongs to neither of the two claims this view is built around. A census has no limit and no total — nothing to compare it against — so it borrows neither vocabulary and simply lists what was there. What governs it is the same rule everything else here obeys: | Rule | Why it is the honest-gauge rule again | |---|---| | **Only this snapshot** | Never a remembered list and never the vendor's catalogue. A name surviving from a previous scan would be a claim about the past presented as the present — the defect the relay's age exists to prevent, arriving through a different door. Four Claude sessions above; take them away and the row goes, even though the vendor still has a quota reading. | | **The grid's own normalization** | Through `DisplayModel`, so `claude-opus-5` reads `Opus 5` here and in the MODEL column, and a reader never has to work out that two spellings are one model. | | **Deduped, sorted** | Four sessions, three names. Alphabetical rather than by recency or by session count: ranked ordering reshuffles every time a turn lands, and §7.1 rule 4 budgets the movement on this screen at one cell. | | **Absent renders absent** | A vendor whose sessions carry no model gets **no row** — `gemini` above has a session and no census. Not an em dash: the dash would claim telltale looked at a model and could not name it, when what happened is that the adapter sources no model at all. | | **Overflow is announced** | `+3 more`, never a clipped list. The cap is the *room* rather than a magic number, and names are dropped whole rather than cut, with the marker's own width reserved before the last name is accepted. At the 60-column floor: ` models Haiku 4.5, Opus 5 +3 more`. An ellipsis there would leave the reader unable to tell whether one name went missing or nine. | **It is its own row rather than part of the heading**, and that was the one real layout fork in this pass. The obvious placement is beside the vendor name — `claude · Opus 5, Sonnet 5 quota relayed by the statusline · 2h ago` — and it is wrong for two reasons that point the same way. First, the heading has one job and §7.17 already made it load-bearing: it states the **quota seam**, and the relayed reading's age may never shed from it. A second variable-length vendor-supplied fact on that line puts a census in competition with the one thing on the surface that is not allowed to give way, and at 60 columns the census wins by being at the front. Second, the census is a *session* fact aggregated per vendor while the heading speaks about the *account* — putting them on one line blurs exactly the distinction the block's shape exists to draw. In the label column it instead joins `5h`, `7d` and `spent` as a fourth labelled fact about one vendor, which is what the column is for. **Within the block it comes first**, before the quota windows and before spend: it names who did the work and the rows under it say what that work cost, subject before predicate. It does not tear the heading from its evidence, because the heading states a provenance ("relayed by the statusline") rather than a number. #### The title carries the room's second rule weight `fleet usage` and ` claude` started in the same column at the same weight, so the body's **title read as a peer of one of its entries** — §9.23's finding one surface over, where a turn page's outline whispered while its entries shouted. The HUD had one rule weight and was asking it to be both the frame's edge and a gauge's empty track. `RuleHeavy` is `━` and `=`, **council's own pair** rather than a second one: §7.1 principle 5 is that these are one product, and a second heavy-rule character would be a second alphabet. `=` is the one unclaimed mark left in the HUD's reduced set — `-` is the light rule and the gauge track and the fact separator and a spinner frame, `#` the gauge fill, `|` the separator, `>` the ellipsis, `~` the reset, `!` the warning, `]` the cursor, `Y` the fan-out, `*`/`o`/`.` the state dots, `_` the caret — and `TestTheHeavyRuleHasAnUnclaimedASCIIPartner` enumerates that list so the next glyph cannot be added without meeting it. **The rule goes ON the title, not under it**, and that is council's ruling rather than a preference: §9.11 spent a whole item removing a heading followed by a horizontal rule, on the finding that such a rule says nothing the heading had not, and ruled that a heading carries its own. Here it also avoids a specific defect — the frame's own light full-bleed rule sits one row above, and a *heavier* line three rows inside a lighter outline is §9.26's hierarchy argument inverted. It costs **zero rows**, which is what makes it affordable on a body with a line budget, and it yields to the legend rather than the other way round: the note is fitted first and the rule takes what is left, because a legend is a statement and a rule is chrome. At the 60-column floor that leaves ten cells; below `usageRuleMin` the line simply has no rule. **The vendor headings deliberately get no rule, not even the light one**, and `TestOnlyTheUsageTitleDrawsTheHeavyRule` asserts the weight as a *count* on the rendered frame — one line, one run, and zero on every other body. §9.26's argument is that a second weight is worth exactly what it is scarce; a rule on every block would spend it five times a screen to restate what an indent, a blank row and the identity hue already say. #### Air and alignment: what was already right, and what is now pinned The columns were already shared — one `usageLabel` cell and one `usageGap` for every fact row, so `claude`'s `5h` gauge starts where `codex`'s `7d` gauge and `agy`'s `gemini-weekly` gauge start — and the blocks were already one blank row apart. This pass **pinned that rather than built it**, which is the honest description: `TestUsageFactsShareOneColumnGridAcrossVendors` now walks the rendered body at four widths and fails if the label column, the gauge, the percentage or the reset countdown lands in two different columns across vendors, and if any label runs into the value column. Before it, the models row could have been added with a layout of its own and nothing would have noticed. Two things were considered and **declined**. A second blank between blocks: §9.11's threshold is that every deliberate blank this product draws is exactly one row, and a one-row gap is a boundary placed between two things meant to be kept together — two rows is nothing the design asked for. And a light rule under each vendor heading, for the scarcity reason above; air is the boundary strength this body can afford, and it already has it. > **The first of those was reversed on 2026-08-09** — narrowly, and only where the rows are > otherwise going to waste. See *the page stops trailing off*. The rule under each vendor > heading stands declined. #### An old reading has to look old (the 19-hour incident) **What happened, 2026-08-09.** A Claude relay entry written nineteen hours earlier reported 15% of the seven-day window. It rendered at full confidence beside a live gauge and was read as current. The account was at 44%. Nothing in it was dishonest. The age was on screen — `· 19h ago`, exactly as §7.15 requires — and every part of the state is one the product genuinely reaches: the five-hour window was gone because `quotacache` drops a window whose reset has passed, the entry survived because it was inside the 24h ceiling, and the reset it reported was still four days out so nothing upstream had any reason to touch it. **The age was present and it was not loud.** A muted four-character suffix is the same weight as every other piece of chrome on that line. ``` telltale │ 7 sessions │ codex 1 gemini 1 agy 3 cursor 1 grok 1 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── fleet usage quota is a reading against a limit; spend is a count with none ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ claude quota relayed by the statusline · ⚠ 19h ago · older than the fleet's shortest quota window 7d ██▉───────────────── 15% ↻ 4d07h codex quota read from its own store, this scan models gpt-5.1-codex 7d ███████████████───── 79% ↻ 22h48m gemini no quota reaches disk anywhere telltale can read models gemini-3-pro agy quota relayed by the statusline models Gemini 3.6 Flash (High) gemini-weekly ███████▎──────────── 38% ↻ 3h00m spent uncached in 1.2M · out 13.1k · summed across 2 sessions on disk, this scan cursor no quota anywhere · its store holds experiment values, not usage models composer-2.5 grok no quota anywhere · no window, no ordinal, no reset time on its disk models grok-4.5 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── esc close ↑/↓ scroll ``` **Five hours, and the number is argued rather than tuned.** `quotaAgeWarn` is the shortest quota window telltale has measured anywhere in the fleet — Claude's `five_hour` (§3.1) — and therefore the shortest span over which a vendor is known to reset a limit *wholesale*. A relayed reading older than that has outlived the fastest-moving quota this product knows about: whatever window it reports, an entire window of the shortest kind could have opened and closed since it was taken, so the reader may no longer assume the number describes now. Below five hours the reading is old and still bounded by something; above it there is no window short enough to bound it. It pairs with `quotaAgeShown` (5m, where the age starts rendering at all) as the two boundaries of a relayed reading's life on screen. Three thresholds it deliberately is **not**: - **Not the per-block shortest window**, which was the first candidate. `model.QuotaWindow` carries no duration — a label, a percentage and a reset time — so a per-block rule would have to infer a length by parsing `"5h"` or `"seven_day"`, and a threshold derived from a display string is the class of guess §4a.1 rejects. It would also have failed on the very reading that prompted this: the expired 5h window was already gone, so the surviving block reported only `7d` and a per-block rule would have stayed silent at nineteen hours. - **Not a second reason to drop a reading.** `quotacache` owns expiry (§7.15) and this view renders everything it is handed, changing only how loudly. Dropping earlier than the reader does would hide a measurement telltale holds. - **Not a freshness gauge.** There is no denominator for "how fresh", which is the spend line's argument in a different costume. **Word first, glyph second, hue third**, in that order, because §7.1 rule 2 is that colour never carries a distinction alone. Past the threshold the heading's dress gains the warning glyph beside the age *and* the reason in words — `· ⚠ 19h ago · older than the fleet's shortest quota window` — and the whole statement renders in `SevWarn`, which §7.5 already defines as the token for warning notices and which the footer's own `⚠ last scan 1m ago` already uses. In the reduced set that line reads `claude quota relayed by the statusline - ! 19h ago - older than the fleet's shortest quota window`. It says "the fleet's" rather than "its" on purpose: five hours is the shortest window in the fleet, not necessarily in this block, and an agy weekly reading six hours old has not outlived its own window. Trading one over-confident render for one over-stated warning is not a trade. The shed cascade is unchanged in grammar — the age is the fact and never sheds, the sentence explaining it is decoration and does, so the barest level is `⚠ 19h ago`. What a narrow terminal loses is the argument, not the alarm. And the **reading itself is untouched**: the 15% keeps its own severity hue, because that is a statement about the account and this is a statement about the measurement. **There is exactly one step.** §9.26's lesson is that a second level is worth what it is scarce, and the only boundary above this one that is not invented is `quotacache`'s 24h drop — which is a disappearance rather than a louder warning. #### Amended 2026-08-09: the header escalates too The pass above deliberately left the header alone, and recorded that as unfinished work rather than as a boundary. It was the wrong half to fix first: the reading that was acted on was read off the **glance** line, so the product ended up loud on the surface a reader opens on purpose and quiet on the surface they merely look at. The header now escalates on the same threshold, in the same order, off the same constants — `quotaAgeWarn` and the reason string are read from §7.17's declarations rather than restated, because two copies of `5 * time.Hour` is how the two surfaces would come to disagree about one reading. Past the threshold the header's age suffix becomes `· ⚠ stale 19h ago`, in `SevWarn`, with `· older than the fleet's shortest quota window` beside it at the most dressed level. The reduced set renders `- ! stale 19h ago`. The reading itself is untouched — 15% keeps its own severity hue, for §7.17's reason: that is a statement about the account and this is a statement about the measurement. **The header keeps a WORD where the view's barest level does not**, and that is the one place the two surfaces deliberately differ. §7.17's shed bottoms out at `⚠ 19h ago`, because by then the reader has opened a body on purpose and the sentence above it is still on screen. The header has no sentence anywhere and is read at a glance, so the level that survives every shed has to state the verdict without the reader knowing what `⚠` is supposed to cost them. `stale` is that word: it names a state rather than restating the duration (`19h old` would be the number a second time), and it is the word this codebase already spends on a reading that has outlived its currency — `DotStale`, and the footer's own `⚠ last scan 1m ago`. **What sheds is the argument and only the argument.** `ageReason` joins the dress ladder as a sixth level above the existing five and is the first thing dropped, ahead of forecasts: it is by a wide margin the longest clause on the line, and it is the only part of the escalation whose absence costs a reader something they cannot otherwise see. Everything below it — glyph, word, age — survives to the barest level, so a narrow terminal loses why the reading is distrusted, never that it is. Same grammar as the view, one surface over. #### Everything else is inherited, not invented - **`u` opens and closes it; `esc` closes it.** The esc chain gains a step and now reads usage → detail → help → clear the query → quit. One body at a time: opening any of the three closes the others, enforced in `Update` rather than in `Render`, because a pane that appears only because it won an ordering is a pane nobody can predict. Find mode closes it too — but only in that direction, since once the mode has the keyboard `u` is a letter (§7.8). - **`u usage` joins the footer hints** beside `enter detail`, and sheds on the same tier boundary for the same reason: below 80 columns the footer keeps only the keys nothing else can teach. It is on the help overlay's keys page at every width. - **The drift notice renders under this body too** (§7.3), like every other. A warning that comes and goes depending on which pane is open is one a reader cannot trust to be there. - **Overflow scrolls** with the help overlay's vocabulary — `↑`/`↓` move the body, bounded against its own rendered length — not the grid's `+N more`, which counts sessions. - **The 60-column floor and the height tiers are unchanged**, and nothing here animates (§7.1 rule 4). At the floor the bars are gone and every fact survives, including the relayed reading's age and the spend total's window. - **The view is not narrowed by `v` or by the find query.** Those narrow the *session* list, and nothing on this surface is a session fact — filtering an account reading by a session filter would be the per-row quota §7.1 forbids, arriving by the back door. #### Declined - **Trend sparklines.** The caches hold one reading per vendor, so a trend line would be drawn from data that does not exist. The burn forecast (§7.12) is the honest version of this and it is confined to the one scan-fresh block that can support it. - **Sorting by usage.** Position is the navigation; see the fleet-order note above. - **Per-row quota.** An account fact is not a session fact — §7.1's sixth rule, and the reason this view exists at all. - **A fabricated fleet total.** Different units (percentages of unrelated windows, and raw token counts), different accounts, different vendors. Any single number across them would be arithmetic telltale invented, which is the ADR-001 violation this whole product is built to refuse. - **A spend line for grok**, added 2026-08-09 and the closest call in this section. grok is the only vendor here that writes real money to disk, so it is the one block where a reader might expect a figure — and it gets none. The only cost on its disk is `usage.costUsdTicks` on each `turn_completed` record: per-turn, not cumulative, in an append-only file that reached 818 KB in one session (§3.9a). A tail-window sum is a **lower bound**, and a lower bound rendered next to the word "spent" is a derived number wearing a read one's clothes. The last turn's cost is already a labelled Extra in the detail pane, where its label says which turn it belongs to. Nothing on grok's block says any of this: the heading speaks about quota and never about spend, and buying one vendor an exception would cost every other block its meaning. #### Amendment, 2026-08-09: the owner's ruling on which vendors this speaks for Ruled by the owner: **Cursor's spend display goes; agy and grok come on.** The retirement half is §7.16's amendment. This is the half that adds. **agy's spend line, and why its window reads differently.** The source is the agy adapter's measured per-conversation token counts — `gen_metadata`'s `#1.#4.#2` (uncached input) and `#1.#4.#3` (output), guarded by the `thinking + answer == output` identity §3.8 requires and the adapter asserts. Those were already on screen per row as display-only extras; what is new is `model.Session.Tokens`, the same two numbers as integers, summed per vendor by the view. The integers exist because a sum of pre-rounded display strings is a sum of roundings; the adapter sets both in the same branch from the same variables, so a row and the fleet total cannot disagree about what was counted. **This is a scan, not a meter, and the wording has to carry that.** Cursor's total was a file that only ever went up. agy's is a sum over the conversations that are on disk at this moment, so *deleting a conversation makes it smaller*. §7.16's rule — the sum never prints without its window — therefore binds against a different window: - the wording is `summed across 2 sessions on disk, this scan`, and it never says "since ". A "since" is a meter's claim and this is not a meter. - the shed cascade drops words and never facts: `summed across N sessions on disk, this scan` → `summed across N sessions on disk` → `across N sessions on disk` → `N sessions on disk`. The **count** survives every level because it *is* the window, and **"on disk"** survives every level because it is the difference between this sum and a monotonic one. "summed" and "this scan" are what give way, in that order. - the count is of the sessions that **contributed a measured reading**, not of the vendor's sessions. The fleet fixture's agy has three conversations and the line says two: the third has not called a model yet, carries no counts, and is in neither the sum nor its window. Folding it in as a zero would put a session in the denominator that contributed nothing to the numerator — §4a.1's rule applied to a window instead of to a cell. - a generation that **failed its self-check** is dropped by the adapter and named in that row's Diagnostics, which is where a reader finds out a total is over fewer generations than the conversation ran. That is inherited behaviour, not new, and it is the one place this surface is quieter than the row it summed: the count says how many sessions, not how many generations inside them were refused. Recorded as a limitation below. - the label is **`uncached in`**, not `in`. §3.8 marks the cache-read component's field number lower-confidence and the adapter refuses to fold it into a rounder total; labelling the number `in` would quietly promote a partial figure to a whole one, and the fleet line is the one place a reader could not catch that. **And it wraps rather than sheds, which nothing else on this surface does.** Every other line here sheds decoration — a gauge that re-states a number still on screen, a phrase around an age. This line is facts only: two counts, a label that has to say "uncached", and a window that may not go. At 60 columns those do not fit on one row and there is nothing left to spend. The header solved that class of problem by dropping whole vendor blocks, because a header has a hard one-or-two-line budget; **this is a body and it scrolls**, so it can pay a second row, and a second row costs a reader nothing next to a sum that has stopped saying what it summed. The window hangs under the counts in the same indent and re-runs its own cascade there — a row to itself buys back dress the shared row could not afford, so the *narrow* render says more about the window than a cramped single-line one would have. It carries no leading glyph: the mid dot separates facts on a line everywhere else in this product, and giving it a second job as a continuation mark is the failure §9.26 is a whole section about. `usage-floor` (60 columns): ``` agy quota relayed by the statusline models Gemini 3.6 Flash (High) gemini-weekly 38% ↻ 3h00m spent uncached in 1.2M · out 13.1k summed across 2 sessions on disk ``` **grok's block** is the absence table's third structural row, above. It qualifies for a block on sessions alone, and it would have rendered the un-surveyed fallback (`no quota telltale can read`) — honest, and a step down from a sentence that names what was measured. It now says `no quota anywhere · no window, no ordinal, no reset time on its disk`, from §3.9a's sweep. It has no spend line, for the reason in Declined above. Added by the reading pass: - **A freshness gauge.** No denominator for "how fresh", so a bar would invent one — the spend line's argument, one field over. - **A second escalation step for the relay's age**, and a per-block threshold parsed out of a window label. Both above. - **A light rule under each vendor heading.** The air-and-alignment note above. (A second blank row between blocks was on this list too, and came off it — see below.) #### Amended 2026-08-09: the header stops repeating the page **The defect, from driving the real thing.** With the usage body open the header still drew the full quota strip — two cramped rows of `ag 3p-5h 0% ↻2h10m 3p-weekly 0% … cc 5h 40% ↻2h17m 7d 62% … cx 7d 37% … cursor spent in 47.8k out 66 …`. Every fact on them is stated properly in the blocks four rows below, with a label column, a 20-cell gauge and the provenance sentence the strip has no room for. It was the same measurement twice, once cramped and once legible, with the cramped copy on top, and it was the ugliest thing on the screen. **Over the `u` body the header collapses to identity and session counts** — `telltale │ 144 of 1380 sessions │ claude 961 codex 299 …` — and the quota strip does not render. The grid keeps today's header exactly. Two rows come back to the body. This **reverses #163**, which ruled the header untouched when this view landed. #163 was right for the grid and wrong for this page, and the principle worth keeping is the one that distinguishes them: > A **glance** surface may not repeat the **read** surface it sits above. The glance line earns its rows by being the only statement of a fact. Over the grid it is: account quota appears nowhere else, so the strip buys a fact that is otherwise off screen. Over the page built to state those same facts at length, it buys nothing and costs two of the rows that page is short of. Identity survives the collapse for exactly that test — the session census is not restated anywhere below, because the usage blocks speak about accounts and never about sessions. One consequence worth naming: at the 60-column floor the recovered row is enough for `grok`'s census row to fit, which it did not before. The fix pays for itself in content on the narrowest terminal telltale renders on. #### Amended 2026-08-09: the page stops trailing off **The defect, same session.** On a tall terminal the content stopped around 40% of the way down and the rest was blank. Nothing was wrong with any line; the page simply had no bottom. A surface built to be *read* trailed off like a truncated file. Council solved this exact problem twice (§9.23's contiguous rails, §9.11's boundary-strength grammar) and neither answer ports: the HUD has no rails, and importing them would give this one body a vocabulary no other surface in the product speaks. The fix is in the HUD's own language, and it is two moves that only work together. **1. The closing rule hugs the content.** The frame's bottom rule was already being drawn — sixty rows below the last vendor block, where it reads as the terminal's edge rather than as the page's. Moved to where the content stops, the same rule makes the body a bounded region with a visible bottom edge, and the leftover rows fall *outside* it: unused terminal, not unfinished page. This is §9.11's boundary-strength grammar with the weight the frame already owns, moved — no new glyph, no new hue, no second rule. The grid is untouched, because its body is a list and a list that ends early has ended; a hard edge under a row area still open for more rows would claim something false. **2. The blocks breathe into two-row gaps.** This is the reversal of the air-and-alignment ruling above, and it is narrow. That ruling was made under a **line budget**, on the assumption that a row spent on air is a row taken from a fact. On a tall terminal the assumption is false — the rows are there, unspent — and the real choice is air versus void. So the second row is never a constant and never a fiat: it appears only when the widened gaps consume fewer rows than the page would otherwise leave blank at the bottom, which makes it a **redistribution of air the page already has**. Short terminal, tight page, exactly as before. The gap before the *first* block never grows, so the page still starts where the page starts — the §7.x anchor rulings hold and this is not vertical centring. **It stops at two, and the cap is the argument.** Distributing all the surplus — justifying the blocks down to the closing rule — was the obvious alternative and is declined twice over. It makes gap height a function of terminal height and block count, so the distance between two vendors would encode nothing while looking like it encoded something. And it puts every gap on a variable, so one grok session appearing reshuffles the vertical position of every block below it, against §7.1 rule 4's one-cell churn budget. Air that says "these are separate things" has to be a constant to say it. Also declined: **inventing content to fill the space** (a fleet total is §7.17's own rejected list; anything else would be a number nobody measured), and **centring the block vertically**. The `usage-tall` golden pins both halves at 52 rows: ``` telltale │ 7 sessions │ codex 1 gemini 1 agy 3 cursor 1 grok 1 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── fleet usage quota is a reading against a limit; spend is a count with none ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ claude quota relayed by the statusline · 2h ago 5h ████████──────────── 42% ↻ 2h13m 7d █▏────────────────── 6% ↻ 5d00h codex quota read from its own store, this scan models gpt-5.1-codex 7d ███████████████───── 79% ↻ 22h48m gemini no quota reaches disk anywhere telltale can read models gemini-3-pro agy quota relayed by the statusline models Gemini 3.6 Flash (High) gemini-weekly ███████▎──────────── 38% ↻ 3h00m spent uncached in 1.2M · out 13.1k · summed across 2 sessions on disk, this scan cursor no quota anywhere · its store holds experiment values, not usage models composer-2.5 grok no quota anywhere · no window, no ordinal, no reset time on its disk models grok-4.5 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── esc close ↑/↓ scroll ``` #### Per-vendor hues: ratified, and exactly as far as council's went Answered 2026-08-09 (San), and it is a **yes**. The question §7.17 left open — whether council's seat-hue exception (§9.28) extends to this surface's vendor names — turned on whether the HUD has the concept the exception was granted for. It does, on this body and nowhere else: **six vendor blocks stack in one column**, each a heading with a paragraph under it, so position answers nothing about which vendor a reader is looking at. That is the exact condition §9.28 named. The grid does not qualify and does not get it — a row's vendor is already answered by a two-letter tag in a fixed column. It arrives on the terms the open question itself set, and every one of them is council's rather than a new argument: - **Names only.** The vendor NAME in each usage-block heading, and nothing else. Not the seam sentence beside it (chrome, and past `quotaAgeWarn` a warning that must keep `SevWarn`), not the gauges, percentages or countdowns (severity owns those), not the spend line, not the models census (theme's identity hue, the same token the grid's `MODEL` column spends), not the grid rows, and not the header's quota block — the glance surface names its vendors by tag in a fixed order, so the question the hue answers is not the one it is asking. `TestTheVendorHueIsSpentOnlyOnTheVendorName` walks the rendered frame and fails if a hue reaches the header, a fact row or the footer. - **The assignments are council's, matched by literal value** — claude `5`, codex `6`, agy `4`, cursor `12`, grok `14`, gemini falling back to the identity hue. A reader who learned in the room that magenta is Claude must not meet a second colour for Claude one keypress away. `TestVendorHuesMatchCouncilsSeats` writes council's numbers out and fails when one copy moves without the other — the shape `TestStripTagsMatchTheHUDSpelling` already uses for the two-letter tags, and for the same reason: the seam between the two surfaces is the normalized session model and `internal/theme`'s numbers, and reaching across it for a rendering detail is the coupling that seam exists to prevent. - **The map does NOT go into `internal/theme`**, and the stdlib rule is not why. These are plain strings and would compile there. The reason is theme's own contract — one hue, one meaning, across every surface that imports it — and `internal/statusline` has no vendor blocks. Two packages holding the same map is the honest cost of keeping that contract intact, and the parity test is what makes the cost bounded. - **4-bit indices, severity and chrome off limits.** `1`/`2`/`3` and their bright twins `9`/`10`/`11` are the ramp this very surface draws its percentages in — a vendor heading wearing red would read as an account in trouble — and `0`/`7`/`8`/`15` are the gauge track and the terminal's own fore/background. `TestNoVendorHueIsASeverityOrChrome` fences it. - **`PlainStyles`-identity by construction, and zero golden churn.** One `retint` helper returns the base style untouched when `Plain` is set, so no golden on this surface can see the feature exist; **any golden diff on this change is a bug**, and none was produced. `NO_COLOR` needs nothing new — `colorprofile` downsamples inside Bubble Tea exactly as §7.5 describes. The **honest weakness is council's too**: `4`/`12` and `6`/`14` are two pairs of one hue at two intensities, and some schemes render each pair close. Council carries the distinction on the two-letter tags; here the thing being tinted is the vendor's full name, spelled out at the head of its own block, so this surface has more carrying it than the room does. Declined inside the ratification, and the list is closed: - **The hue on the grid rows and the header.** Position and the two-letter tag already answer "which vendor" on both, so it would be the circus row §9.28 refused for council's column headers, spent on a question the layout had already settled. - **A hue on the vendor tag as well as the name**, wherever the two appear together — double the ink for a distinction the name already carries. Council's own declined list. - **Gemini given a hue of its own.** The legal set has one index left (`13`), and spending it to complete a set is a palette entry with no argument behind it. Gemini takes the identity-hue fallback and looks exactly as it always did. #### Known limitations - **An aged-out relay reading is indistinguishable from one that never arrived**, by the trade in the table above. Both render as "no quota relayed yet". - **The header and this view still say some of the same things twice — just never at the same moment.** The 2026-08-09 amendment above removed the on-screen overlap by collapsing the header while this body is open, so no reading appears in two places in one frame any more. What survives is the *maintenance* half of the limitation: two renderers still speak for the same measurement, and a future change to one has to be made in both. `quotaVendors` being shared by both surfaces is what keeps them from disagreeing about *which* source speaks for a vendor; nothing yet keeps them from diverging in tone. - **The void fix is decided per frame, so a resize can change the gaps.** `usageAir` widens the blocks apart only when the taller layout still fits, which means dragging a terminal across the threshold moves every gap by a row at once. It is a resize, not a tick — §7.1 rule 4 budgets *frame-to-frame* churn on a still screen and this cannot fire on one — but it is a visible jump at the boundary rather than a smooth one, and smoothing it would mean the variable-height gaps the amendment declines. - **The absence sentences are per-vendor literals**, so a seventh vendor arrives with the fallback wording (`no quota telltale can read`) until someone measures its seam and gives it a sentence. That is the honest default — it claims nothing about a seam nobody has looked at — but it is a step down from the five that name theirs. **`grok` was that case live** for one day: it landed as the fleet's sixth vendor (#183), flowed into the block layout and the shared column grid with no change to either, and took the fallback sentence because nobody had measured what its store says about an account. §3.9a's sweep supplied one on 2026-08-09 and it now names its own seam; the fallback is back to being a path nothing currently takes. - **The spend line's count is of sessions, not of generations.** A conversation whose read dropped some generations for failing the `thinking + answer == output` self-check still counts as one contributing session, and its partial sum is in the total. The drop is named in that row's Diagnostics and nowhere on this surface, so a reader looking only at the fleet line cannot tell a whole conversation from a partly-refused one. The alternative — excluding the whole conversation — would discard generations that passed their own check and undercount by an amount nothing on screen could name, which is the worse of the two silences. - **The spend line's window shrinks silently when a conversation is deleted.** "on disk" is the only thing saying so. There is no honest alternative from a passive read: the scan cannot know what was on disk yesterday without keeping its own history, and a cache of previous scans would be telltale asserting a past it did not measure at the time. - **One over-age block costs the whole header line its gauges.** The escalation is decided per block, but the dress cascade is decided per LINE — one level has to fit every block on it — so the eight columns `⚠ stale` adds to a single vendor can push the line down a level that every vendor pays for. That is exactly what the `usage-stale-relay` frame above shows: at 120 columns with three vendors the header was already flush against its budget, so escalating Claude's reading dropped the bars from agy's and Codex's fresh blocks too. The trade is deliberate and it is the cascade's existing grammar rather than a new rule — the percentage beside each bar carries the reading, and a quietly stale number is a worse failure than a missing bar — but it does mean the frame that demonstrates the fix is also the frame that pays the most for it. A per-block dress would fix it and is declined: blocks of different heights and vocabularies on one line is a grid a reader has to reconstruct, and §7.2's whole argument is that they should not have to. - **The models census counts sessions, not turns.** A model that ran once six hours ago and a model that has been running all morning are the same entry in the row. There is no per-model activity anywhere the HUD reads, so weighting the list would be an invention — but a reader who takes the order for importance will be wrong, which is part of why it is alphabetical rather than ranked. ### 7.18 The scan keeps up: what a poll actually costs, measured (2026-08-09) The footer's `⚠ last scan Ns ago` (`view.go` `staleAfter = 3s`) was on permanently on the owner's machine. That notice was **correct** — the scan really was taking longer than three seconds against a 1 s `pollInterval` — which is the worst version of the problem: a true signal that fires constantly stops being read, and the one field the HUD has for saying "do not trust what you are looking at" becomes wallpaper. Raising `staleAfter` was considered and rejected outright; it hides the measurement instead of fixing what it measures. **Measure first.** Profiled against the live corpus on 2026-08-09 (Windows 11, i7-7700K, 1,404 sessions across six vendors — 967 claude, 346 codex, 65 agy, 53 grok, 7 cursor, 1 gemini). Nothing from that corpus is in this repository; only the timings are. | vendor | refs | discover | read (all refs, gate 8) | per-read p50 / p95 | |---|---|---|---|---| | claude | 967 | 29 ms | **2.612 s** | 12.8 ms / 63.7 ms | | codex | 346 | 27 ms | 138 ms | 1.0 ms / 13.2 ms | | agy | 65 | 1 ms | 42 ms | 2.5 ms / 14.4 ms | | grok | 53 | 5 ms | 11 ms | 1.0 ms / 3.3 ms | | gemini | 1 | 0 ms | 3 ms | 3.2 ms | | cursor | 7 | 0 ms | 2 ms | 1.0 ms | Whole `Scan`, warm: **1.84 s / 1.91 s / 3.37 s** over three consecutive runs. Four of the five things worth suspecting were already fine, and saying so is the point of recording this: - **The vendors are already concurrent.** `Scan` fans out one goroutine per adapter and `readAll` fans out per ref behind a semaphore of 8. Serialization was not the finding. - **The reads are already bounded.** Head 64 KB + tail 128 KB held at this corpus size — 164 MB read against 693 MB on disk. Cursor's SQLite store was already snapshot-cached on a two-stat check, and grok's `updates.jsonl` never showed up in the profile at all. - **The scan is already off the UI goroutine** (`scanCmd`), so a slow scan lagged the display; it never froze input. - **The 8-hour idle filter is genuinely expensive** — 1,235 of 1,404 sessions were read and then hidden — but see the rejected optimization below. The finding was **JSON parsing, and specifically re-parsing work that had not changed**. Broken down over claude's 967 files: open 73 ms, stat 14 ms, head read 185 ms, tail read 364 ms, subagent stat pass 321 ms, and **`json.Unmarshal` 2.65 s** — 46,727 records, every second, of which on a typical tick approximately none had been written since the last one. A pure `os.Stat` pass over the same 967 files costs 50 ms. **The fix** is a per-transcript parse cache in `internal/adapter/claudecode`, keyed on the file's `(size, mtime)` — the same shape `internal/adapter/cursor` already uses for its store snapshot. The honesty constraints decided its boundaries, not convenience: - **`last_activity` is not cached.** The §6 Q8 ruling makes it `max(mtime, newest record timestamp)` and *both* inputs move, so only the record-timestamp half — a pure function of the bytes the parse read — is stored. The mtime half is re-stat'ed and the max re-folded on every read. This is the field the display's whole staleness story hangs off; freezing it would have been the exact self-defeating fix. - **The sub-agent count is not cached.** It is a function of `now` as much as of the disk (§7.13's recency horizon): a fan-out expires with no file changing. It runs on every read, cache hit or not. It was always a stat pass and stays one. - **Diagnostics and degradation replay verbatim.** A hit reports the same torn records and the same drift verdict a fresh read would. `drift.Watch.Fold` only reads its watch, which is what makes replaying a stored one sound. - **Absence stays absence.** The stat happens *before* the cache lookup, so a deleted transcript is `ErrSessionGone` and never a replay; and a field the parse did not source is stored as empty and rebuilt as nil. - Entries are pruned in `Discover` against the live set, so a HUD left running for a day does not accumulate one per session that ever existed. The residual risk is every mtime cache's: a rewrite landing on byte-identical size *and* identical mtime is invisible. NTFS timestamps are 100 ns, the vendor appends rather than rewrites, and the cursor adapter already accepts this trade. **Rejected: skipping the read for sessions the idle filter will hide.** It is the biggest apparent win on the table — 1,235 of 1,404 rows — and it is a lie. Q8 exists precisely because NTFS defers mtime while a writer holds the file (observed lags of ~100 s hot, ~20 min closing), so a stat-only recency prefilter would hide the hot sessions the HUD exists to watch. The scan also cannot know the filter: `a` toggles `ShowAll` and the rows must already be there. The cache buys the same speed with none of that. **Result**, `BenchmarkScan` in `internal/hud` over a synthesized 1,400-session corpus (generated in the test, never committed; five runs each, same machine): | | before | after | |---|---|---| | warm scan (steady state) | 798 ms median (533–931) | **82 ms median (64–174)** | | cold scan (first, after launch) | 896 ms median (708–1156) | 994 ms median (655–1200) — unchanged within noise | On the live corpus the warm whole-`Scan` went from 1.84–3.37 s to **181–204 ms**. The cold scan is untouched by design: the first frame must genuinely read everything, and that is what the spinner is for. The benchmark's corpus is Claude-shaped only, which is a deliberate narrowing recorded here so nobody reads it as a whole-fleet figure: claude was 967 of 1,404 sessions and 2.6 s of the 2.9 s of read time, and it is the only adapter carrying this cache. The other five together were under a fifth of the budget. It runs at full scale in CI (~30 MB of temp files) and drops to 50 sessions under `-short`. Codex's 138 ms is the next-largest item and is now co-dominant with everything else put together. It is left alone: the scan is an order of magnitude inside its budget, and the same cache would need its own correctness argument against its own read path. ### 7.19 `w`: the week page — the slow windows, one line per vendor (2026-08-09) The owner's question, verbatim: "one view of the weekly usage for these models — it would help with scoping work." §7.17 already holds every reading that view needs and spends a block per vendor to say it, with the census, the spend lines and the five-hour windows in between. A scoping glance wants none of those; it wants the slow pools, one line each. So `w` opens a LENS over §7.17's data — the same `quotaVendors`/`usageBlocks` assembly, so the two surfaces cannot disagree about an account — rendered as a tight table. **Which windows, and why the rule is honest.** The page shows every window the vendor itself names weekly, plus the vendor's longest. Neither leg infers a duration, which is the constraint that shaped this section (§4a.1; quotaAgeWarn's ruling that a length parsed out of "5h" or "seven_day" is a guess wearing a fact's clothes): - the `-weekly` suffix is vendor vocabulary read verbatim — agy names its buckets ("3p-weekly", "gemini-weekly" observed, §3.8) and quotacache carries the names as ids unchanged. Reading the vendor's own suffix is reading, not translating. - the LAST window rides `model.QuotaWindow`'s ordering contract — "display order, shortest first" — so it is the vendor's longest pool by structure rather than by arithmetic. Claude's slice ends on `seven_day`, Codex's on `secondary`. No id is parsed for a length, and each row renders the vendor's own label beside the reading, so the page never states a duration the vendor did not. One edge is deliberate: a vendor whose only surviving window is short — Claude relayed after quotacache dropped an expired `seven_day` — shows that window under its own label. It is the longest reading telltale holds, and the label says how long it is. **What is kept off.** SPEND does not appear, and not because it is unimportant: a spend total's accumulation window is "sessions on disk, this scan" (§7.16) — not a week, not any calendar span — and rendering it under a page titled "this week" would claim a window the number does not have. The u page renders spend correctly, one key away. A vendor with no reading keeps §7.17's absence sentence rather than an em dash, because a dash would say "no reading now" about vendors that are structurally unreadable (§4a.1's three kinds of nothing). Sorting by remaining headroom was considered and rejected: readings, absences and two-window vendors do not order on one axis, and a page that reshuffles when a percentage moves spends §7.1 rule 4's churn budget to encode nothing the percentages do not already say. **The relayed age rides every row.** This page has no vendor headings to carry §7.15's age, so it rides each row as a suffix, and past quotaAgeWarn it escalates in the header's own grammar — the word (`stale`), the glyph, then the hue as the second signal. The REASON sentence does not travel here; §7.17 carries the argument, this page carries the alarm: ``` claude 7d ██▉───────────────── 15% ↻ 4d07h · ⚠ stale 19h ago ``` **The frame, generated by the build** (`internal/hud/testdata/golden/week.txt`; the week-stale variant is the golden quoted above). Both agy weekly pools under one vendor name, the second on a continuation row; the four five-hour buckets in the fixture render nowhere; 3p-weekly's measured 0% draws a full empty track while the three absence vendors draw sentences — zero and absent, still different states on the page built for a glance: ``` telltale │ 7 sessions │ codex 1 gemini 1 agy 3 cursor 1 grok 1 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── this week each vendor's longest window, and every window it names weekly ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ claude 7d █▏────────────────── 6% ↻ 5d00h · 2h ago codex 7d ███████████████───── 79% ↻ 22h48m gemini no quota reaches disk anywhere telltale can read agy 3p-weekly ──────────────────── 0% ↻ 6d23h gemini-weekly ███████▎──────────── 38% ↻ 6d23h cursor no quota anywhere · its store holds experiment values, not usage grok no quota anywhere · no window, no ordinal, no reset time on its disk ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── esc close ↑/↓ scroll ``` **Known limitations, named:** - **"This week" is the page's name, not each row's claim.** Codex's `secondary` window is selected for being the longest, not for being seven days; if the vendor ever reports a different length there, the label beside the reading says so and the page needs no code change. The title states the question the page answers, the labels state the facts. - **A new agy bucket that is weekly but not named `-weekly` would miss the suffix leg.** It would still appear if it is the slice's last window; otherwise it waits for the registry to learn the vendor's new word, which is the same posture every verbatim- vocabulary surface here takes (§7.15's convert rule). ### 7.20 `--hide`: the standing hide list (2026-08-10) The owner's request, near-verbatim: hide gemini and cursor, because only the CLI vendors are in use these days. The `v` filter cannot answer it — a filter narrows one launch to one vendor, and this is the opposite ask: every launch, every vendor EXCEPT two. So the HUD takes a hide list: `--hide gemini,cursor`, with the env var `TELLTALE_HUD_HIDE` as the flag's default. The env var is the standing per-machine preference (TELLTALE_ASCII's precedent) and the flag always wins, including `--hide ""` to see everything for one launch without unsetting the variable. **Where the hide is applied.** To the snapshot, as each scan lands — not in Render. The grid, the vendor lines, the fleet quota strip, the `u` page and the `w` page all read the same snapshot, so stripping it once is what keeps five surfaces from ever disagreeing about who is hidden. The header census comes from the same slices, so its counts match the rows for free. **How this survives the honesty rule.** A monitor that silently hides rows is a liar (§7.1), and a hidden vendor has no backstop at all: it leaves the "N of M" census entirely, and no keypress re-reads the choice. So the footer states `hidden gemini cursor` for the whole run, and the notice outranks the filter and the query in the drop order — only the two ⚠ facts sit above it. The `v` cycle skips hidden vendors, because a filter that can only ever select an empty grid is a dead stop on a one-key cycle; a `--vendor` naming a hidden vendor at startup is refused loudly rather than opened onto a contradiction. **What this is not.** Not an uninstall: the adapters stay registered, and the vocabulary (`agy`/`antigravity`, `cursor`/`composer`) is parseFilter's own, so the two flags cannot disagree about what a vendor is called. The list is deduplicated and sorted at parse time so the footer's wording is stable no matter how it was typed. `--hide all` is refused: a HUD told to hide every vendor is a request to not run it. ### 7.21 The event sink — every hook, one durable stream (2026-08-11) **What it is.** `telltale events` is a loopback-only HTTP server plus a durable log. An emitter POSTs one hook event to `/events`; the sink appends it to a JSONL day file under `~/.telltale/events/`, and rebroadcasts it to every WebSocket client on `/stream`. Two read endpoints serve a future viewer: `GET /events/recent?limit=N` (newest first) and `GET /events/filter-options` (the DISTINCT of the three tag axes: `source_app`, `session_id`, `hook_event_type`). The design is a re-implementation from a published observability shape (an emitter, one events table, a live stream); no code was taken from the reference, which carries no license. **Why it exists.** The fleet runs four vendors and each one's hooks fire into that vendor's own log, in that vendor's own shape. The sink is the one place a hook event from any vendor can land in one shape: any process that can pipe JSON is a source. `tools/emit-event.py` is the reference emitter (stdlib-only Python, no dependency to install): it reads the hook payload on stdin, promotes the fields a reader filters on (`tool_name`, `tool_use_id`, `error`, `agent_id`, `agent_type`, `stop_hook_active`), stamps epoch-millisecond time, and POSTs. Its hard rules are the hook contract: a quarter-second connect probe before the POST, a 5 second timeout on the POST itself, no retry, and exit 0 on every path — a sink that is down costs the agent the probe and one stderr line, never a failed turn. There is no summarization pass anywhere: the payload travels and is stored verbatim. **Trap 3 — "5 second timeout" never bounded the down path, and `localhost` was the wrong host (measured 2026-08-15, Windows 11, during a fleet-wide slow-hook audit).** A refused connect is not a timeout: Winsock retries it internally for ~2 seconds per address family before surfacing WinError 10061, so the original emitter — one POST, no probe — cost ~4.4s of blocked agent time on EVERY hook event while the sink was down (transcripts recorded p50 4.3–4.5s per PostToolUse across Read/Bash/Edit/Write). The `localhost` default doubled the damage and taxed the up path too: the sink binds `127.0.0.1` only, and Windows resolves `localhost` to `::1` first, so every event paid the full retry cycle against `::1` before trying the address the sink actually listens on. Hence the probe (a 0.25s bounded connect answers "is anyone listening"; measured: down sink now costs ~0.26s plus interpreter start) and hence the default URL naming `127.0.0.1`. **Distribution is one edit per repo.** The wiring pattern is a hook command of the shape `python3 /tools/emit-event.py --source-app ` — `--source-app` is the only per-repo change. Claude Code payloads carry `hook_event_name` and `session_id`, so no other flag is needed there; a wrapper for a vendor whose payload lacks the name passes `--event-type`. **Why the store is JSONL, not SQLite.** The reference keeps an `events` table with indexes on the three tag axes and the timestamp. This repo takes no dependency for a storage path (decisions/001), and `internal/sqlite` is a byte-level READER with no write path — so the contract the indexes serve is met with stdlib parts: one JSONL file per UTC day, the retention window held in memory, distinct-value sets kept beside it. The two queries the endpoints need — last N by arrival, DISTINCT of three columns — are exactly what that shape answers. The WebSocket is hand-rolled for the same reason (`internal/eventsink/ws.go`): the sink speaks one direction of one frame type, and that is a page of checked stdlib code, not a dependency. **Retention, which the reference does not have.** `--retain ` (default 30). The sweep runs at startup and then hourly: memory drops events past the window, and a day file is deleted only when its whole day is past it. The file-name pattern is affirmative (`YYYY-MM-DD.jsonl`), so the sweep can never delete a file it did not write. **The boundary this moves, said out loud.** This is the first content-bearing store under `~/.telltale/` — rows carry the hook payload verbatim, not numbers-and-keys. Three facts keep it inside the read/write contract: it is its own foreground mode the operator starts (the `otel` precedent — the gauges gain no write), it binds loopback only and refuses any other host at startup, and nothing in the gauges reads or renders these files. CLAUDE.md's boundary section names the exception. **Dark by design, and the v1 gate.** v1 is gate-held (§1) and this subsystem touches no gate surface: no council code, no HUD or statusline render, no new seat. The sink runs dark — events accrue and stream, and nothing in telltale displays them yet. A viewer is a later call site, not a re-plumb; §7.16's held-display precedent is the model. **Verified live, 2026-08-11, Windows 11.** A fake PreToolUse payload piped into `tools/emit-event.py` against a running `telltale events`: the POST returned the stored row, the row landed in `~/.telltale/events/.jsonl`, `GET /events/recent` returned it, `GET /events/filter-options` listed its three axes, and a WebSocket client connected to `/stream` received `{type:"initial"}` on connect and `{type:"event"}` on the insert. The sink-down path was exercised the same day: with no server listening, the emitter printed one stderr line and exited 0. **Retention verified 2026-08-11, same box.** The startup sweep deleted a synthetic `2026-06-30.jsonl` staged beside the live day file and logged `retention sweep deleted 1 day files`, and the same run left `2026-08-11.jsonl` byte-identical by SHA-256 — so the sweep drops a day past the 30-day window and keeps a day inside it. **Live drive, 2026-08-11/12.** The runs above used a piped fake payload. A real vendor-invoked firing is now recorded. On the Windows reference box (Claude Code 2.1.226→2.1.228), an interactive Claude Code session (v2.1.228) in the repo directory ran one Bash tool call. The project-local PostToolUse hook posted to the sink, and the sink stored row id 4 with `source_app` `"telltale"`, the real session id, and the verbatim PostToolUse payload (`tool_input`, `tool_response`, `transcript_path`, `cwd`). The sink had run since about 11:58 local; the row landed 2026-08-12T02:38Z (timestamp 1786502282870). The drive found two traps, and both are measurements, not readings of a document. **Trap 1 — headless print mode does not run PostToolUse hooks.** In `claude -p` print mode, PostToolUse hooks NEVER ran: not from the project's `.claude/settings.local.json` with workspace trust accepted, not passed through `--settings`, with and without a `matcher` key. The measurement is a breadcrumb test. Two hooks wrote a file by two different mechanisms (`cmd /c echo` to a file, and `python -c` writing a file), and zero files appeared. So the hooks were not invoked at all — this is not an invoked-and-failed hook. **Do not generalize this to the council gate's stream-json surface.** §9.8's probe measured PreToolUse firing over `--input-format stream-json` the same night. The two runtimes differ; measure each one. **Trap 2 — `uv run` in a hook command exits before the script runs.** A hook command of the form `uv run /tools/emit-event.py` exits 2 with the error ``No environment file found at: `.env` `` on any machine where `UV_ENV_FILE` is set globally. A PostToolUse hook failure is silent in normal use, so the hook looks wired and stores nothing. The working shape is a direct interpreter invocation. `tools/emit-event.py` is stdlib-only and needs no `uv`. The `telltale events` usage text now recommends the direct form. **Amended 2026-08-16 — a taken 4519, measured and then named.** §7.16a's amendment of the same day fixed this shape for `telltale otel grok`, and left the identical residue here. **Measured 2026-08-16**, Windows 11, `main` at `1995b34`, with a throwaway listener holding 127.0.0.1:4519: `telltale events` printed one line on stderr and exited 1. ``` telltale events: listen tcp 127.0.0.1:4519: bind: Only one usage of each socket address (protocol/network address/port) is normally permitted. ``` The failure was already loud and already correctly coded — nothing pretended to collect and nothing hung. What the line did not carry is what to do next, and there are three parts to that: who probably holds the port, that `--addr` moves this side (the flag already existed and was already documented), and that moving this side ALONE stores nothing. **The likely holder is a different one than the collector's, and that is the whole reason this message is not a copy of §7.16a's.** 4318 is OTLP/HTTP's registered port, so a collision there is very probably a rival receiver — Jaeger, the OpenTelemetry Collector, a vendor agent. 4519 is telltale's own and nothing else in this fleet claims it, so the only holder telltale can defend naming is **a `telltale events` sink the operator already started**. The message says that, and it says the cheap thing first: check for the running sink before moving the port, because if one is already listening the emitters already reach it and a second sink buys nothing. On a port the operator chose it names no holder at all — telltale cannot know who took 4600, and it says so rather than guessing. **What else must move, and why dropping that half is worse here than for the collector.** The emitters are the other side: `tools/emit-event.py` defaults to `http://127.0.0.1:4519/events` and takes `--server-url`, so a moved sink needs `--server-url http:///events` added to the hook command in **every** repo's `.claude/settings.json` — a second per-repo edit beside `--source-app`. A sink moved alone listens forever and stores nothing. §7.16a's version of that reads like "grok spent nothing"; this one is quieter still, because the emitter's hook contract (measured 2026-08-11, and hardened by Trap 3 on 2026-08-15) is a 0.25s probe, one stderr line and **exit 0 on every path**. A half-moved sink therefore produces no failure anywhere: the fleet simply looks like one where no hook ever fired. That is why the redirect is in the error text and in the flag help, not left to this document. **The redirect is measured here, not merely prescribed.** §7.16a had to mark its `OTEL_EXPORTER_OTLP_ENDPOINT` line unverified, because moving grok's own exporter needs a fresh instrumented capture. This side owns both halves, so there was no excuse to leave it at a prescription. **Measured 2026-08-16**, Windows 11, same build: `telltale events --addr 127.0.0.1:4520` bound and logged `listening on 127.0.0.1:4520`; a synthetic PreToolUse payload piped into `tools/emit-event.py --source-app tt-moved-port --server-url http://127.0.0.1:4520/events` exited 0 with no stderr; the sink logged `stored #1 tt-moved-port/sess-moved-live PreToolUse`; `GET /events/recent` returned the row with the payload verbatim; and `2026-08-16.jsonl` appeared in the store. The run used a redirected home directory, so it wrote to a temporary store and not to the operator's own. The suggested port is the failed port plus one, and telltale does not scan for a free one: a port free at the scan is not free at the bind, and a suggestion that looked verified would be the dishonest one (§4a.1). The loopback bind stays absolute across the flag — `TestAMovedPortIsStillLoopbackOnly` pins that, and it matters more here than for the collector, because these rows carry hook payloads verbatim rather than four token counts. `TestAHeldPortSaysWhoProbablyHasItAndWhatElseToMove` and `TestTheDefaultPortCollisionNamesASinkAlreadyRunning` pin both branches of the message, and `TestAMovedPortBindsAndStoresAnEvent` pins that the way out works. **`internal/bindaddr`, added by the same change.** The busy-port detection, the loopback test and the plus-one suggestion now live in one package that both §7.16a's collector and this sink call; the two messages stay in their own packages, because what a collision MEANS is per-mode and only the mechanism is shared. The extraction is not tidiness. The detection carries a measured Windows fact that the portable-looking version gets wrong — `errors.Is(err, syscall.EADDRINUSE)` is FALSE there for a real collision, because the bind returns errno 10048 (`WSAEADDRINUSE`) while Windows builds define `syscall.EADDRINUSE` as one of Go's synthetic `APPLICATION_ERROR` constants (536870914, measured go 1.26). Windows is the primary target (ADR-002), so a second copy of that check is a second place for a later reader to simplify it back to one arm and break the only platform CI runs. `TestARealCollisionIsDetectedOnThisPlatform` provokes a real collision rather than constructing an error value, so whichever arm a platform needs is the arm its suite exercises. **Amended 2026-08-16 — "it binds loopback only" was not containment.** Three facts were offered above as what keeps a verbatim content store inside the read/write contract, and the second of them was the load-bearing one. It did not hold against a browser. Measured the same day (§7.24): a page on another origin posted a forged event into this sink, and — the worse half — opened `ws://127.0.0.1:4519/stream` and was handed the `initial` snapshot, every retained hook payload verbatim. A WebSocket handshake is exempt from CORS, so no content-type rule or preflight was ever going to reach that path; the refusal had to move into the handler. Every endpoint now refuses a request carrying `Origin`, the stream included and **before** the upgrade, because a page that reaches `onopen` has already been handed the snapshot. `/events` additionally requires `Content-Type: application/json`, which is what `tools/emit-event.py` was measured sending. The containment sentence above should now be read as three facts plus a fourth: no gauge reads these files, the operator starts the mode, the bind is loopback, **and a web page is not a sender.** **Amended 2026-08-17 — the sink gets its first reader, and it is its own mode.** "Dark by design" above named a viewer as a later call site. This is that call site. `telltale events view` lists what the sink stored, filters it by the three tag axes or by day, and follows the store live. `internal/eventview` holds it. **It is its own foreground mode, and that is the whole reason it may exist.** The four facts the paragraph above just finished restating are what contains a verbatim content store, and the fourth of them is that no gauge reads these files. A reader wired into the HUD would have spent that one: hook payloads would land on a surface that redraws on every tick, in a process the operator did not start for this. A separate mode spends none of the four. `telltale snapshot` (§7.22) and `telltale otel grok` (§7.16a) set the precedent, and `TestNoGaugeReadsTheEventStore` asserts it over the transitive import graphs of `internal/hud`, `internal/statusline` and `internal/snapshot` rather than leaving it to a reviewer's memory. **It reads the day FILES, not the sink's own endpoints, on three grounds.** The sink serves `GET /events/recent`, `GET /events/filter-options` and `/stream`, and this reader uses none of them. - **Trust: §7.24 already settled it, and it settled it the other way round from the intuition.** That section measured a plain file write planting the same row a POST plants with MORE control over it, and stated the boundary once: a program running as a principal the store's ACL admits is trusted by these listeners exactly as far as the filesystem trusts it. So the HTTP path grants this reader nothing the file path does not. The endpoints are not closed to it either — `internal/localonly` refuses a request carrying `Origin`, and a local program sends none — which is the point: a viewer that connected would be served, and would gain nothing by it. - **Availability, which is what actually decides it.** The sink is a foreground mode the operator starts. Its endpoints answer only while that process is alive, and only over the window that process loaded at startup. The day files outlive it. A reader that needed a running sink would be dark in exactly the case it is reached for: after the fact. - **A checkable boundary rather than a promised one.** With no network call in the mode at all, "it makes no network call" is an import-graph fact. `TestTheViewerOpensNoSocket` asserts it against `go list`, and says in its own comment what it does not cover. **What it costs, and what the cost buys back.** Follow mode POLLS the day files on an interval instead of receiving a push, so the interval is the honest latency bound and `--interval` names it rather than the banner claiming "live". The trade is not one-sided: the sink's `broadcast` drops a subscriber whose buffer fills, by design, while a file tail cannot miss an event that way, because the file is the durable record and the tail only moves forward through it. One `Tailer` serves both the startup listing and the follow loop, so an event stored between the two can neither be missed nor printed twice — the seam a separate priming read would have had to choose a failure mode for. **Keys on the row, content behind a flag.** A row carries the arrival id, the stamp the emitter sent, the three tag axes, and the promoted fields (`tool_name`, `tool_use_id`, `agent_type`, `agent_id`, `stop_hook_active`). The payload prints only under `--payload`. So does the promoted `error`, and that split is the one judgement call worth recording: the other promoted fields are keys, while an error message is free text the hook was handed, which makes it content in the same sense the payload is. The row still prints the WORD `error`, because whether a row has one is what decides if the reader asks for the body. Note that `session_id` is a key HERE and content in §7.22: the snapshot renders no session id at all, because a gauge rollup has no use for one, while the three tag axes are this subsystem's whole filter surface and the sink already serves them at `/events/filter-options`. Two different contracts, each stated where it binds. **Plain text and no colour, following `doctor` rather than the TUI.** This output is read piped into a file and pasted into an issue by someone asking why a hook stored nothing, so every distinction it makes is carried by a word — which satisfies the colour rule by having no first signal that is not one. `--ascii` and `NO_COLOR` are therefore not flags on this mode. Nothing is truncated either: the widest column is a 36-character session id, and that is the field a reader carries into `--session`, into a vendor's own log, into an issue. A clipped one is a session nobody can correlate. **Zero and absent stay different here too** (§4a.1). A row with no timestamp renders the word `absent`, never `1970-01-01` — epoch zero through a date formatter produces a date that reads like a measurement. A `stop_hook_active` of `false` renders as a value and a missing one renders as nothing at all. Both are pinned: `TestAnAbsentTimestampIsNotNineteenSeventy` and `TestAMeasuredFalseIsNotAnAbsentField`. The partial-read rule holds as well: an unparseable line costs that line, is counted, and is reported on screen, so a store that is 40% unreadable cannot look like a quiet fleet. A half-written line is not unreadable and is not counted as such — `internal/jsonl` holds it back until its newline arrives, which matters because the sink appends while this reader is reading. `--day` selects a FILE, which is the day the SINK recorded the row, not the stamp the emitter sent. The two differ across UTC midnight and whenever a sender's clock is off. The file name is the fact this reader can check; the stamp is the sender's claim, and it is in its own column to be compared against. **What it deliberately does not do.** It writes nothing, and does not create the store directory — a reader that created it would make "the sink has never run here" unanswerable on the next run. `TestTheViewerWritesNothing` hashes the store before and after. It renders no gauge, adds no seat, and touches no council code, so it stays clear of the v1 gate (§1) for the same reason the sink itself did. And it does not summarize, group or count anything about the payloads: the sink stores them verbatim and this prints them verbatim or not at all. **Verified live, 2026-08-17, Windows 11**, built from the branch, with a redirected `USERPROFILE` so nothing touched the operator's own store. `telltale events` bound 127.0.0.1:4530; a synthesized PostToolUse payload piped into `tools/emit-event.py --source-app tt-live-probe --server-url http://127.0.0.1:4530/events` was logged as `stored #1 tt-live-probe/11111111-cccc-4ddd-8eee-000000000009 PostToolUse`. The sink process was then **stopped**, and `telltale events view --payload` still listed that row with its payload byte-identical to what the emitter sent. That last step is the availability argument, measured rather than asserted: with the sink gone, every endpoint this reader could have used was gone with it. Follow mode was driven the same day against a synthesized store: it printed the retained tail oldest-first, and a row appended to the day file appeared within one 500ms interval. ### 7.22 `telltale snapshot` — the read mode whose reader is a program (2026-08-11) **What it is.** `telltale snapshot` runs one scan and prints the fleet's current gauge state as a single JSON document on stdout, then exits 0. It is the same scan the HUD runs — `hud.Scan` plus the account relay — reshaped by `internal/snapshot` instead of rendered into a frame. Three flags: `--vendor` (one vendor only, the HUD's vocabulary), `--compact` (one line instead of indented), `--timeout` (default 10s). **Why it exists.** The fleet's own agents are now a reader of this data, and neither existing read surface serves them. The statusline answers one vendor in one line of styled text. The HUD is a full-screen TUI that runs until you quit it, and its output is a frame of box-drawing characters an agent would have to scrape. Both are built for eyes. An agent that wants to know whether anything is close to its context window, what the fleet has spent, or which vendor stopped reading, had no answer that was not a screen-scrape — so it either did not ask, or it read `~/.telltale/` directly and coupled itself to a cache format that is nobody's API. **A separate mode, for `doctor`'s reason.** What it prints goes somewhere else. The HUD owns the alternate screen; this writes to a pipe and returns. Nothing on this path enters the TUI, renders a gauge, or touches council — which is what keeps it clear of the v1 gate (§1). It is additive: no existing surface changes. **The schema, and the four rules that shape it.** The document is `{schema_version, generated_at, scan_error, fleet, vendors[]}`. - **Zero and absent stay different** (§4a.1). A measured zero is the number `0`. An absent value is `null`. No sentinel numbers, and **no `omitempty` anywhere**: an optional key is always present, carrying `null`. That last part is the sharper half — a key that vanishes when its value is absent makes "no reading" and "this schema moved under me" the same observation for the consumer, which is the zero-vs-absent collapse one level up. `internal/snapshot/testdata/golden/zero-vs-absent.json` pins it beside the HUD's golden of the same name, and `TestZeroIsANumberAndAbsentIsNull` asserts the two states differ in JSON *type*, not merely in bytes. - **A derived value says so.** Each vendor block carries `estimated`, the sorted list of `model.Field` names whose value here an adapter computed rather than read. It is the JSON form of the render layer's `~`. - **"Can't know" is not "absent now".** Each vendor block also carries `unsupported`, the fields that vendor exposes nothing for, ever. A `null` on a field named there is a capability statement; a `null` anywhere else is this moment's reading. The HUD spends a whole column-drop rule on this distinction; JSON gets it for two lists. Both lists cover the SESSION-sourced fields only, and the first live run is what settled that: an adapter's quota capability describes what a session exposes, while the `quota` array comes from the account relay, so listing quota in both put `agy` in the document with two relayed windows and the word `quota` under `unsupported` — two true statements that read as one contradiction. - **Definitive empty states.** A list with nothing in it is `[]`, never `null`. A reader must never handle two spellings of "nothing". **Pre-computed aggregates, because the alternative is every consumer doing the same arithmetic.** `fleet` carries the session count, the liveness census (`live`, `idle`, `stale`, and `unknown` as its own count rather than folded into `stale`, which is an age claim those rows cannot support), the vendor census by status, the highest context percentage anywhere, and the total cost. `context_pct_max` is a max and not a mean: the fleet question is "is anything close to its window", which an average over idle sessions hides. **Two things are deliberately absent.** There are **no per-session rows**. Partly that is the rollup being the product — an agent wants one answer per vendor, not a list to fold — and partly it is the read/write boundary: a session's honest identity is its name and its workspace path, and this surface renders numbers and keys, never content. `TestNoSessionContentReachesTheDocument` plants markers in every content-bearing field of a session and requires that none survives, including the session id. And there are **no token counts**. The relay is wired and the HUD reads it, but the DISPLAY is held by the owner (§7.16's amendment, applied to grok in §7.16a). A JSON field is a rendering; adding one here would end that hold as a side effect of a different feature. `TestSpendIsNotRendered` pins the omission so the day the hold lifts is a decision, not a drift. **Quota comes from the account relay and never from a row** (§7.15, §7.1's sixth rule). Hanging a window off a session would assert a per-session limit no vendor publishes. `quota_read_at` travels with it, because a quota figure without the age of its reading is a number the consumer cannot judge. **It writes nothing.** The gauges' contract holds on this path with one item spare: it reads vendor stores and the quota relay, calls no network, reads no credential, and does not even write the quota relay — it renders no quota of its own to relay. **Verified live, 2026-08-11, Windows 11.** `telltale snapshot` against the reference box's real stores returned a document carrying all six adapters, 1423 sessions with a liveness census that summed to that count, `agy`'s two relayed quota windows with their reading time, and `estimated: ["subagents"]` on claude against `estimated: []` on cursor. The zero-vs-absent pair appeared in that real document without being staged: `agy`'s `3p-weekly` window carried `"used_pct": 0` — a measured zero — beside five vendors whose `quota_read_at` was `null`. `--compact` returned the same document on one line and `--vendor codex` returned that vendor alone. `--json`, a positional argument, `--vendor chatgpt` and `--timeout 0` each printed a corrective error, no document, and exit 1. That run is also what found the quota-capability contradiction described above; the contradiction was fixed and the run repeated. **The stated reader has now consumed it, 2026-08-12.** At 2026-08-12T01:59Z an agent session ran `telltale snapshot --compact` (binary built at main `01770ec`) and answered real fleet questions from the parsed JSON: 6 vendors watching, 1,443 sessions, 1 live, `context_pct_max` 75.8 on codex, and `agy` quota with 4 windows of which `gemini-weekly` carried 11.9 `used_pct`. Zero-vs-absent held in the document the agent read: `cost` was `null` everywhere, and a `used_pct` of 0 was the number 0. **Amended 2026-08-16: the contract is published, and CI is its first measured consumer.** Everything above was asserted by this package's own tests. A schema that only its author validates against is a schema nobody has tested, so the contract now lives in a file a consumer can fetch, and the gate reads that file rather than a second copy of its rules. - **`docs/snapshot.schema.json`** is the document's JSON Schema, draft 2020-12. It documents what the shipped binary emits, measured against the real output of a live run and against the four goldens. Where the code and an intention differ, the code wins. - **The `test` job validates the BUILT binary's `snapshot --compact` output against it**, then validates the four golden documents too. The binary's own document runs on a machine with no vendor CLIs, so the goldens are what cover the shapes that machine cannot produce: a vendor the operating system refused with its message, a drifted store, a scan error, and the zero-vs-absent pair. The step also re-asserts the "writes nothing" half, as a before/after of `~/.telltale`. It runs after the quota relay smokes, so the runner's document carries relayed quota. Measured 2026-08-16 against that exact sequence: `agy` arrives with two windows and a `quota_read_at`, one of them a `used_pct` of 0 — a measured zero on a bare runner, unstaged — while `claude` arrives with `[]` and a null read time, because full.json's reset stamps are absolute and now in the past. The step asserts that at least one vendor still carries a relayed window, so the day that stops being true is a red build and not a quietly narrower gate. - **The gate is proved non-vacuous on every run.** Three mutations break the real document one way the contract forbids, and the validator must reject all three: a dropped optional key (the `omitempty` regression), the string `"n/a"` where a nullable number belongs (the sentinel regression), and a `schema_version` bump nobody wrote a schema for. A gate that has never failed is a gate nobody has measured. - **The validator is pinned Python, not a Go module.** `go.mod` carries no direct dependency outside the TUI stack, and this document records that choice four times over — the SQLite reader (§3.2), the zstd reader, the OTLP listener (§7.16a) and the event emitter (§7.21) each refused a library. A schema module used only by a test would still enter the shipped module's graph. Hand-written Go assertions were the third option and are not a schema gate: they restate the contract in a second place, and two statements of one contract drift. Run the gate the way CI runs it: ``` python -m pip install jsonschema==4.23.0 go build -o telltale.exe ./cmd/telltale ./telltale.exe snapshot --compact | python tools/validate-snapshot.py - python tools/validate-snapshot.py internal/snapshot/testdata/golden/*.json python tools/validate-snapshot.py --mutate drop-key ``` **What the schema deliberately does not promise.** Each of these is a limit of the format rather than an omission, and a consumer that reads the schema as a stronger claim will be wrong. - **It does not promise which kind of `null` you got.** The schema guarantees that a measured zero is the number `0` and an absent value is `null`, because the two have different JSON types. It cannot tell a reader that the `0` in front of it was measured. `estimated` and `unsupported` are what carry that, and `TestZeroIsANumberAndAbsentIsNull` is what keeps the emitter honest about it. - **It does not bound any value.** `context_pct_max` has no maximum and `cost_usd_total` has no minimum, because nothing in `internal/snapshot` enforces one. A bound the code does not hold is a claim the schema cannot back (§4a.1). - **It does not close the object.** Every object sets `additionalProperties: true`. That is the schema agreeing with `SchemaVersion`'s own rule: a field added at the end is not a break, because every reader parses by name. `additionalProperties: false` would make the schema call a break what the emitter does not, and would redden a consumer's pipeline on a release that added a field. `required` is what carries the weight instead — a renamed or dropped key still fails, which is the regression that matters. - **It does not enumerate the vendors.** A seventh adapter adds a value, not a field. `status` IS enumerated, because those four words are the whole of `hud.VendorStatus` and a fifth would be a new state a reader must handle. - **It does not promise the order of `vendors`.** The entries arrive sorted by vendor id today. Nothing asserts that, so the schema claims nothing about it. **Amended 2026-08-16: a human-visible consumer ships beside the machine one.** CI validating the document proves the shape holds. It does not show anyone what the document is for, and a contract whose only consumer is its own gate is a contract nobody has used. `tools/fleet-prompt.ps1` is that second consumer: one PowerShell function, one `snapshot --compact` call, one parse, one line of prompt text. - **What it renders.** The vendor count that is watching, the session and live counts, any vendor whose status is not `watching` named with that status, the highest context percentage with the vendor holding it, the busiest relayed quota window, and the words `scan degraded` when `scan_error` is not null. - **What it demonstrates, which is the reason it exists.** The two rules the schema states survive the trip into a caller. A null drops its whole segment rather than printing 0 or a dash, which in PowerShell means `if ($null -eq $v)` and never `if (-not $v)` — `-not 0` is true, so the idiomatic spelling is exactly the one that erases a measured zero. A figure whose vendor lists `context_pct` in `estimated` keeps the render layer's `~`. An unknown `schema_version` returns an empty string, because a prompt segment that guesses at a contract it does not know is worse than one that is absent for a release. - **Windows PowerShell 5.1, not 7.** That is what a Windows 11 box has before anyone installs anything, and this is the primary platform (ADR-002). No ternary, no null-coalescing, no `ConvertFrom-Json -Depth` and no `-AsHashtable`. Output is ASCII and holds no ANSI escapes; colour is the caller's, and the line has to read on a console that has none. - **`scan_error` renders as two words and never as its own text.** The message can be long enough to break a prompt, and it can name a path. A prompt segment is the wrong surface for a diagnostic string. **Driven 2026-08-16, Windows 11, PowerShell 5.1.26100.9168**, against the real stores at branch `snapshot-example` off main `08c300f`. The live document (1554 sessions, six vendors watching, `context_pct_max` 75.8 on codex whose `estimated` holds `context_pct`, and agy's four relayed windows) rendered: ``` tt 6 watching | 1554 sessions, 3 live | ctx ~75.8% codex | quota 12.2% agy/gemini-weekly ``` The four goldens drove the shapes a healthy fleet cannot produce, through the same `-FromFile` path. `watching.json` produced `attn 1: codex unreadable` beside `ctx ~61.3% claude`; `drifted.json` produced `attn 1: claude drifted` and `scan degraded`; `empty-fleet.json` produced counts and nothing else, no `ctx` segment at all, because every reading in it is null. `zero-vs-absent.json` is the pair that matters and it held: its `context_pct_max` of 0 rendered `ctx ~0%`, a printed zero, against empty-fleet's null rendering as no segment whatsoever. A `schema_version` of 2 returned the empty string. The live document is quoted in the pull request unredacted, and no redaction was needed: the surface renders numbers and keys, so the real output carries vendor ids, counts, percentages and timestamps and no session name or workspace path. That is the "never content" claim above, observed rather than asserted. ### 7.23 The drop-file relay — a row telltale did not measure, and says so (2026-08-16) §4 promises a documented adapter interface, and §4a.7 works one example through. That promise answers a contributor who will write Go. It does not answer the other question the launch post raised: what happens to a tool telltale ships no adapter for, and is not going to? The owner ruled out a plugin runtime and lifted this feature's demand gate on 2026-08-16. A runtime would execute a stranger's code inside a process whose entire contract is "reads, never writes, no network, no credentials" — every guarantee in this document would become a guarantee about someone else's plugin. The middle path is a **drop file**: the vendor, or a script the user writes, puts one small JSON document under `~/.telltale/dropfile/.json`, and telltale renders it as a fleet row. The reader opens one file and can do nothing else. `docs/dropfile.md` is the spec; this section is the reasoning. **It adds no write exception.** The three sanctioned writes (§7.15, §7.16, and council's `room.json`) are untouched. `internal/adapter/dropfile` creates nothing, writes nothing and removes nothing; the directory is the operator's to fill, and a missing one reports the vendor absent. The read/write boundary in `CLAUDE.md` needed no fourth bullet, which is the strongest evidence this was the right shape. #### The problem this format has and no other adapter has Every other adapter reads a store its vendor wrote while doing its own work. Nobody had a motive to write a flattering number into it, and each adapter's package doc names the live corpus every field was grepped out of. A drop file is written FOR telltale, by whoever, and every value in it is that writer's claim. **telltale measured that a file exists, when it was last written, and what it says. It did not measure the session.** So §4a.1's rule bites in a new place. The rule has always been about a *value* — measured, inferred, or absent. Here the whole ROW has a provenance, and a drop-file row must never be readable as a measured one. #### Why the `~` estimate marker is the wrong mark The obvious move is to mark every claimed value with the estimate marker and be done. It is wrong twice over, and the second reason is the one that settles it. It states a falsehood about mechanism. `~` means `CapDerived`: the adapter computed the value from something that is not the value, and the snapshot's `estimated` array says so in exactly those words. This adapter computes nothing — it reads the number the writer wrote, verbatim. Marking these rows `~` would claim telltale did arithmetic it did not do. And it collapses two provenances into one spelling. "telltale inferred this" and "somebody asserted this" are different claims that a reader would discount differently, and after the collapse no reader could tell which one they had. That is the failure §4a.1 exists to prevent, imported through a new door — the same shape as the zero-versus-absent collapse, one level up. The rejected alternative is recorded rather than merely dismissed: a per-cell mark of any kind, `~` or a dedicated glyph. Beyond the two objections above it repeats one fact in every cell of a row where the fact is uniform, and §4a.2 already refused a per-field glyph for degradation on the narrower ground that "we failed to read it" starts to read as data. **Provenance is a property of the row, so it is marked once, on the row.** #### The three marks, and none of them is a colour 1. **The vendor id is `self-reported` for every drop file.** The grid's identity column reads `SR` and the header census reads `self-reported 2` in full. `SR` is the only tag in `vendorTag` that is not a vendor abbreviation, deliberately: every other tag answers "which tool did this row come from" because telltale measured that tool's store, and here the only answer telltale can give is where the numbers came from. 2. **The claimed tool leads the row's label** — `windsurf: refactor the parser`. `SR` is shared by every drop file and cannot separate windsurf from aider, so the row itself names its claimant rather than hiding it in a pane nobody opens. 3. **`telltale snapshot` emits `"self_reported": true`** on the vendor entry, beside and never inside `estimated`. Both HUD marks are plain ASCII words, so `--ascii` and `NO_COLOR` change neither, and `testdata/golden/self-reported-row.txt` — which renders through `PlainStyles` — is the whole assertion rather than half of it. The drop-file vendor gets **no hue of its own**: a hue is a vendor identity (§9.28) and these rows have none to give, so it takes the identity-hue fallback as Gemini does. #### Impersonation is unrepresentable, not merely rejected The format has **no field for a vendor id**. A drop file cannot claim to be Claude Code, because there is nowhere in the document to make the claim — the id is a constant the adapter owns. Nor can it choose its own row: the session id comes from the FILE NAME, so the operator's filesystem decides what a row is called and a document cannot rename itself onto another one's row. That is the same "the allowlist is the struct" mechanism `internal/cursorhook` uses against a payload carrying reply text and an email address, and it is why the planted-credential test can assert absence rather than sanitization: a key with no destination is dropped by `encoding/json` before this package sees it. #### One vendor id, not one per tool Per-tool ids were considered and rejected. They would put a writer-chosen string in the column that says what telltale measured — a file naming itself `claude` would draw a row indistinguishable from a measured one — and `Capabilities()` is per-adapter, so a set of tools sharing one adapter could not honestly declare different capability sets anyway. With one id, `Capabilities` describes the FORMAT, which for this adapter is the only source there is: `unsupported` names the fields the format cannot express, and a file that omits `cost_usd` yields "absent now" rather than "can't know". Both statements are true, which is what the merge costs and what it buys. #### The one claim telltale can check, it checks A writer that stops writing leaves a file that keeps asserting whatever it last said. A file claiming `last_activity` of "now" would render live forever over a dead session, and `--vendor self-reported` would become the way to pin a row to the top of the grid. telltale cannot check a cost or a context percentage against anything. It CAN check `last_activity`, because **a file cannot have activity newer than its own last write**, and the mtime is telltale's own measurement rather than the writer's claim. So a claim ahead of the mtime is replaced by the mtime, and the substitution goes in `Diagnostics`. The value stays present and is NOT marked `Degraded` — a measured mtime is a better answer than absence, and §4a.2 requires degraded fields to be absent. Staleness proper reuses `internal/quotacache`'s rule and its constants rather than inventing a boundary: 24 hours to expire, five minutes of future-skew tolerance, and the reading's age travelling with it past five minutes. A file outside those bounds draws no row at all, which is that package's own ruling that the honest display for "no reading" is absence. #### Absence has two spellings on input and one on output §7.22 emits every key with an explicit `null`, because a reader parsing a document must not have to tell a missing key from a changed schema. An INPUT cannot hold its writer to that: a key omitted and a key written `null` both have to mean absence, or the format fails documents over fields the writer simply had no value for — which is §4a.5's partial-read rule broken at the door. So the semantic convention is snapshot's, exactly: **zero is a number, absence is nil, and no sentinel number stands for either.** Only the syntax differs, two accepted spellings in rather than one guaranteed spelling out. `TestZeroIsAMeasurementAndAbsentIsNil` walks all three cases, and `testdata/golden/self-reported-row.txt` pins the render: a claimed 0% draws a full empty track beside a claimed `$0.00`, and a row claiming neither draws the absent marker in both cells. #### The render, generated ``` telltale │ 3 sessions │ claude 1 self-reported 2 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── SESSION MODEL CONTEXT COST AGE ● CC │ telltale C:\src\code Opus 5 ███▊──────── 34% $1.20 │ 12s ● SR │ windsurf: refactor the parser C:\src\code gpt-5-codex ──────────── 0% $0.00 │ 40s ◐ SR │ aider C:\src\work — — │ 9m ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── q quit / find enter detail u usage w week v vendor s sort a all ? keys ``` The header census says the word in full and the identity column echoes it per row. The last two rows are the zero-versus-absent pair carried inside the claimed rows: windsurf claims a measured 0% and $0.00 and draws a full empty track beside a printed zero, and aider claims neither and draws the absent marker in both cells. This golden renders through `PlainStyles`, so every mark visible here survives `--ascii` and `NO_COLOR`. #### What the format cannot express, and why each door was never cut - **Quota.** An account property (§7.1), sourced from the statusline relay (§7.15). A session-shaped document has no account to speak for. The usage page says "no quota by design" for this vendor rather than borrowing one of the six sentences that report a seam which came up empty — this is a door never cut, not a seam that returned nothing. - **Liveness.** §4a.4 allows a hint only from a positive vendor signal the HUD cannot see. This is the field where a false claim is most tempting and least checkable, so there is no field for it. Liveness is classified by the HUD from `last_activity`, bounded by the mtime clamp above. - **Token counts.** They feed §7.17's fleet spend sum. A claimed count summed beside measured ones would yield a total carrying no mark at all. #### No canary, no pin, no drift watch `internal/adapter/pins` gets no row, and `internal/adapter/drift` is not wired in. Both exist because an adapter reads a private, undocumented format that a vendor may move without saying so, and the pin names the build the field map was surveyed at. **This format is telltale's own and it is published**, so there is no vendor build to pin and no shape to watch drift away from. `schema_version` does the job a canary does elsewhere, and it does it better: a document whose contract number this adapter does not speak is skipped whole rather than read leniently, because guessing that the field names still mean what they meant is inventing every value at once. #### Schema version `self_reported` is a field ADDED to §7.22's vendor object, so `SchemaVersion` stays 1 under that document's own rule: a field added is not a break, because every reader parses by name, and every object sets `additionalProperties: true`. Nothing already in the document changed meaning and no key left. `docs/snapshot.schema.json` carries the key in `properties` and in `required` — the emitter always emits it — and `tools/validate-snapshot.py` passed against the four goldens and against the built binary's live output. ### 7.24 Who may push to a loopback listener (2026-08-16) **The question.** §7.16a and §7.21 each open a listening socket, and each treats the loopback bind as the thing that contains it. Both write a store the product treats as measured: `telltale otel grok` folds a push into `usage/grok.json`, which is the file the HUD reads as grok's measured spend, and `telltale events` stores hook payloads verbatim. Neither listener asked who was pushing. In a product whose whole claim is that a displayed value came from measured vendor output, an unauthenticated input to that value is a real seam, so it was measured rather than argued about. **The measurement.** Windows 11, telltale built from `main` at `65a113c`, go 1.26, Chrome 151.0.0.0 headless. Every run used a redirected `USERPROFILE`, so nothing touched the operator's own stores. The collector ran on 127.0.0.1:41318 and the sink on 127.0.0.1:41519; the browser probe page was served from a separate origin, `http://127.0.0.1:41999`. *A local program forges a measured total.* A stdlib Python script built an `ExportLogsServiceRequest` carrying one `grok_code.api_request` record and POSTed it to `/v1/logs`. The collector answered 200, logged `counted api request`, and wrote: ``` {"vendor":"grok","since":"2026-08-16T20:21:11.601084-04:00","written_at":"…", "requests":1,"input_tokens":999000,"output_tokens":888000, "cache_read_tokens":777000,"reasoning_tokens":666000} ``` Nothing distinguishes that from a real export. **The forger needs no secret**: the port, the path and the record shape are all published here and in the package docs, and this repo is public. *A plain file write forges the same total, and more of it.* A hand-written `grok.json` claiming 4242 requests and 111111111 input tokens, with a `since` six hours back, was accepted whole by `readEntry` — the same function the gauge read path `ReadAll` uses. The proof is that a later relayed request accumulated **onto** it (`requests` 4242 → 4243) and kept the hand-written `since`. So the file write is not merely equal to the POST, it is **stronger**: a POST may only add four non-negative counts to whatever window is open, while the file writer picks the window's start, its request count and every total outright. The sink's store measured the same way — a line appended straight to `.jsonl` was served by `GET /events/recent` after the sink reloaded. *A web page forges into both stores, with no local code at all.* This is the result that decided the design. A page on `http://127.0.0.1:41999`, driven by a real headless Chrome, posted a forged OTLP record into `usage/grok.json` and a forged hook event into the sink. It works because neither handler read `Content-Type`: a `text/plain` body is one of the three CORS-safelisted media types, so the request is a *simple* one the browser sends outright, with no preflight to refuse. The response is unreadable to the page and that changes nothing — the write has already happened. *A web page reads the verbatim store.* Worse than the write, and it is the sink's, not the collector's. A WebSocket handshake is exempt from CORS entirely, and `upgrade` never looked at the request. The same page opened `ws://127.0.0.1:41519/stream` and was handed the `initial` snapshot — the last hundred stored events, hook payloads and all. §7.21's containment claim was "it binds loopback only"; against a browser that claim was doing no work. **The refuted argument, and the part of it worth keeping.** The tempting reading of the first two results is that the seam does not matter: any local program can write `usage/grok.json` directly, so a secret on the HTTP path buys nothing, and the honest fix is a documented boundary. **The measurement refutes that as stated, and confirms half of it.** The two paths do not have the same senders. A web page reaches the socket and cannot reach the file — it is not a program on this machine at all. The file's reach is bounded by an ACL (on the reference box `C:\Users\sanle\.telltale\usage` grants Full to the owner, SYSTEM and Administrators, and Modify to two further profile-inherited principals); the socket's reach, before this change, included every page the operator visited. So the boundary was not a boundary yet. The half that survives is the local half: a program running as a principal the ACL admits is trusted completely, and no HTTP-path secret would change that by one byte. **What was built: make the socket's senders equal to the file's.** The fix is not authentication and does not pretend to be. It removes the one class of sender the filesystem already excludes, and then states the rest of the boundary plainly. `internal/localonly` carries the check, extracted the way `internal/bindaddr` was for the same two modes. - **Refuse any request carrying `Origin`.** This is the arm that generalizes and the only arm that can cover the WebSocket handshake. Measured: every browser request carried `Origin`, including the handshake — which carried **no** `Sec-Fetch-*` header at all, which is exactly why the check reads `Origin` and not `Sec-Fetch-Site`. Neither real sender carried it: `tools/emit-event.py` arrived as `Python-urllib/3.14`, `Content-Type: application/json`, no `Origin`; the exporter-shaped request the same way with `application/x-protobuf`. A page cannot suppress the header, because the user agent attaches it rather than the script. - **Require the media type the measured sender sends** — `application/x-protobuf` for the collector (§7.16a's capture pins grok's own exporter to it), `application/json` for the sink (`tools/emit-event.py`). Parameters are ignored, so `; charset=utf-8` passes. Measured from the other side: with a non-simple `Content-Type` Chrome sent only an `OPTIONS` preflight to each endpoint and **no POST followed**, because neither server answers a preflight with CORS headers. - **Both arms, because they fail in different directions.** `Origin` names the sender class but rests on a header a future browser could stop sending on some path nobody has measured; the media type rests on nothing about browsers at all, and turns any such page into one that must preflight. Neither is load-bearing alone. The sink applies the `Origin` arm to its **read** paths and its stream as well, not only the POST: a page that cannot plant a row can still ask for the rows already there, and those rows are content. A refusal is `403` — not `401`, because no credential would help, and a 4xx rather than a 5xx so an OTLP exporter stops retrying instead of looping against a door that will not open. **No token, and the reason is the measurement, not the effort.** A shared secret was the obvious shape and it was refused three times over. It closes nothing against the principal that matters, because a program that can read a token file beside the store can write the store directly — measured above, and pinned by `TestTheFileWriterSetsWhatTheRelayCannot`. It would put a secret on disk next to the thing it protects, for a reader that already has the disk. And it would gate the collector's only working path on an **unmeasured** knob: §7.16a already had to mark `OTEL_EXPORTER_OTLP_ENDPOINT` unverified because moving grok's Rust exporter needs a fresh instrumented capture, and `OTEL_EXPORTER_OTLP_HEADERS` is the same instrument problem. A wrong guess there makes the collector count nothing while looking healthy, which is §7.7's worst failure. **OS-level peer verification was refused too**: loopback TCP peer identity on Windows needs `GetExtendedTcpTable`, which is not stdlib (decisions/001, the same rule that hand-rolled the OTLP and SQLite readers), and it does not even answer the browser case — the peer there is `chrome.exe`, a perfectly legitimate local program. An executable allowlist is a different and worse contract. **The trust statement, stated once so it can be quoted.** A program running on this machine as a principal `~/.telltale/`'s ACL admits is trusted by these listeners exactly as far as it is trusted by the filesystem, because it can plant the same row either way and the file write is the stronger of the two. That is a deliberate boundary, not an oversight, and `internal/usagecache/trust_test.go` pins it so a later session adding a bearer token walks past the reason it buys nothing. What is **not** trusted, as of this change, is a web page. **Verified after the change, same box, same probes.** The measurement is only worth what the re-run says, so the identical pages were driven at the rebuilt binary. The collector logged `refusing a request from a web page (Origin: http://127.0.0.1:41999)` and counted nothing; the sink logged the same refusal twice, once for the POST and once for the stream handshake; and the WebSocket probe, which had reported `EXFILTRATED: {"type":"initial",…}` before the change, now reports `BLOCKED (error)` and `CLOSED code=1006` with no snapshot delivered. The other half held too: an exporter-shaped push was still counted (`counted api request — in 42 · out 42 · reasoning 42 · cache read 42`) and `tools/emit-event.py` still stored a row and exited 0. **What this does not claim.** The capture is one machine, one day, one browser engine (Chromium 151). Firefox and Safari were not driven; the `Origin` behaviour they rest on is the Fetch and RFC 6455 requirement rather than a measurement here, and §3.4's discipline applies before extending the claim. Nothing here defends against a local program, and it is not meant to. The gauges' no-network rule and the loopback-only bind are untouched and remain absolute: this change only narrows who may talk to a socket that was already loopback. ### 7.25 `telltale mcp` — the same document, in front of an agent (2026-08-18) **What it is.** `telltale mcp` serves [§7.22](#s7-22)'s snapshot document over the Model Context Protocol on stdio. An MCP client starts the process, speaks JSON-RPC down its stdin and reads frames off its stdout, and calls one tool — `fleet_snapshot` — which runs one scan and returns the document. One flag: `--timeout` (default 10s), which is the deadline for each CALL rather than for the process. Nobody types this command; it goes in a client's config once, e.g. `claude mcp add telltale -- \telltale.exe mcp`. That spelling is the CLI's own — read off `claude mcp add --help` at Claude Code 2.1.233 on 2026-08-18, where `claude mcp add my-server -- my-command --some-flag arg1` is the documented stdio form — and not a shape assumed from memory. Running it is the operator's to do; see the closing paragraph. **Why it exists, given that §7.22 already shipped.** The snapshot mode answered "a reader that is a program" and it answered it well: CI validates it, and `tools/fleet-prompt.ps1` renders it. Both of those readers are SCRIPTS somebody wrote on purpose. The reader this repo actually has most of is an agent, and the honest difference is not that an agent cannot shell out — several can — it is that a command nothing told it about is a command it does not run. A tool its client LISTS is the version it reaches without being asked, with the argument names and the value rules in front of it. That is the same gap §7.22 described one layer up: the data was honest and reachable, and nothing was reaching it. **It is a fourth READER, not a second document.** `internal/mcpserver` calls `snapshot.Encode` on the document `snapshot.Build` produced from the scan `cmd/telltale` runs for `telltale snapshot` — one scan path (`scanDocument`), one adapter roster (`allAdapters`), one vendor vocabulary (`snapshotAdapters`), one serializer. The tool result's text content is byte-for-byte what the CLI prints. That is the whole design, and `TestTheToolResultIsTheSnapshotDocumentUnchanged` is what holds it: every honesty property this surface claims is a property of those bytes — a measured zero is `0`, an absent reading is `null`, no optional key is omitted, `estimated` names what an adapter computed, `unsupported` names what a vendor can never source, `self_reported` names an entry whose writer claimed it. A second serializer here would be a second statement of that contract, and two statements of one contract drift — which is §7.22's own argument for a published schema over hand-written assertions, applied to itself. `structuredContent` carries the same document as JSON beside the text, for a client that reads it that way. It is the identical value marshalled twice by one package, so the two can never disagree. No `outputSchema` is declared beside it, deliberately: the document's schema is published at `docs/snapshot.schema.json` and CI validates the shipped binary against that file, and embedding a copy in the binary would be exactly the second statement just refused. **One tool, not a family.** Every fleet question — what is close to its window, what has been spent, which vendor stopped reading, how old is the quota reading — is already one parse of one document. A second tool answering a subset would have to re-serialize part of it, and that is where the zero-vs-absent rules get restated ([§4a.1](#s4a-1)). **The tool DESCRIPTION carries the honesty rules, because the model reads it.** A caller that does not know them reads a `null` as a zero and an estimate as a measurement — the collapse this repo exists to prevent, moved one process outward into the agent. So the description says, in the text the model sees before it decides to call: null is absent and never 0, `estimated` means telltale computed it, `unsupported` means the vendor can never report it, and `self_reported` means its writer claimed it. `TestTheToolListNamesTheOneTool` pins all four words. **Stdio only, and that is what makes [§7.24](#s7-24) not apply.** telltale's two other machine-facing surfaces listen on loopback, and §7.24 exists because a loopback bind is not containment on its own — a measured headless Chrome planted a usage row and read the whole event store. This mode binds nothing. The client owns both pipes and starts the process, so there is no third party to refuse and no `Origin` to check. `TestTheServerOpensNoSocket` asserts the direct imports rather than the transitive graph, and says so: the graph already reaches `net` through `internal/hud`'s TUI framework, and linking that code is not calling it ([ADR-002](#adr-002)'s distinction). **It writes nothing**, with the gauges' contract one item spare for §7.22's reason: it renders no quota of its own to relay. `TestTheServerWritesNothing` drives a whole session with the home directory redirected and compares the tree before and after. **The protocol surface is four methods and stops there.** `initialize`, `tools/list`, `tools/call`, `ping`, plus the notifications a client sends and expects no answer to. Absent: resources, prompts, completion, sampling, subscriptions, and any server-initiated request. Their absence is *stated* in the capabilities object rather than discovered by a client that tried one. Two shapes are refused with the reason rather than half-answered: a JSON-RPC batch array (removed from MCP in 2025-06-18, and this server answers revisions on both sides of that), and a message with an id and no method, which is a response to a request this server never sent and would be a protocol error to answer. **Version negotiation echoes what the client asked for, within a list.** `supportedVersions` is `2024-11-05`, `2025-03-26`, `2025-06-18`, `2025-11-25`; the narrow surface above is spelled identically across all four, so echoing is a true statement rather than a compatibility guess. An unknown request gets `2025-11-25`, the latest supported, which is the lifecycle's own rule. **`2026-07-28` is deliberately not on the list, and it is the newest revision, so the omission is the interesting one**: that revision requires a server to implement `server/discover`, and this one does not. Claiming the version would be claiming a method a client is entitled to call — ADR-001's failure in protocol form, a capability asserted rather than built. The same revision's versioning section says a client may invoke methods inline instead, which is the path that works here. **Two error channels, and the split is about who can fix it.** A bad `vendor` argument and a scan that failed come back as a tool RESULT with `isError` set, because the model asked and the model can correct. A tool name this server never listed, an unknown method and a malformed line come back as JSON-RPC errors, because those are the client's plumbing. A failed call carries no `structuredContent` at all — there is no measurement behind it, and an empty document would be a fleet with nothing in it, which is a different claim. **Hand-written JSON-RPC over `encoding/json`, not the official MCP Go SDK.** This follows the module's standing position rather than inventing one: `go.mod` carries no direct dependency outside the TUI stack, and this document records the same refusal for the SQLite reader ([§3.2](#s3-2)), the zstd reader, the OTLP listener ([§7.16a](#s7-16a)) and the event emitter ([§7.21](#s7-21)). Four methods and one tool is a smaller surface than the SDK's own API. **Verified live, 2026-08-18, Windows 11, against the built binary** (branch `mcp-server` off main `4b58258`, `go build -o telltale.exe ./cmd/telltale`), driven by a scripted stdio client that writes request lines and reads response lines. Seven messages in, six responses out — the seventh was the notification, which is answered by silence — exit 0, empty stderr, ~2.3s for the whole session including one scan of the real stores: - `initialize` at `protocolVersion: "2025-06-18"` echoed that version and returned `capabilities: {"tools":{}}` with `serverInfo`; `notifications/initialized` produced no response at all, which is the half a client would hang on. - `tools/list` returned the one tool. - `tools/call fleet_snapshot` returned a document of 8 vendors and 1,523 sessions, and `structuredContent` parsed EQUAL to the text content. **The zero-vs-absent pair appeared in it unstaged**: `agy`'s `3p-weekly` window carried `used_pct` as the integer `0` — a measured zero — beside `gemini-weekly` at `0.2`, while `cost_usd_total` was `null` fleet-wide and five vendors carried `quota_read_at: null`. `codex` carried `estimated: ["context_pct"]` against `cursor`'s `[]`, and the drop-file entry carried `self_reported: true`. - **That document was written to a file and validated against `docs/snapshot.schema.json` with `tools/validate-snapshot.py`: `ok`.** The published contract holds on the new surface, checked with the gate the CLI's own document is checked with rather than with a second reading of it. - `--vendor chatgpt` came back as a tool result with `isError: true` naming the accepted words; `fleet_quota` came back as JSON-RPC `-32602` naming the tool that does exist; `resources/list` came back as `-32601` naming the methods that do. - **Writes nothing, measured on the real store**: `~/.telltale` held the identical 35 files, sizes and modification times before and after a full session. CI drives the same sequence against the built binary on every run, and pipes BOTH spellings of the tool's document — the text content and `structuredContent` — through the same validator (`.github/workflows/ci.yml`). **That gate was proved non-vacuous the way §7.22's schema gate was**: one sentence was removed from the tool description, the binary rebuilt, and the gate failed naming the missing word; the sentence was restored and it passed again. **What is NOT verified, stated so nobody reads the run above as more than it is.** **No third-party MCP client has connected to this server.** The drive was telltale's own scripted client, which proves the framing, the document and the error channels, and says nothing about how a shipped client negotiates a version, orders its requests, or renders a tool result. Wiring one up writes an entry into the operator's own client configuration, which is his to make; until he does, this section claims a correct server and not a working integration. `STATE.md` carries the debt. ### 7.26 `telltale history` — what one vendor spent, day by day, from its own files (2026-08-29) Every reader this product has answers about NOW. The statusline and the HUD render a scan; §7.22's `snapshot` and §7.25's `mcp` serve one scan to a program; §9.42's `doctor` reports a preflight. None of them can answer *where did last week go*, and the reason is structural rather than an omission: the HUD's Claude read is a head+tail parse — 64 KiB in, 256 KiB back (§3.1) — because §7.18 measured a whole-corpus walk at 164 MB and 46,727 records per second against a 1 s poll. A history needs every record of every file. **Measured on the owner's own corpus, 2026-08-29: 701 transcripts, 143,304 records, 10.4 s for a 30-day walk.** That is four orders of magnitude off a poll tick, so this is a foreground mode that reads once and returns, on `doctor`'s precedent, and never a page in the HUD. #### It is SPEND-shaped, and §7.16 and §7.17 are the vocabulary Nothing here is new grammar. §7.17's table of two claims puts this surface entirely in the right-hand column, and the consequences are the ones already ruled: - **No gauge, no percentage, no bar, no countdown, no ceiling.** There is no denominator anywhere in a token count. `TestTheFrameBorrowsNoneOfQuotasVocabulary` is `TestUsageSpendBorrowsNoneOfQuotasVocabulary` on this surface, and CI asserts the same properties on the built binary's stdout. - **A sum never prints without its window** (§7.16's accumulation ruling). Every count carries its day; the report carries the span it walked and the zone it resolved days in. - **Never a fleet total.** §7.17 already rejected one as "arithmetic telltale invented"; this mode goes further and reports **one vendor per run**, so the arithmetic is not available to make. The vendor's name is on the title, on the window line and on both refusal paragraphs. **One rule here is narrower than §7.17's and is new.** The four columns are not added together either. Input, cache read, cache write and output are four separately billed categories and telltale holds no price, so their sum would be a number that reads like a bill and is not one — the §7.16 derived-`inputTokens` trap arriving through addition instead of through a vendor's own arithmetic. `TestNoTotalIsRenderedAnywhere` computes what such a total would render as and fails if it appears anywhere in the frame. #### Absent is not zero, on two axes §4a.1's rule, applied to a table instead of to a cell: | | What it means | What it renders | |---|---|---| | a request whose usage block reported zeros | a measured zero — the request happened | `0` in all four columns, `1` request | | a workspace with no request that day | it sent nothing | **no row** | | a day inside the window with nothing at all | nothing was written that day | **no row** | The failure this refuses is the natural implementation: pre-seeding a bucket for every (day, workspace) pair so the table comes out rectangular. **A rectangular table here is a table of claims** — every zero cell in it would assert a request nobody sent. What makes a missing day readable instead is the window line: the report states the span it walked, and a sentence under the table says in words that a day with no row carried no token-bearing record. `zero-vs-absent.txt` is the golden, named after `internal/hud`'s for the same reason. #### The derived value is the DAY, and it is disclosed in words rather than with a `~` The vendor writes an instant. A calendar day is that instant resolved in a time zone, and the zone is a choice telltale made — so the window line reads `days resolved in EDT UTC-04:00`, and the OFFSET is on it because a zone abbreviation alone is ambiguous across regions. This is deliberately not a `~`: that marker means an estimated VALUE (§4a.1), and a day bucket is an exact reading under a stated convention. Marking it `~` would say the count might be wrong, when what is conventional is which side of midnight it fell on. Two related refusals, both measured rather than assumed: - **A record with no readable timestamp is in no day.** It is counted and named in diagnostics, never folded into today — that would move a measurement onto a day nothing said it belonged to. - **A record stamped ahead of the clock cannot be dated**, on every adapter's own `futureSkew` rule. A skewed clock must not be able to invent a day's spend. #### The survey: which vendors could support this, and which cannot The mode covers **claude only**, and the six it does not cover are named on **every run**, each with the reason. That block is not documentation politeness — it is the whole of what stops a table headed with one vendor's name from being carried away as a fleet answer, and it prints unconditionally for the reason `doctor` prints its three-state legend every time. The survey is a **source read of this repository's adapters** at the revision it was written on — each adapter's record struct, its package doc, and the live-corpus verdicts those docs already carry. It is **not** a fresh measurement against a live vendor, and the difference is stated because CLAUDE.md's measured-claims rule makes it load-bearing: the version pins these verdicts rest on are the adapters' own `VerifiedAgainst` constants. Two questions, in order. A vendor joins only when both answer yes, and **the second is the one that surprised the survey**: a count with no timestamp is not a smaller history, it has no day axis at all. **Amended 2026-08-29, the same day: the source read was caught out on grok, exactly where the caveat above says it can be.** grok's row said the record carried a total and no date. A live re-measure at grok 1.0.5 read a real `turn_completed` record off disk and found a full input/output/cache split beside the envelope's own `timestamp` ([§3.9a](#s3-9a)'s 2026-08-29 block). The reading of `internal/adapter/grok` was correct — that struct does parse `totalTokens` alone. The error is that **a record struct is an allowlist, so what it omits is a decision and not an absence**, and the verdict reported the omission as the file's shape. The row below is corrected and grok stays uncovered, because nobody has built or measured the coverage; it is no longer refused on fields. The general rule this buys is worth more than the row: **before a vendor is built here, re-read its records, not its struct.** | vendor | counts on disk? | dated? | verdict | |---|---|---|---| | **claude** | **yes** — `message.usage` carries four raw counts (`input_tokens`, `cache_read_input_tokens`, `cache_creation_input_tokens`, `output_tokens`) on every assistant record | **yes** — RFC3339 `timestamp`, and `cwd` on the same record | **COVERED.** Four billed categories, dated, per request, per project. The only vendor with the cache split that makes four honest columns. Pinned at Claude Code 2.1.233 (§3.1) | | codex | yes — a `token_count` event carries `info.last_token_usage` (this turn) beside a cumulative `info.total_token_usage` | yes — the rollout envelope's own `timestamp` | **the next slice.** Two things owed: which of the two a day may sum is a ruling nobody has made, and there is no cache split, so a codex block carries two columns where claude's carries four | | gemini | partial — `tokens.input` is `promptTokenCount`, which [§3.7](#s3-7)'s adapter labels a context-occupancy proxy, and the cached subset is not separable from what it parses | yes | refused on UNITS. Summing an occupancy proxy per day counts one conversation's prefix once per turn, under a header that would read like uncached input | | agy | yes, and the best-guarded in the fleet — `gen_metadata` carries uncached input and output per generation behind the `thinking + answer == output` identity §3.8 requires | **no** — the reverse-engineered field map carries no per-generation timestamp | refused for want of a DAY. Real numbers, no axis to put them on | | cursor | **no** — `tokenCount.inputTokens`/`outputTokens` were 0 in 310 of 310 message rows; declared `CapNone` (§7.16) | n/a | nothing to read. This vendor's counts arrive by hook, not on disk | | grok | **yes, and this row said "a total only" until 2026-08-29** — `turn_completed`'s `usage` carries `inputTokens`, `outputTokens`, `cachedReadTokens` and `cacheCreationTokens` beside `totalTokens`, measured at grok 1.0.5 and on disk since 1.0.0 ([§3.9a](#s3-9a)) | **yes** — the envelope's own `timestamp`, which `internal/adapter/grok`'s struct does not parse | **the old verdict described the adapter's STRUCT, not the record**, and both halves of it were wrong about the file. Not refused on fields any more; simply not built. One unit trap is owed first: `inputTokens` INCLUDES the cache read here and claude's `input_tokens` excludes it, so the four columns are not the same four | | pi | yes — `message.usage.{input,output}` per assistant message, with a `cwd` | yes — record `timestamp` | datable, second after codex. No cache split, and it carries `usage.cost.total` per message — money, which this mode renders nowhere and would have to rule on | `self-reported` is absent from the table and that is not an omission: §7.23's drop-file rows are what a tool said about itself, with no session store behind them to walk. Giving it a "not covered" verdict would imply a file this mode could learn to read. **The block renders in fixed fleet order**, the same order §7.17's blocks and the header's per-vendor counts walk. Ordering it by how close each vendor is to coverage reads better as a roadmap and was **declined**: it would make one list in this product order vendors by a property no other list orders them by, which is the reshuffle §7.17 spends a paragraph refusing. The roadmap signal lives in the words — codex's verdict says it is the next slice. #### The layout, and what each choice is for The generated render (`internal/history/testdata/golden/ledger.txt`, at the default 100 columns): ``` telltale history — what claude spent, day by day, read from claude's own session files read from C:\src\home\.claude\projects window 7 local days, 2026-08-23 through 2026-08-29, days resolved in TST UTC-05:00 read 41 transcripts, 39,184 records DAY WORKSPACE IN CACHE READ CACHE WRITE OUT REQUESTS SESSIONS 2026-08-24 C:\src\code\telltale 1,204 1,903,551 62,004 13,118 14 2 2026-08-27 C:\src\code\notes-api 96 0 4,102 812 3 1 C:\src\code\telltale 22,140 8,830,112 511,903 140,277 191 5 2026-08-29 ...rkspace-path\telltale 3 12 0 44 1 1 ``` - **Plain text, no colour, no TUI** — `doctor`'s argument (§9.42), and it applies harder here: a history is read in a pipe and pasted into a message. Every distinction the report draws is a WORD, which satisfies §7.1 rule 2 by having no first signal that is not one. `--ascii` and `NO_COLOR` have nothing to switch off, so neither is a flag: a flag that does nothing is a promise that something was configurable. - **`REQUESTS`, not `TURNS`.** One turn can produce several API requests, so "turns" would be a count telltale did not take. The column is the number of records that carried a usage block, which is exactly what was counted. - **Counts are exact and grouped, never floored.** `theme.Tokens` floors to `1.9M` on the gauge surfaces because a header line has no room for digits and rounding *up* would invent tokens nobody was billed for (§7.16). This is a table with the room, so it rounds nothing at all — strictly the more honest of the two, and affordable only here. - **The day is drawn once per day**, not restated on a second workspace row: repeating the date makes the eye read a second reading where there is one. - **Rows are day-ASCENDING**, so today lands at the bottom, next to the prompt the reader is looking at. Sorting by spend was refused for §7.17's reason — position is the navigation. - **The workspace column is the only one allowed to give way**, and it truncates from the LEFT: a path's identifying half is its tail, and the marker sits at the front where it says "something was removed" before the reader has read the value. Below the floor the table overruns the wrap column rather than letting numbers collide (`narrow.txt` is the golden at 60 columns). - **A record whose own record named no `cwd` gets its own `(no cwd)` bucket**, never a neighbour's. Attributing it to the last workspace seen in the file would be a guess, and a guess in the project column is indistinguishable from a reading once it is on screen. #### The read/write boundary **This mode writes nothing at all.** It reads one vendor's store, calls no network, binds no port, reads no credential, and relays no quota — it renders none, which is `snapshot`'s own argument for holding the contract with one item spare (§7.22). It joins statusline, hud, snapshot and mcp as a **reader**; `CLAUDE.md` names it in that list. `internal/history/boundary_test.go` is the mechanical half, on `internal/eventview/boundary_test.go`'s precedent: `go list` answers what this package imports, and the gate fails if it ever reaches `quotacache`, `usagecache`, `eventsink`, `eventview`, `council`, `net/*`, `os/exec` or a TUI module. The check is on DIRECT imports and says so — this package imports `internal/adapter/claudecode` for `Discover`, and an adapter's dependency graph is not this mode's write surface. **Content cannot reach a rendered value**, by the technique `internal/cursorhook` uses against a payload carrying a user's email beside four numbers: the record struct IS the allowlist, and `encoding/json` drops every field with no destination. The only strings that survive a parse here are the workspace path and the timestamp. Diagnostics carry counts and never bytes. #### What it reuses, and the one thing it does not Sessions are discovered by `internal/adapter/claudecode`'s own `Discover`, so a session this mode counts is a session the HUD would draw and the two cannot come to disagree about what a session is — including the two traps that function encodes (the glob is not recursive, and a basename is validated as a UUID; recursing inflated the live session list 2.4× and double-counted every token). Records are framed by `internal/jsonl`, so the U+2028 trap and the 1,004,230-byte record are handled in the one tested place. What it does NOT reuse is that adapter's head+tail parse. A ledger needs every record, so it walks whole files through `jsonl.Scan`. Three record classes are refused and each is counted in diagnostics rather than dropped silently: unparseable records, `` records (Claude Code's own locally generated notices, which carry a zeroed usage block and would otherwise look exactly like the measured zero above), and inline `isSidechain` records — 0 of 179,614 in the live corpus, so a non-zero there is a vendor change worth seeing rather than a routine skip. #### Verified against the built binary, not only the suite The suite drives the host in-process and across a process boundary it starts itself, which is two thirds of the claim. The last third is that `telltale.exe` really does this, so the shipped binary was driven by a client that shares none of its code. **2026-09-01, Windows 11 Pro 10.0.26200, `telltale.exe` built from this branch.** The host was started as `telltale council host --pipe \\.\pipe\telltale-council-smoke --workspace --vendor claude --read`, and a **PowerShell** `NamedPipeClientStream` — no Go, no telltale code — opened the pipe and wrote one line. What came back, whole: ``` {"kind":"welcome","protocol":1,"host_pid":15900} {"kind":"room","room":{"version":1,"workspace":"…\councilhost-smoke","turn":0,"posture":"read", "seats":[{"vendor":"claude","binary":"…\claude.exe","phase":"idle","drivable":true}]}} ``` Four things are verified there and each was a separate way to be wrong. The descriptor admitted a same-user client. `host_pid` matches the process that was started, so the client-side `GetNamedPipeServerProcessId` check ran against the real server. The seat resolved a real binary and drew `idle` — **no vendor was spawned**, because nothing was dispatched, which is the "never start a vendor to see whether it answers" rule holding across the new process boundary. And when the PowerShell client disposed its stream, **the host exited on its own**, which is detach being unexposed, measured rather than asserted. #### Known limitations, named - **The window is complete or it says so.** A walk stopped by `--timeout` prints what it read and marks the report incomplete, in its own paragraph: the ROWS stay true and the WINDOW stops being, and a reader who misses that sentence would read a lower bound as a total. - **A day is a local calendar day.** A session that crossed midnight in another zone lands where this machine's zone puts it. The offset is on screen; nothing converts. - **`REQUESTS` counts records, not API calls, if the vendor ever writes two records for one call.** No such case is known at 2.1.233; it is named because the column's honesty rests on the vendor's record-per-request shape rather than on anything telltale can check. - **Nothing here is cached.** Every run re-walks, at the cost measured above. A cache would be a ledger that can disagree with the files it came from, and the mode is not on a tick. ### 7.27 `telltale council ls` — the saved room, read and never opened (2026-09-01) `telltale council` is the only way to see what the saved room holds, and it is an expensive way. It enters the alternate screen, it detects every vendor, and after this section's sibling (§9.52) it also rebuilds the seats. An operator who only wants to know *what is saved* must not pay for a room to find out. `telltale council ls` answers that question and does nothing else. The mode also has a job that outlives this question. §9.52's rebuild and the later host work both need one place that reports what is on disk. A discovery surface built under a deadline, beside the feature that needs it, is a surface that inherits that feature's shape. This one is built first and alone. #### It is the SIXTH reader, and it holds the same contract CLAUDE.md's read/write boundary lists five readers: `statusline`, `hud`, `snapshot` (§7.22), `mcp` (§7.25) and `history` (§7.26). This is the sixth, and it is closest to `history` — it reads no scan at all. It reads exactly one file, `~/.telltale/council/room.json`, through `LoadRoom`, the same loader the room itself uses. - It **writes nothing**. Not the file it read, not a cache, not a lock. - It **spawns nothing**. No vendor process starts. `exec.LookPath` is the deepest it reaches, and that resolves a name against `PATH` rather than running a program. - It **binds nothing**. No port, no pipe, stdout only. - It **relays no quota**, so it holds the contract with the same one item spare that `snapshot` and `history` hold it with, and for the identical reason: it renders no quota of its own, so it has none to relay. The reason this is stated at the same length the other five state it is that council is the product's one ratified exception. A council sub-mode that read like a gauge but wrote like the room would be the exception growing by accident. This one is a gauge. #### The three states a seat can be in, and why two of them are not one §4a.1's rule is the whole of the per-seat output. A saved room names a set of vendors, and each one is in exactly one of three states: | what is true | how it renders | why it is its own state | |---|---|---| | a session id is saved, and this machine can run the vendor | `saved` | the only row a rebuild can act on | | a session id is saved, and this machine cannot run the vendor | `saved, not installed here` | the id is real and unreachable from this box. A room opened here will not rebuild it. | | no session id is saved for this seat | `no thread saved` | a measured absence: the seat was in the roster and never answered, or its thread was cleared | Collapsing rows two and three would tell an operator on a second machine that a conversation is gone when the id is on disk and the vendor is missing. That is the same class of error as a column of dashes for a vendor that could never fill it. #### What it deliberately refuses to say **It never claims a thread is alive.** Nothing this mode can read proves that a vendor still holds a session. Only the vendor answers that, and only when a process asks it to resume. So the word is `saved`, never `live` and never `resumable`. §9.52 states the same limit from the room's side: a launched process is not a proven thread either. **It never verifies an id against a vendor.** Verification means a spawn, a spawn means a vendor process, and on some seats a resume attempt is billable. A read mode that quietly spends money is not a read mode. The cost of the honest answer is one word: `saved`. **It prints no content, because the file holds none.** `room.json` is session ids, a workspace, a brief PATH and a handful of scalars (`resume.go`'s `SavedRoom` doc comment is the contract). This mode cannot leak a conversation because it has no conversation to leak. **It does not restore the posture it prints.** The saved posture is a record of what the room stood in, and `reattach` already refuses to re-apply it. A reader that printed it as though it were the posture of the next launch would undo that ruling in a listing. #### Shape Words and no colour, on `doctor`'s and `history`'s precedent. Every fact carries its own label, so no column alignment has to survive a narrow terminal. **It takes no flags at all, and that is a decision.** The other readers take `--root` to point at a corpus of *vendor* stores. This mode reads telltale's own state, so `--root` here would have to mean a different thing, and one flag with two meanings across two modes is worse than no flag. It is also a two-word mode rather than a `--list` on the room, on the `hook cursor` / `events view` / `otel grok` precedent: none of the room's flags apply to it, and a flag would have to explain why it ignored every one of them. A refused file prints the reason `LoadRoom` gives, in `LoadRoom`'s own words, and exits 0. A damaged file on disk is a state to report, not an error to fail on — the same ruling the room makes when it opens anyway and says why. No saved room at all prints one sentence naming the command that makes one. ### 7.28 `telltale council host` — the room in a process of its own (2026-09-01) **What this section rules.** Council runs the room and the screen in one process. This section splits them. A HOST process owns the vendor processes, the pipes and the room state. A CLIENT process connects to the host and renders. The two speak newline-delimited JSON over a Windows named pipe with an explicit security descriptor. **What this section does NOT ship, stated first so nobody reads a promise into it.** Detach is not exposed. No key detaches the room, no command rejoins one, and nothing survives the client. The client starts the host, drives it, and kills it on exit. `telltale council` is unchanged: the daily command still runs the single-process room, and every golden file in `internal/council` still describes it. This is the process boundary and its transport, built and measured, with the feature that needs them deliberately withheld. The ownership inversion is the risk; shipping it alone is how the risk gets reviewed alone. #### Why a host must PARSE, and why "hold the processes" is not a smaller version of it The tempting cheap host holds the child processes and lets the client keep reading them. It does not work, and the reason is mechanical rather than aesthetic. `pumpStdout` (`internal/council/runner/runner.go`) drains each child's stdout continuously. Nothing else drains it. If nobody reads, the operating system's pipe buffer fills, and the next write the vendor makes blocks. The vendor then stops mid-turn. So a host that only holds processes is a room that silently stops working the moment the reader goes away, which is the opposite of the property the split is for. The second cheap version fails on the same fact from the other side. "Do not kill the seats" leaves the pipe handles owned by the process that made them (`cmd.StdoutPipe()` in `session.go`). The agents then survive as processes nobody can read and nobody can write. That is worse than killing them, because it spends quota with no channel to see it on. **So the host parses. That is what makes it a host and not a babysitter, and it is not optional.** #### No pseudo-console, and this is a property of council's seats A terminal multiplexer hosts a terminal. tmux and Zellij emulate one, which is most of their size. Council's seats are not terminals. They are line-oriented JSON processes on anonymous pipes (`session.go`), started with `CREATE_NO_WINDOW` and `HideWindow` so that they have no console at all (`proc_windows.go`). There is no screen to emulate and no size to track. `CreatePseudoConsole` is therefore **out of scope by ruling, not by omission.** A later session must not reach for it by reflex. Council would need it only to host an interactive vendor TUI, and council drives no vendor that way. Repaint at any width is already free. `Render` is pure over `State`, and `TestRenderIsPure` holds it there. A client hands its own width to the same pure function. The hardest problem in a Unix multiplexer — tell the server the new size, resize the pty, repaint — does not exist here. #### The transport: a named pipe, and [§7.24](#s7-24) is the reason [§7.24](#s7-24) measured a loopback bind and found it was not containment. A headless Chrome, on a page the operator merely visited, planted a forged row in `usage/grok.json`, planted an event in the sink, and read the sink's whole verbatim store over `ws://127.0.0.1:41519/stream`. The fix was to refuse any request carrying `Origin` and to require the measured sender's media type. **This socket is strictly worse than those two if it is reached the same way.** It carries transcript content in both directions, and it accepts dispatch commands. The room writes by default, and three of the four seats are batch CLIs with no channel to ask permission on. A page that can post a turn into a hosted room can spend the operator's quota and edit their working tree. §7.24's two-arm check would have to hold perfectly, on a surface worth far more than four token counters. **A named pipe removes the class instead of filtering it.** No URL scheme addresses `\\.\pipe\...`. `fetch`, `XMLHttpRequest` and `WebSocket` cannot reach it. That makes `internal/localonly`'s check unnecessary here rather than merely satisfied, and unnecessary is the stronger of the two positions: the check §7.24 exists to perform has nothing left to perform. **This is a successor to §7.24's ruling and not an exception to it.** §7.24 narrowed who may talk to a socket that was already loopback. This section picks a transport that the excluded sender cannot address at all. Loopback TCP is therefore **refused**, and the refusal is recorded here so a later session does not re-derive it. A file-based transport is refused too, on three counts: it cannot carry a live stream without polling, it cannot answer "is the host alive" without a liveness heuristic, and it reopens "who may write this file" with no bounding principle. A pipe's answer to the last one is an ACL the operating system enforces. **The dependency ruling: `golang.org/x/sys/windows` only, and no `go-winio`.** x/sys/windows is already a direct dependency and carries `CreateNamedPipe`, `ConnectNamedPipe`, `DisconnectNamedPipe`, `CreateFile`, `CancelIoEx`, `CreateEvent`, `GetOverlappedResult`, `SecurityAttributes` and `SecurityDescriptorFromString`. That matches this repo's recorded habit of a page of checked stdlib code over a dependency (`decisions/001`). **AMENDED 2026-09-01, the same day, and the amendment is the interesting part.** This section first ruled that overlapped I/O was the one thing that would justify a dependency and that the design did not need it — a blocking `ConnectNamedPipe` on one goroutine, one client at a time, because a second simultaneous client is refused anyway. **That was reasoned, and it was wrong.** A synchronous pipe DEADLOCKED the room on its first streamed frame. Measured on Go 1.26.6, Windows 11 Pro 10.0.26200: the host's reader was parked in `ReadFile` waiting for the client's next command, the host's writer was parked in `WriteFile` holding a 277-byte room frame, and the client was parked reading. **Windows serialises every operation on a synchronous handle**, so a read that is waiting blocks a write on the same handle until it finishes. The mistake was not about pipes; it was about the shape of this protocol. One client at a time bounds CONCURRENCY and says nothing about DIRECTION, and this host is full duplex on one handle by construction: it pushes room frames while it waits for commands. Half duplex would mean the room could only draw when the operator typed, which is not a room. So both ends open with `FILE_FLAG_OVERLAPPED` and go to `os.NewFile`, which associates them with the Go runtime's completion port. The only hand-rolled overlapped work is `ConnectNamedPipe`, about ten lines, and it pays for itself twice: `CancelIoEx` on that pending connect is how `Close` wakes a waiting `Accept`, which retired a self-connect hack the synchronous handle had forced. The dependency ruling is unchanged. **The general lesson worth keeping: "one client at a time" is a statement about how many peers, never about how many directions.** **The wire format is newline-delimited JSON, one frame per line.** Not for speed. The project already has the parsers, the framing rule (§4's JSONL framing rule), and the habit across every seat protocol (`runner/protocol.go`). #### The security descriptor, measured rather than cited **The default is a leak, and it must never be used.** With `lpSecurityAttributes == NULL`, `CreateNamedPipe` gives the pipe a default descriptor. Microsoft's own page says that descriptor grants read access to the Everyone group and to the anonymous account. This repo's rule is that a claim about behaviour is measured and not read off documentation ([ADR-001](#adr-001)), so it was measured. **MEASURED 2026-09-01, Windows 11 Pro 10.0.26200.** A pipe created with a NULL `SecurityAttributes`, its DACL read straight back off the handle with `GetSecurityInfo`: ``` D:(A;;FA;;;SY)(A;;FA;;;BA)(A;;FA;;;S-1-5-21-...-1001)(A;;FR;;;WD)(A;;FR;;;AN) ``` `WD` is Everyone and `AN` is ANONYMOUS LOGON, each carrying `FR` — `FILE_GENERIC_READ`. The documentation page is confirmed on this machine. What council's pipe carries instead, read back the same way: ``` D:P(A;;FA;;;SY)(A;;FA;;;BA)(A;;FA;;;S-1-5-21-...-1001) ``` `GA` is mapped to `FA` at creation, which is why the applied string and the object's own string are not byte-identical, and why a test must compare ENTRIES rather than the whole line. **The trap this measurement caught, recorded because it will recur.** The first version of `TestTheDefaultDescriptorIsTheLeakWeRefuse` looked for the raw SID `S-1-1-0` and came back CLEAN — it would have reported the pipe safe while Everyone was on it. Windows renders a well-known SID as its two-letter alias in an SDDL string, and which spelling you get is the API's choice rather than the object's state. **Both spellings are checked.** A security assertion that matches on one rendering of an identity is not an assertion. That test is the NEGATIVE CONTROL for the test beside it. Without it, the positive test only proves that a string was applied, and not that the string prevents anything. **The descriptor this design applies:** ``` D:P(A;;GA;;;SY)(A;;GA;;;BA)(A;;GA;;;) ``` - `D:P` makes the DACL protected. No inherited entry is added. - `SY` is LocalSystem. `BA` is the Administrators group. The third entry is the literal SID of the account the host runs as, read with `Token.GetTokenUser()`. - **Administrators are admitted deliberately.** An administrator can already open the host process for debug, so a denial on the pipe would be theatre. Stating that is better than a descriptor that looks stricter than it is. `gatehook.go` takes the same posture about `0600` on Windows. - **The literal SID, and not `OW` (CREATOR OWNER).** `OW` is a placeholder that an object substitutes at creation, and whether a named pipe substitutes it the way a file does was not measured. The literal SID needs no such answer. `TestThePipeCarriesTheExplicitDescriptor` creates the real pipe, reads its DACL back with `GetSecurityInfo`, and asserts three things: the three intended SIDs are present, `S-1-1-0` (Everyone) is absent, and `S-1-5-7` (ANONYMOUS LOGON) is absent. **Who can still connect: any process that runs as the same user.** That is exactly the boundary §7.24 ratified and pinned. Quoted from it, because the sentence is the contract: *a program running on this machine as a principal `~/.telltale/`'s ACL admits is trusted by these listeners exactly as far as it is trusted by the filesystem.* The descriptor above makes the pipe **equally** permissive as `~/.telltale/`, and never more. **Another user on the machine: no, with this descriptor. Yes, read-only, with the default.** That difference is written down here rather than left to a reader to work out. #### Peer verification, in both directions §7.24 had to REFUSE operating-system peer identity for loopback TCP. It needs `GetExtendedTcpTable`, which is not stdlib, and it does not answer the browser case anyway, because the peer there is a legitimate `chrome.exe`. **On a named pipe the check is available and it does answer.** This is the strongest single argument for the transport, over and above the browser argument. - **Server side.** The host calls `GetNamedPipeClientProcessId`, opens the client process, reads its token user, and refuses any client whose SID is not the host's own. - **Client side.** The client calls `GetNamedPipeServerProcessId` and checks the same way. This is the anti-squatting arm. The classic Windows attack is a lower-privilege account that pre-creates a well-known pipe name, so that a later server or client attaches to theirs. **Measured at `golang.org/x/sys` v0.47.0:** `GetNamedPipeClientProcessId` and `GetNamedPipeServerProcessId` are both exported. `ImpersonateNamedPipeClient` is **not**. The design uses the process-id route for both arms rather than adding a `syscall.NewLazyDLL` binding, and the choice is an improvement rather than a fallback: impersonation changes the calling thread's token and must be reverted on every path out, and a missed `RevertToSelf` leaves a thread running as somebody else. `FILE_FLAG_FIRST_PIPE_INSTANCE` is set on the create. A name another process already holds then fails the create outright instead of adding an instance to a pipe somebody else owns. #### Containment: a nested room job, and the property that must not be lost `proc_windows.go` states the property in its own words: `JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE` covers "the case we cannot code around" — if telltale dies, the handle closes and Windows reaps the whole tree. **That property is preserved and moved one process outward. It is not weakened.** The per-seat job stays exactly as it is, so that seat eviction and turn cancellation still kill exactly one tree. One ROOM job is added. It carries `KILL_ON_JOB_CLOSE`, the host assigns ITSELF into it first, and then every seat lands in the hierarchy under it. The host is the only holder of that handle. | how the host ends | what happens to the seats | |---|---| | clean quit | the host kills them deliberately, then exits | | panic or unhandled fault | the process dies, the handle closes, Windows reaps every seat | | `taskkill /F`, Task Manager End Task | the same — the handle closes, every seat is reaped | | the machine loses power | nothing survives anyway | **The nested-job reap is measured, not read.** Microsoft's *Job Objects* page states that a nested job with `KILL_ON_JOB_CLOSE` terminates its processes and its child jobs when the last handle closes. `TestAHardKilledHostReapsEverySeat` (`internal/councilhost/roomjob_windows_test.go`) runs it instead. The test binary re-executes itself as a stand-in host, which builds the room job, assigns itself, starts a grandchild seat behind a per-seat job, and reports the grandchild's pid. The test then calls `TerminateProcess` on the stand-in host, which is what `taskkill /F` does, and asserts the grandchild is gone. `TestTheRoomJobHoldsTheHostAndTheSeat` pins the structure with `IsProcessInJob`, so that a pass cannot be earned by the per-seat job alone. #### The crash blast radius, and the three things that bound it **The baseline is identical to a telltale crash today.** Every seat dies, the conversation in RAM is lost, and `room.json` survives. That is the current behaviour with a process boundary moved, and it is not a regression. What is new is that **the operator cannot see it happen**, because the host has no terminal. Three mitigations, and all three are required. 1. **The client renders the death.** A broken pipe is reported as *the host exited, and the seats went with it*. It must never render the same way as an ordinary disconnect. Two states render two ways is [§4a.1](#s4a-1)'s whole discipline, applied to a process. 2. **`room.json` stays the floor.** The host writes it on the same schedule the single-process room does. A hard-killed host therefore leaves the session ids behind, and the next `telltale council` falls into the existing resume path. **This work is strictly additive: it never removes the fallback.** 3. **No auto-start, ever.** No Windows service, no Run key, no scheduled task, and no restart-on-crash supervisor. The operator starts the host and the operator ends it. A host that resurrects itself is a host the operator cannot reason about, and it would make `telltale doctor` dishonest. #### The read/write boundary: what the host writes, and what it must never write **Transcript content is never persisted. Not under `~/.telltale/`, not in a temporary file, not compressed, not encrypted, and not "only the last N turns".** It lives in host memory and it dies with the host. `resume.go` already ruled this for the same data: every vendor stores its own history against its own session id, so a second copy here would be a private conversation in a place the user did not ask for. The rule does not change because the process holding the data changed. The one precedent for verbatim storage is the event sink, and CLAUDE.md records that its containment "is scope, not redaction" — its own foreground mode, started by the operator, read by no gauge. A host is started by a client on the daily path, so that grant's own reasoning refuses to extend here. The host writes what council already writes: `council/room.json`, session ids and workspace, never content. It adds one file, `council/host.json`, and the argument for it is that it is the same class of file, in the same directory, holding the same class of value, for the same purpose — the keys that let a later launch find what is already there. It holds `version`, `pid`, `pipe`, `started_at`, `workspace`, `seats` and `turn`. Four of those are already in `room.json`; the other three are process facts. `resume.go` already states the leak profile of this exact shape — which directory was worked in, when, and a set of opaque ids, and not a word anyone said — and a pid does not change that sentence. **Liveness: the file says WHAT, the pipe says WHETHER.** A pid is reusable, and a stale `host.json` is the normal case after a hard kill. So `host.json` is never read for liveness. A host is running if the pipe opens, and that is not a heuristic. #### The spawn guard, extended in the same change `internal/council/main_test.go`'s `TestMain` makes the council package's spawn vars fail closed. It exists because the opposite default was measured starting `codex exec --json -s danger-full-access` from a plain `go test` run — a live agent turn, with full write access, on the operator's own account. **CI can never catch that class**, because CI has no vendors installed and nothing dispatches. A host spawns from a DIFFERENT PROCESS and a different package, so it is outside that wrap. Extending the guard is therefore part of this change and not a follow-up. Two new spawn paths exist and both are guarded: - **The host's own vendor spawn.** `internal/councilhost` has its own `TestMain`, which wraps its spawn vars on the same rule: a binary this machine can resolve panics, and names the call site and the full argv. The rule is copied rather than re-invented, because it is the same question the operating system is about to ask. - **The client's spawn of the host.** This one is sharper. It starts `telltale.exe council host`, which resolves on any machine that built the binary, and that host then starts real vendors — so an unguarded test would launch billed turns two processes away from the assertion. It goes behind a var in package `council`, guarded in that package's `TestMain` and stubbed in `countSpawns` with the restore added to the existing `t.Cleanup`. #### Naming: "reattach" is already taken `reattach` means *resume the vendor session ids saved in `room.json`*. It carries that meaning in the code (`Reattachment` in `resume.go`), in the golden files (`testdata/golden/reattached.txt`), and in the demo script. Overloading it would make the room's own notice ambiguous at the exact moment the notice is trying to be honest about what was restored. **Ruling: `reattach` keeps its current meaning. The verb for a live host is `rejoin`.** Two words, two facts, so a notice can say which one happened. Renaming the old one to `resume` is cleaner and is not taken here: it is golden-file churn across at least four files, and it is the owner's call. `host` is the noun, over `server` and over `daemon`. `server` is the word this product's thesis refuses out loud. `daemon` is a Unix word on a Windows-first product ([ADR-002](#adr-002)). A host holds the room. #### Known limitations, named - **Host memory is unbounded.** A room accumulates turns for as long as it lives. A turn ceiling is owed before detach ships, and the drop must be stated in the header rather than applied silently, on the retention discipline `telltale events` already has (§7.21). - **One client at a time, and the OPERATING SYSTEM refuses the second.** The pipe is created with one instance, so a second open comes back `ERROR_PIPE_BUSY` and the client renders that as "one client at a time". An earlier draft of this section said the refusal would name the holder's pid. It does not: naming it would mean keeping a second instance open purely to answer on, and a second accept path on a security-sensitive surface costs more than the pid is worth. Multi-client attach is a tmux feature, it is not free, and it must not be acquired by accident. - **A stale host is deliberately not mitigated.** A host nobody returns to keeps running and can keep spending. It must NOT self-terminate on idle: a detached room that dies on its own is precisely the failure the operator cannot see. The answer is discovery, and discovery ships before detach does. - **Unix is not built here.** The equivalent is a Unix domain socket in a directory at mode 0700, and it is stdlib. It carries an asymmetry that must be measured into `PARITY.md` first: `proc_unix.go` records that on macOS a process group does not bind lifetimes, so a `kill -9` on a host LEAKS every seat there, while on Windows it does not. That is the reverse of the usual direction. *Amended 2026-09-02: built, and the asymmetry is in `PARITY.md` — [§7.30](#s7-30).* - **A GATED room is refused, not hosted.** A gated seat blocks on a question, and carrying that question and its answer over this wire is a card, a keystroke, and a reply written back down the vendor's own stdin. None of that is built, so `councilhost.New` refuses `PostureWriteGated` outright, and a gate arriving mid-turn puts a sentence on the seat's card. A blocked seat and a slow seat must not render alike. - **A CONVERSATIONAL seat is drawn as undrivable.** cursor-agent's ACP server cannot be handed a turn by writing a line — its turn cannot be built until the vendor has answered a request of the room's own ([§9.36](#s9-36)) — so the host names it and refuses it rather than dispatching into a column that would never finish. - **The unwatched-write ruling is still owed, and it is not owed YET.** Detach plus a write posture plus `--auto` is a risk shape the docs have no ruling on: the room writes by default, and three of the four seats cannot ask permission. It is not owed here because nobody is unwatched — the client holds the room for the whole of its life. It must be ruled in the same change that exposes detach, and never as a default that happened. - **The room state on the wire is the host's own projection, not `council.State`.** `State` carries pointers and rendered projections, and folding it onto a wire format is the work that makes council's own `Model` the client's renderer. That is the next slice, and it is named here so that nobody reads this rung as having paid for it. ### 7.29 detach, rejoin and kill — the room outlives the terminal (2026-09-01) **What this section rules.** [§7.28](#s7-28) built a host and withheld the feature that needs it. This section exposes detach: a client leaves and the host keeps the seats, a later client rejoins the same live process, and `telltale council kill` ends it on purpose. It also carries the unwatched-write ruling §7.28 said was owed with it and never before. **What a reader must not read into it.** No key in the room's TUI detaches anything, and no golden file moved. The reason is stated at the bottom under *the TUI has no host to leave*, and it is a fact about what §7.28 built rather than a shortcut taken here. #### The four verbs, and the one that was reserved §9.52 reserved `rejoin` and spent nothing on it. This section spends it, and the three words still name three different facts: | word | what it is a fact about | when it is true | |---|---|---| | **reattach** | the FILE | the room read `room.json` and holds the saved ids | | **rebuild** | the PROCESS | the room launched a NEW vendor process on a saved id | | **rejoin** | the PROCESS THAT NEVER STOPPED | a client reached a host that was already running | `kill` is the fourth and it is not one of that family: it is a verb about the host, not about a thread. `stop` was refused because it understates what happens — this terminates agent processes that hold live sessions and spend quota. #### Detach is an explicit FRAME, and a bare disconnect still ends the room The tempting version reads a closed pipe as a detach. It is refused, and the refusal is the whole safety argument of this section. **A client that died is not a client that left.** A crash, a `taskkill` on the terminal, a power-off — every one of those closes the pipe exactly the way a deliberate detach does. A host that could not tell them apart would keep a room running on an INFERENCE, and the room it kept running would be one nobody chose to leave. That is the state this product refuses everywhere else: §4a.1 rules that two facts render two ways, and here two facts must not even reach the same code path. So the wire grows one frame, `KindDetach`, and the rule is: | what the client does | what the host does | |---|---| | sends `KindShutdown` | kills every seat and exits — unchanged from §7.28 | | sends `KindDetach` | keeps every seat, re-arms the pipe, waits for the next client | | closes the pipe with neither | kills every seat and exits — unchanged from §7.28 | The third row is the one worth reading twice. **A bare disconnect keeps the meaning §7.28 gave it**, so nothing about a crashed client changed when detach arrived. That is what makes this change additive rather than a re-definition of the existing behaviour. #### The host survives the client's PROCESS, and the mechanism was already there Two properties have to hold, and only one of them is new code. **The socket close is handled by the frame above.** That is the new part. **The process exit needed nothing**, and this is the measured half. `spawn_windows.go` starts the host with `CREATE_NEW_PROCESS_GROUP | CREATE_NO_WINDOW`, and both flags already do the work a detach needs. `CREATE_NO_WINDOW` gives a console application **its own** console with no window, rather than attaching it to the client's — so closing the operator's terminal sends that terminal's `CTRL_CLOSE_EVENT` to the processes on ITS console, and the host is not one of them. `CREATE_NEW_PROCESS_GROUP` keeps the client's ctrl+c out of the host for the same reason one process group down. Windows does not end a child when a parent exits, so nothing else was holding the host to its client except the shutdown-on-disconnect rule above. **`DETACHED_PROCESS` is still NOT used, and §7.28's deferral is now a decision.** §7.28 named it as what a host "wants" and refused to reach for an unmeasured flag. It stays unused, because the property it would buy — no console at all — is one `CREATE_NO_WINDOW` already delivers for this purpose, and the two flags are mutually exclusive. Swapping a measured flag for an unmeasured one to buy a property already held would be the guess [ADR-001](#adr-001) refuses. `TestADetachedHostOutlivesItsClientProcessAndStillReapsEverySeat` measures the claim rather than restating it: a real host in a real process, a real client that detaches and then EXITS, and the host still answering a second client afterwards. #### The containment property is not weakened by detach, and that is measured too §7.28's table says a hard-killed host reaps every seat, and `TestAHardKilledHostReapsEverySeat` runs it. A detached host is the case that table was written for and never covered: the host is now the only process left, so if detach cost the room job anything, nothing at all would reap the seats. **The same test carries both halves, and that is deliberate rather than tidy.** They are one story — a host that outlived its client is exactly the host whose containment has nobody left to check it — and splitting them would have meant two stand-in hosts measuring one process's life. So the second half runs on the first half's host: a stand-in process runs a REAL `Host.Serve` (real `NewRoomJob`, real `Listen`, real handshake) and starts a stand-in seat with **no per-seat job of its own**, so the room job is the only thing that can reap it. The client process detaches and exits, a second client rejoins from the test, and only then does the test call `TerminateProcess` on the host the way `taskkill /F` does and assert the seat is gone. **The seat is asserted ALIVE before the kill, and that assertion is the control.** A reap test whose subject had already died would pass while measuring nothing, which is the failure mode a containment test can least afford. The no-per-seat-job detail is what makes it a measurement rather than a ceremony, and it is carried over from §7.28's own test for the same reason: the per-seat job also carries `KILL_ON_JOB_CLOSE` and its handle also dies with the host, so a seat wrapped in one would be reaped by the mechanism that already existed. #### Rejoin: the file says WHAT, the pipe says WHETHER, and the PID says WHO §7.28 ruled that `host.json` is never read for liveness and left the probe unbuilt, because nothing needed it and the obvious implementation was wrong. That ruling stands and the probe is built here, on three readings that are three different questions: 1. **WHAT would I be rejoining?** `host.json`: pid, pipe name, start time, workspace, seats, turn. Numbers and keys, and the file is unchanged from §7.28. 2. **WHETHER a host is running.** `WaitNamedPipe` with a zero timeout, which asks whether the NAME exists and **does not connect**. That distinction is the whole reason discovery.go refused to build this before: a probe that dials CONSUMES the host's single pipe instance, and the host reads the close as its client leaving. Three answers, and all three are used — `ERROR_FILE_NOT_FOUND` means no host, success means a host with nobody attached, and `ERROR_SEM_TIMEOUT` means a host whose one client seat is already taken. 3. **WHO that pid is.** A pid is reusable, so the pid alone answers nothing. The probe opens the process, reads its image name with `QueryFullProcessImageName`, and requires the same executable name this binary runs under; then it reads the process creation time with `GetProcessTimes` and requires it to be no LATER than the `started_at` the file claims. A recycled pid fails the first check; a different telltale that took the number fails the second. **`WaitNamedPipe`, `QueryFullProcessImageName` and `GetProcessTimes`, measured at `golang.org/x/sys` v0.47.0:** the last two are exported and are called directly. `WaitNamedPipe` is **not** exported, so it is bound with `NewLazySystemDLL` — the same call `IsProcessInJob` in `roomjob_windows.go` already makes, and the one `decisions/001` sanctioned for the hand-rolled OTLP reader, the byte-level SQLite reader and the hand-rolled WebSocket. A page of checked code rather than a dependency. **All three readings must agree before a client rejoins.** A file with no pipe is a host that died. A pipe with a pid that is not telltale is a name somebody else took, and `Dial`'s existing server-side peer check is the second arm on that. Neither of those is an error to fail on; both are states to render. #### The four states this feature can leave an operator in, and none of them render alike §4a.1's rule is the whole of this subsection. `rebuilt`, `survived` and `died` must never render alike, and detach adds a fourth that must not look like any of them. **You left** — printed by the client that detached: ``` detached. the host keeps the seats and the conversation, and it is pid %d. `telltale council` rejoins it. `telltale council kill` ends it, and every seat with it. ``` **You came back and it was still there** — the `rejoin` case, and the only one of the four in which a vendor process was never restarted: ``` rejoined the host that was already running — pid %d, started %s. the seats kept working while you were away. nothing was rebuilt, and no session was resumed. ``` **You came back and it was gone** — the `died` case: ``` the host you left is gone, and the seats went with it. it was pid %d, started %s, and nothing on screen could say when it ended. the room's session ids are still in %s, so `telltale council` rebuilds those seats. ``` The third line points at §9.52's rebuild, and it uses §9.52's own word. A room that told an operator their conversation was gone when the ids are on disk would be the same error `council ls` refuses to make about a vendor that is missing from one machine. **You asked to leave and the room refused** — the unwatched-write ruling, below. The `rejoin` sentence carries the clause `nothing was rebuilt, and no session was resumed` because that clause is the entire difference between this state and §9.52's `rebuilt`. Without it the two are one sentence apart and an operator would read a rebuild as a survival, which §9.52 calls the most expensive lie this surface can tell. #### THE UNWATCHED-WRITE RULING (owed by §7.28, paid here) **A room that writes to the workspace without asking does not detach. The host refuses it, and it says why.** The risk shape §7.28 named is detach plus a write posture plus `--auto`. On a hosted room those three collapse into one condition, and the collapse is worth stating rather than leaving a reader to derive: - §7.28 already refuses `PostureWriteGated` outright — a gated seat blocks on a question this host cannot carry. So a hosted room is never gated. - A hosted room that is not read-only is therefore an **ungated** write room: every tool call runs with nobody to ask. That is exactly what `--auto` means on the room's own surface (`dispatch.go`'s `seatPosture`: write plus not-asking is `PostureWrite`). So the condition is one word — the room's posture — and the refusal is keyed on it. | posture | detach | |---|---| | read | allowed | | gated write | allowed by this ruling, and **unreachable**: §7.28 refuses to host a gated room at all | | write (ungated, which is `--auto`) | **REFUSED** | The refusal sentence, verbatim, and it is one sentence on purpose: ``` this room writes to the workspace without asking, so it will not detach: telltale never leaves an agent working while nobody is watching. ``` The remedy is a second line rather than a longer sentence, because §9.17's tell is that a refusal without a remedy is this room's stated defect and a run-on sentence is not a remedy: ``` the room is still here and still yours. open it with `telltale council --host --read` to get a room you can leave. ``` **Three things this ruling is NOT.** It is **not** a claim that the write posture is unsafe. The room writes by default and that ruling stands ([§7.28](#s7-28)'s parent, and the `--read` opt-out's own reasoning). What changes is only whether the operator may walk away from it. It is **not** a supervisor. Nothing watches the room for the operator, nothing re-approves anything, and nothing self-terminates. The refusal is a refusal. It is **not** the option the costing recommended. The scope ladder that produced this rung recommended allowing the detach and reporting afterwards what happened while nobody watched — turns dispatched, tool calls approved, files touched. **That was reasoned and it is overruled by the owner** (2026-09-01). The report it proposes is a record of an act that already happened, and the product's whole claim is that it does not act unwatched; a receipt is not consent given in advance. The recommendation is recorded here so a later session does not re-derive it as new. **Enforced in the HOST, never in the client.** The host is the process that would keep running, so it is the process that must refuse. A check in the client alone would be a check a second client could simply not make. `TestAWritingRoomRefusesToDetach` pins it against the host, and `TestAReadRoomDetaches` is its positive control — without that pair, a refusal that refused everything would pass. #### `telltale council kill` — the fifth surface, and it is the room's own executioner Sub-noun before the flag set, matching `ls` and `host` and matching `hook cursor` / `events view` / `otel grok`. It takes no arguments, for §7.27's reason: none of the room's flags apply. It reads `host.json`, runs the same three-part probe rejoin runs, and then calls `TerminateProcess` on the pid. **That is deliberately the blunt instrument and not a shutdown frame.** Three reasons: 1. It is the mechanism §7.28 already MEASURED. The room job carries `KILL_ON_JOB_CLOSE`, the host holds the only handle, and `TerminateProcess` is what `taskkill /F` does — so `kill` leans on the property `TestAHardKilledHostReapsEverySeat` proves rather than on a second path that would need its own proof. 2. A shutdown frame needs the pipe, and the pipe may be held by a client. A `kill` that could not end a room BECAUSE somebody was in it would be useless for the case it exists for. 3. The word says so. `kill` is honest about ending agent processes mid-turn. It **refuses** rather than guessing when the probe disagrees with the file: a pid that is not a telltale process is reported and not terminated, and a stale `host.json` is removed with a sentence rather than acted on. Killing a pid the file names and the probe cannot confirm is the one failure this command could make that nothing could undo. #### `telltale council ls` gains a live-host section and stays the SIXTH READER §7.27's contract is unchanged and is re-stated here because the temptation to break it is exactly what a new section invites. - It **still writes nothing** — including no cleanup of a stale `host.json`. A reader that tidied would be a writer. **The ROOM removes that file instead**, on the died path, and the asymmetry is the point: council is already the ratified writer of that directory, and `host.json` is a file this same feature added, so a room that removes a record of a process it has just proved is gone is tidying its own state. `TestCouncilLsLeavesAStaleHostFileAlone` pins the reader's half. - It **still binds nothing and connects to nothing.** The liveness probe asks whether a NAME exists; it does not open the pipe. That is why the probe had to be built the way it was: a dialling probe would have made `ls` capable of ending the room it was listing. - It **still spawns nothing.** - It **still relays no quota**, so it holds the boundary with the same one item spare. What it prints is what the probe measured, and the three states stay apart in the same way the seat states do: a live host, a `host.json` whose process is gone (named as stale, with the remedy), and no file at all. #### The TUI has no host to leave, so no key was added and no golden moved **`telltale council` runs the single-process room, and it always has.** §7.28 built the host beside it and wired no daily path to it, so there is no host for a key in the TUI to detach from. A `d` in that room would have to either do nothing or lie, and a hint on the help panel for a key that does nothing is the honest-gauge failure this project exists to prevent, spent on its own surface. So the way into a hosted room is **`telltale council --host`**, and it is an opt-in flag rather than a change to the daily command: | command | what happens | |---|---| | `telltale council` with no live host | the single-process TUI room, unchanged | | `telltale council --host` | a hosted room, drawn by §7.28's plain client, which you can leave | | `telltale council` with a live host | **rejoins it** — §7.28's plain client again, with the rejoin notice | | `telltale council` with a dead host's file | prints the died notice, then opens the ordinary TUI room, which rebuilds from `room.json` (§9.52) | | `telltale council` with a live host somebody else is in | **refused**, naming `telltale council kill` | The last row is a refusal rather than a fall-through, and that is the load-bearing one. Falling through would open a SECOND room over the same workspace on the same saved session ids, which is two rooms rebuilding one conversation — worse than any refusal. **The renderer is §7.28's plain-text `Render`, and that is stated rather than hidden.** §7.28 calls it "a legible proof that the wire carries a whole room — not a second council TUI", and making council's own `Model` the client's renderer is still the next slice. So a rejoined room looks different from the TUI room, and the client's banner says so in its own words rather than letting the operator discover it. **The client is line-oriented, deliberately.** A blank-prompt line dispatches a turn; `/detach` leaves; `/quit` ends the room; `/interrupt` abandons the turn in flight. Keys would need the TUI, the TUI needs `Model` on the wire, and that is the slice this rung is not. #### What this rung deliberately does not do **It persists no transcript.** Unchanged, and it is the rule this whole area is built around. A rejoining client is handed the host's CURRENT projection over the wire, not a replay from disk. The host holds the conversation in memory and it dies with the host. `resume.go` ruled this for the same data, §7.28 restated it, and a second client arriving does not make a second copy any more acceptable. **It does not bound host memory.** §7.28 named this and said a turn ceiling was owed before detach shipped. It is **not paid here**, and saying so is better than a ceiling picked without a measurement: nothing has yet measured what a room accumulates per turn, so a number here would be a guess presented as a limit. A detached room that runs for a week grows, and the honest statement is that nobody has measured how fast. **It adds no auto-start, no service and no supervisor.** §7.28's third mitigation is unchanged and this rung is where it would have been tempting. The operator starts the host and the operator ends it. **It does not make a stale host self-terminate.** §7.28 refuses that outright — a detached room that dies on its own is precisely the failure the operator cannot see — and detach is the rung that makes the refusal cost something. What answers a stale host is discovery, and discovery shipped first, on purpose. **It is Windows-only, like the host it extends.** The Unix equivalent carries the asymmetry `PARITY.md` already owes: on macOS a process group does not bind lifetimes, so a `kill -9` on a host LEAKS every seat there. A detached host makes that asymmetry worse rather than equal, which is a reason to measure it before building rather than to build and label. *Amended 2026-09-02: no longer Windows-only — [§7.30](#s7-30) builds the Unix transport, measures the asymmetry on Linux, and records it in `PARITY.md`.* **No vendor has been dispatched to through a host, detached or not.** Every seat spawn in this package's suite is stubbed, by design — the spawn guard exists to stop a test spending a real turn — so neither CI nor a session can close that debt. It is the operator's to pay, and `STATE.md` carries it. ### 7.30 the room outlives the terminal on macOS and Linux too (2026-09-02) **What this section rules.** [§7.28](#s7-28) built the host on a Windows named pipe and a Job Object and withheld the Unix side until one measured asymmetry reached `PARITY.md`. [§7.29](#s7-29) exposed detach on the same terms. This section builds the Unix transport and the Unix containment, so that `telltale council --host`, `/detach`, rejoin, `telltale council ls` and `telltale council kill` do on macOS and Linux what they do on Windows — and it states, rather than labels, the one thing they do differently. The owner's daily machine is a MacBook, so before this the crew tool's central feature refused on the machine it was for, and `council ls` there printed `telltale council --host opens a room in one` as the remedy for an absence: a true absence with a false remedy. That sentence is unchanged and is now true on every platform; the fix was the transport, not the words. **What a reader must not read into it.** No frame was added to the protocol, no field to `host.json`, and nothing about what reaches disk changed: the room's conversation still lives in host memory and dies with the host. `protocol.go` is untouched. The seats' own containment (`runner/proc_unix.go`, a process group per seat) is untouched too, and the reason is the next heading. #### The containment is the host's SESSION, and it is not a Job Object The Windows room job has one property no Unix primitive has: `JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE` binds every seat's **lifetime** to the host's last handle, so a host that dies by any route takes its seats with it, and a child cannot leave the job by anything it does. `runner/proc_unix.go` measured the Unix side on 2026-08-17: a process group NAMES a set of processes and does not bind their lifetimes — SIGINT, SIGTERM, SIGHUP and SIGKILL on the parent, and the `Setpgid` child survived all four. So the Unix containment is built from the two facts that do hold, and `roomjob_unix.go` states the difference in as many words: - **Membership is inherited and cannot be lost by accident.** The client starts the host with `Setsid`, so the host leads a session of its own with no controlling terminal; every seat, and every shell a seat starts, is born carrying that session id, and it changes only when something in the tree calls `setsid(2)` itself. **That is the one escape, and it is the measured difference: a child that calls `setsid` leaves the session and nothing here can see it go. A Job Object child cannot leave.** Nothing council dispatches is known to do this; a vendor that daemonised its tool calls would be a `PARITY.md` finding. - **Lifetime is not bound, so the host reaps.** `NewRoomJob` installs a SIGTERM/SIGINT handler and ignores SIGHUP; `Serve` turns a caught signal into `Shutdown`, which kills every registered seat through its per-seat group and exits on the ordinary path, removing `host.json`. `telltale council kill` sends SIGTERM to the host, SIGTERM to every other process in the host's session at the same time, waits a bounded grace for the host, and then sweeps the session with SIGKILL — so a host that could not run its handler still has its seats ended by the command that ends it. On a **dead** host's stale file, `kill` sweeps the dead session too and prints how many processes it ended, because a dead leader's session id still names its orphans and the kernel does not hand that pid out again while any of them carries it. **Why the session and not the host's process group.** A room-level container must hold every seat; a per-seat container must kill ONE seat's whole tree on an interrupt or an eviction without touching the others. `runner/proc_unix.go` builds the second from a process group per seat, and it is the containment the ordinary room measured and relies on. Seats that joined the host's own group would give that up: `kill -TERM -pgid` on one seat would signal every seat. A session sits one level above the process groups and holds all of them — the same nesting the Windows side has, the room job over the per-seat jobs, in the platform's own hierarchy. **What is left uncovered, stated once here and once in `PARITY.md`:** `kill -9` the host and walk away, and the seats keep running — holding sessions and spending quota — until the next `telltale council kill` sweeps the dead session. `TestASigkilledHostLeaksItsSeatAndKillSweepsIt` pins BOTH halves: the leak, so the asymmetry cannot be forgotten, and the sweep that answers it. #### The transport: a socket in the council directory, and two locks beside it The socket is `~/.telltale/council/telltale-council-.sock`, in the directory council already writes, and `Listen` refuses a directory whose mode admits another account (`chmod 700` is the named remedy). The node is 0600, and both connect directions read the peer's uid and pid from the kernel — `SO_PEERCRED` on Linux, `LOCAL_PEERCRED` plus `LOCAL_PEERPID` on macOS — so the program on the other end is trusted exactly as far as the filesystem trusts it, which is §7.24's boundary unchanged. A browser has no URL scheme for a filesystem path any more than for `\\.\pipe\...`, so §7.28's transport ruling stands; the gate that enforces it (`TestTheHostBindsNoPort`) moved from "this package does not import `net`" to "this package reaches `net` for Unix domain sockets, by name, and for nothing else", and it walks every source file on every platform's build tags. **One client at a time is enforced by the host here, not by the kernel.** A pipe with one instance refuses a second open; a Unix socket accepts into a backlog whether or not anybody will ever `accept`. So the listener runs its own accept loop for its whole life, hands a connection to `Accept` only when `Accept` is waiting, and answers everything else with `KindRefused` at once, naming the pid that holds the room. A second client is told "one client at a time" instead of sitting in a backlog until the first one leaves. **The probe cannot connect, and the socket alone cannot answer it.** `WaitNamedPipe` asks "is an instance available" and opens nothing. A socket path answers only "does a node exist", and a node outlives a hard-killed host; a probe that CONNECTED would be accepted the moment the host was free, read as a client with no handshake, and would end the room it was listing. So two facts are carried by `flock(2)` on two zero-byte files beside the socket, and a probe reads each by trying a SHARED lock without waiting — it fails at once if the host holds the exclusive one: - `.lock`, held for the listener's whole life: somebody is listening. A second `Listen` fails on it with the sentence `FILE_FLAG_FIRST_PIPE_INSTANCE` earns, and a stale socket node whose lock is free is unlinked safely. - `.held`, held while a client is attached: the room is held. `Dial` reads it before connecting, so a client never connects into a held room; the host's refusal is the second arm. A flock is released when the process exits by any route, SIGKILL included, so a dead host never reads as listening. That is the property a pid in a file cannot have, and it is why the file says WHAT and the lock says WHETHER. **flock, and not fcntl record locks — measured.** The first cut used `F_SETLK` byte ranges on one file, two facts on one node, and every in-process test failed: fcntl locks belong to a PROCESS, so a probe inside the host's own process read its own lock as free, and closing the probe's descriptor released the listener's locks with it. flock belongs to an open file description, so a second open in the same process conflicts exactly as another process would. The cost is a second zero-byte node. **The pid says WHO, in four readings.** `kill(pid, 0)` alone answers "is there a process I may signal" and nothing else, and a pid is reusable. `verifyHostProcess` asks four things: the pid is alive and this user's; it leads its own session (`getsid(pid) == pid`, which every host does); its image name is this binary's (`/proc//exe` on Linux, `kern.proc.pid`'s `p_comm` on macOS, sixteen bytes and compared as a prefix); and it started no later than `host.json` says (`/proc//stat` field 22 against `btime`; `p_starttime` on macOS). A recycled pid fails on the second, third or fourth, and an identity that cannot be read is "not a host", never "gone". #### What is measured, and where **Linux, in this repository's own CI.** The `race` job runs the whole suite on `ubuntu-latest`, and every claim above is a test there: `TestOneClientDrivesAHostedRoomEndToEnd` and `TestAHostAcrossAProcessBoundaryTakesAClient` (the transport, the peer check, the one-client refusal by both arms, a bare disconnect ending the room); `TestADetachedHostOutlivesItsClientProcessAndKillReapsEverySeat` (a client process detaches and exits, the host survives, a second client rejoins and receives the projection, and `kill` reaps a seat reachable only through the session); `TestAProbeDoesNotConsumeTheHostsSocket` (five probes, then a connect, then `PipeBusy` with a client attached); `TestAStaleHostFileIsNeverMistakenForALiveHost` (a pid that ran and exited, reported dead by `Probe` and left on disk, removed by `kill`); `TestASigkilledHostLeaksItsSeatAndKillSweepsIt` and `TestTheIdentityCheckRefusesARecycledPID` (the containment and the identity, on stand-in processes). The built binary was also driven by hand on Linux on 2026-09-02 — `--host --read` with `/detach` on stdin, `ls`, a rejoin, `kill`, and a `kill -9`'d host reported stale by `ls` and left alone — and the host's `ps` line read `SID == PID`, `PPID 1`, no TTY. **macOS is measured by the test suite, and not yet by hand (amended 2026-09-02).** `peer_darwin.go` and `identity_darwin.go` compile under `GOOS=darwin` for both architectures, and every call they make exists in x/sys v0.47.0. The `darwin` CI job's first run on Apple Silicon (the crew PR, run 33657422294) ran this package's `_unix_test.go` suites in-process on the runner: host, client, detach, rejoin, one-client refusal, kill and the stale-file path all passed there, which exercises `LOCAL_PEERPID` on every dial. What that first run also found: a sandboxed home under `/var/folders///T/` put the socket path at 105 bytes, past macOS's 104-byte `sun_path`, and the host refused while the client heard nothing for ten seconds; `PipeName` now retreats to a short per-uid directory when the council directory's path would not fit (pipe_unix.go). Still owed: the built binary driven by hand on a Mac through `--host`, `/detach`, `ls`, rejoin and `kill`, and `p_comm`'s sixteen-byte truncation against a binary named `telltale` under a real `kill` sweep; until then `PARITY.md`'s row says "suites measured, hand cycle owed". #### What this rung deliberately does not do **It does not put the seats in the host's process group.** The heading above says why, and `PARITY.md`'s row names the consequence: `council kill` reaps through the session, and a host signalled by anything else reaps through its handler. **It does not make the host survive its own SIGKILL.** Nothing on this platform can, and the sentence that says so is in three places on purpose — `roomjob_unix.go`, `PARITY.md`, and the test that pins the leak. **It does not add a field to `host.json`.** The socket path goes where the pipe name went. The two lock files and the socket node carry no content; `TestATurnIsNotPersistedAnywhere` asserts the lock files are zero bytes and walks the directory they live in. **It does not change the protocol or the client.** `Client`, `RunClient`, `Rejoin`, `StartHosted`, `KillHostedRoom` and `ls`'s words are the same code on every platform; the platform split is entirely below `Listen`, `Dial`, `ProbePipe`, `verifyHostProcess` and `killProcess`. ### 7.31 the hosted room draws with the room's own columns, and you can leave it from there (2026-09-02) **What this section rules.** [§7.28](#s7-28) built the host and drew its room with a plain text client. [§7.29](#s7-29) exposed detach on that client and said, in its own words, that the TUI had no host to leave. This section makes council's own `Model` and `Render` the renderer of a hosted room. A room that outlives the terminal now draws the five columns, the posture badge on each column, the tool trace, the turn page, the panes ([§9.51](#s9-51)) and the help panel, the same way the single-process room draws them. It also gains `/detach` inside the TUI. The plain client stays as the degraded path, and the paragraph below says when it is used. **What a reader must not read into it.** No golden file of the single-process room moved. Every hosted frame is pinned by a new golden under `testdata/golden/hosted-*.txt`. No transcript content reaches disk: `room.json` and `host.json` stay numbers and keys. No new spawn path exists. The client still reaches the host through `startHostedRoom` and `joinHostedRoom`, which the spawn guard already wraps. #### Design point 1: what crosses the wire Three shapes were possible. 1. **`council.State` on the wire.** The host would own the canonical `State` and the client would only render it. Refused. `councilhost` cannot import `council`. `boundary_test.go` pins that direction, and this section is the reason the direction exists. `State` also carries facts that belong to the client's machine and to the reader: the posture claim for each vendor, the focus, the scroll, the pane sizes, the help page, the draft. A host that owned those would own the reader's eyes. 2. **Events on the wire.** The host would forward every `runner.Event`, and the client would fold them. Refused. A rejoining client needs the whole current room, and the host holds no event log. A replay from the host would be a second copy of the conversation, which `resume.go` refuses for this data. 3. **The host's projection, widened.** The host's `Room` is already the fold. It already travels whole on every frame, so a rejoin already receives the whole current room as its first frame. This section widens `Room` to carry every fact `Render` reads about a seat, and package `council` builds a `State` from it with one pure function. **Ruling: shape 3.** One process owns the conversation, and that process is the host. One process owns the view, and that process is the client. The function that joins them is `stateFromRoom` in `internal/council/hosted.go`. It is pure over its two arguments, the `Room` and the previous `State`. It copies the conversation from the seat and keeps the view from the previous column with the same vendor: scroll, follow, last focus, the quota reading the client read from its own relay. `TestHostedStateIsPureOverTheRoom` pins the purity. What the seat carries now, and did not before: | field | why `Render` needs it | |---|---| | `Turn`, `Prompt`, `Quoted` | the column echo, the band, the turn page, `PageTurns` | | `History` | the transcript, the turn page, the act ledger | | `Acts` as `{Text, Status, Detail}` | the trace's outcome marks, which a flat string lost | | `Started`, `Ended`, `Elapsed` | the header clock, the finished column's figure, the inbox | | `CostUSD`, `CostSession` | the cost cell and its `session` word | | `Settling` | `done · exiting` on a batch seat, [§9.33](#s9-33)'s linger made visible | | `Skipped`, `NoteDetail` | the skip line and the two-line card | `Acts` was a list of rendered strings. It is structured now, because a string had already folded the outcome into words and the TUI draws the outcome as a mark. The plain client renders the same structure through `actLine`, so its output did not change. **Redaction stays a client choke point.** The host does not redact. The client passes every body, act and note through `Redact` and `sanitize` in `stateFromRoom`, on the way into `State`. That is the same rule the single-process room applies in `applyEvents`, applied to a whole string instead of a stream. **The frame cost is named.** A frame now carries every seat's history. The 50 ms coalescing tick bounds the rate, and `maxFrame` bounds the size at 8 MiB. Nothing has measured a room against that ceiling. The host's memory ceiling is still owed ([§7.28](#s7-28)), and this section makes the frame grow with it. #### Design point 2: the crew is the model [§9.54](#s9-54) made the seats busy one at a time. The host still refused a second dispatch while any seat was running. That was the committee, and this section removes it from the host. - `KindDispatch` carries `Seats`. The client resolves the route (`@codex`, `-@claude`, `@all`, the default) against its own `State`, and sends the explicit vendor list. An empty list still means every drivable seat, which is what the plain client sends. - The host refuses per seat. A busy seat is named on the room's notice and the idle seats in the same list still go. The refusal reads the seat's phase and its `Settling` flag, under the room's lock, in the same call that marks the seats it accepts. The old `watchTurn` poll is gone. - A drivable seat the dispatch did not name is marked `not addressed in turn N`, with `Skipped` set, exactly as `sendTurn` marks it. - `KindInterrupt` carries `Seats`. The client's `ctrl+c` interrupts the focused seat when it is busy, every seat when the focused seat is idle, and ends the room when nothing is busy. That is `viewKey`'s own rule, and the footer's cancel cell names which. - The header's `turn N → route · K in flight`, the inbox strip and the seat numbers all compute from `State`, so they work over a hosted room without a second implementation. #### Design point 3: keys, and what a key must not do **`/detach` is a room verb, not a key.** A single keystroke that walks away from five agents is a keystroke that can be pressed by accident. The verb is bare-only, on `/read`'s rule: a sentence that opens with the word is prose. **In a single-process room `/detach` refuses.** The notice says `this room has no host to leave. open one with telltale council --host`. A verb that did nothing would be the honest-gauge failure spent on the room's own surface. The verb sits outside `roomVerbs`, because the refusal line that lists that table fits its narrowest room with no cell to spare (`refuseUnknownCommand`); the hosted help panel teaches it instead. **`q` and `ctrl+c` still END the room.** They send `KindShutdown`, the host kills every seat, and the closing line says so. Nothing about leaving happens on a key. **The footer did not change.** Every hint on the mode line is the same in a hosted room. The two places that say `hosted` are the header, which gains the word `hosted` after the posture word, and the composer's border label, which gains `hosted pid N`. Both are words, so they survive `--ascii` and `NO_COLOR`. The help panel's `/cd` row becomes a `/detach` row in a hosted room, and its `ctrl+c / q` row says that `q` ends every seat. The panel keeps its 16 rows. **What is refused inside a hosted room, with one notice each.** `/cd`, `/seat`, `/unseat`, `/read`, `/write`, `/arena`, `/flow`, `/hand`, `/adopt`, `/retry`, `/trace`, `c`, `u`, `x`, `o`, `a`, `s` and `ctrl+r`. Each of them changes a seat's process, tree, thread or gate, and the client holds none of those. The rebuttal (`ctrl+r`) is owed: a quoting turn hands each seat a different prompt, and the dispatch frame carries one. #### Design point 4: four states, four renders [§4a.1](#s4a-1) rules that `rebuilt`, `survived`, `died` and `refused` never render alike. The TUI keeps every sentence [§7.29](#s7-29) wrote, and adds none of its own: | state | where it renders | the sentence | |---|---|---| | detached | printed after the alternate screen closes | `RenderDetached` | | rejoined | the room's notice line, on the first frame | `RejoinedNotice`, a one-line form of `RenderRejoined` that keeps the clause `nothing was rebuilt, and no session was resumed` | | the host died while you were away | printed before the ordinary room opens, unchanged | `RenderHostDied` | | the host died while you were in it | the TUI quits and prints it | `RenderHostExit` | | refused | the room's notice line | `UnwatchedWriteRefusal`, verbatim, from the host's own `KindRefused` | The refusal is the host's sentence and never the client's. The client sends `KindDetach` in a write room too, because [§7.29](#s7-29) enforces the ruling in the host, and the client draws the reason the host returns. `TestTheFourNoticesNeverRenderAlike` walks the one-line form beside the others. #### Design point 5: the live seat is scoped out `--live` with `--host` is refused before anything opens, with one sentence. A live seat is a pseudoconsole child, and in a hosted room that child would have to live in the host and its cell grid would have to cross the wire on every repaint. That is a second wire format and a second spawn guard, and it is owed. `STATE.md` carries it. #### Design point 6: the plain client stays `councilhost.Render` and `RunClient` are unchanged and still used. The TUI needs a terminal. When `stdin` is not a character device the hosted room falls back to the plain client, so a scripted drive (`/detach` on a pipe, which is how the Linux measurement in [§7.30](#s7-30) ran) still works. `plainClient` in `hostcmd.go` is the one place that decides, and it is a var so a test can force either. The notice tests stay in `councilhost`, because the sentences live there. #### Flags a hosted room refuses, and why `--brief`, `--record`, `--trace` and `--live` are refused with `--host`, one sentence each, before the room opens. The host takes no brief file, holds no recorder and no trace sink, and spawns no pseudoconsole. A flag that was accepted and did nothing would be a promise the room could not keep. `--fresh` is honoured. `--auto` and `--shared-tree` are accepted, because a hosted room already runs ungated and already runs every seat in the workspace. #### The host drives the crew's live shapes through their measured adapters [§9.57](#s9-57) moved codex and grok to request/response live shapes in the registry, and the host marks a conversational seat undrivable. From that merge until this section a hosted room could drive only claude and agy, and nothing recorded it. The host now takes the batch adapter `vendors.LiveFallback` names for such a seat, the way the room retreats to it on a refused handshake (`fallback.go`), and the seat's badge wears that adapter's measured claim. The wire carries `FellBack` so the client can draw it. A conversational seat with no fallback, which is cursor, is still refused in words. The agy seat is not conversational and keeps its live stream shape, which §9.57 lists as unmeasured. #### Known limitations, named - **A hosted room starts every seat fresh.** The host is handed a roster and never a saved session id, so no `Restored` card is drawn and no thread is resumed. This predates this section. It is named here so the missing card is read as honest rather than as a defect. - **The host does not write `room.json`.** [§7.28](#s7-28) said it would, on the room's own schedule. It does not, and no code in `internal/councilhost` reaches `SaveRoom`. A hosted room's session ids therefore never reach disk. `STATE.md` carries it. - **No vendor has been dispatched through a hosted TUI room.** Every seat in both suites is stubbed. The stubbed end-to-end in `internal/council/hosted_e2e_test.go` re-executes the test binary as a real host with an empty roster and drives the `Model` over a real pipe. The owner's second live brief on the built binary is the closer. ## 8. Roadmap (decided 2026-08-01; adoption track added 2026-08-02, ADR-005) Rigor stays the floor; features and front-end craft are the priority axis from here. Each item names its incumbent inspiration and the honest-gauge twist that makes it ours. Sources rule unchanged: a segment ships only when this doc names its source. Read the version numbers below against §1: v1 cuts on the snapshot gates, so an item marked for v1 is not an item waiting on the gauges. They are done. ADR-005 adds a second axis: external adoption is an explicit product goal, and adoptability is a design input rather than a lagging indicator. That does not reorder the feature track below — it adds the adoption track that runs beside it, and one of its items lands *before* v1 is done, which is why this section is no longer titled "after v1". ### Adoption track (ADR-005) ADR-001's sequence stands unchanged — dogfood → eval + design doc → launch post — so these items are ordered by that sequence, not by version number. 1. **Activation slice — runs in parallel with the dogfood window**, i.e. now, not after v1. Four pieces, all packaging rather than capability: prebuilt binaries via goreleaser; one-command install with **scoop/winget first**, per Windows-first (ADR-002); a README hero visual; and a useful **zero-config first frame** — the binary's first run has to show something true without being configured, because an install that lands on an empty screen has spent its only attempt. macOS/Linux binaries are cross-compiled; macOS ships labeled **"smoke-verified on macOS — Windows is the continuously verified target"** and Linux keeps **"built, not verified"**: ADR-002's "no macOS/Linux work until v1" is amended for *distribution only*, the no-porting/no-verification-effort rule stands, and both labels are ADR-001's flagged-limitation pattern applied to a platform instead of a segment. The macOS label is point-in-time and SHA-bearing — the suite, build, statusline smokes, a 53-session live Claude Code read and the HUD all ran on macOS at `052a9d6` (ADR-005, second amendment) — while CI still runs `windows-latest` only, and five of the six adapters have still never met a live macOS corpus. The README positioning line — *"one local HUD for every coding agent you use"* — lands **with** this slice and deliberately not before it: a positioning claim that arrives ahead of a one-command install is a promise the reader has no way to act on. 2. **The launch post tests one hypothesis: cross-harness visibility** — do multi-harness power users want one honest local HUD across the agents they already run? That, and only that, is what the launched product contains. The signal that answers it is evidence someone actually *ran* telltale — a version-bearing bug report, a real-session screenshot, a PR grounded in running it, package-manager feedback, or an unsolicited statement of use. Engagement without that (a comment, a question, a hot take) answers a different question, and is read as such. **Amended 2026-08-15: the hypothesis is room-led, because the product is.** The wording above was written 2026-08-02 and names the HUD, four days before the council-is-the-product ruling (§1); the post that goes out leads with the room, so the signal above is read as evidence about cross-harness visibility of the ROOM — one brief answered by five vendor CLIs side by side, with the gauges as the infrastructure under it. The signal test itself is unchanged, and so is item 3's exclusion. 3. **Needs-input / blocked / done state is the first post-validation feature** — the attention-routing job, and the reason the product is positioned the way it is. It is built where the vendor seams already support it: Claude Code hooks, Codex notify events, agy's `agent_state` (observed live transitioning `tool_use` → `idle`, §3.8), and Cursor Hooks (documented and versioned — §3.9, and the reason the Cursor adapter's `status` field is deferred rather than mapped). It is judged on its own terms rather than on the launch's: the launch explicitly does not claim this ground, so its result is evidence about cross-harness visibility only, and neither validates nor falsifies a capability the launched product did not contain. 4. **The agy disk-seam re-survey RAN the same day this track landed — verdict: OPEN** (§3.8 re-survey block; prompted by ccusage issue #1402). Both claimed surfaces verified against the local 1.1.9 corpus: the advertised transcript.jsonl is real (ADR-004's watch item resolved — the first survey's "never written" was wrong at the same version), and `gen_metadata` token counts decode with a self-checking arithmetic identity. **The next adapter work item is therefore the agy HUD adapter itself** — transcript-first (Name/Workspace/LastActivity/liveness scaffolding from plain JSONL), `gen_metadata` only for Model + token counts, honoring §3.8's build cautions (WAL sidecars, stale summary index, PII fields, assert-the-identity). agy stops being a statusline-only vendor when that ships, and telltale ships the lane's only Antigravity HUD adapter. Liveness and subagents stay structural-only until observed live. — **BUILT the same day** (decisions/006, `internal/adapter/antigravity` + `internal/sqlite`, §3.8 "Adapter built"): four reported fields, zero new dependencies, the token identity asserted at read time and holding 16/16 on the live corpus. agy is now a HUD vendor; the two deferrals stand. 5. **Cursor (Composer) is the fifth vendor and the SIXTH HUD lane** — surveyed and **BUILT the same day** (§3.9, decisions/007, `internal/adapter/cursor`). It is the first vendor to persist its own context percentage, and the first whose store holds live credentials, so the adapter's most load-bearing property is its read allowlist. Two watch items follow it and neither is blocked on effort: **Cursor Hooks** (cursor.com/docs/hooks — a documented, versioned payload carrying `conversation_id`, `model`, `workspace_roots`, `transcript_path`, and context numbers on `preCompact`) is where item 3's needs-input signal should come from for this vendor, rather than reverse-engineering `status` out of the store; and the `cursor-agent` CLI keeps a separate store that is not installed on the survey machine and stays unverified and out of scope until it is. — **The second watch item CLOSED 2026-08-29.** `cursor-agent` is installed on this machine now, its store was surveyed, and `internal/adapter/cursor` reads its per-session manifest. §3.9's 2026-08-29 addendum carries the field map, the `schemaVersion` pin and the composition argument. The Cursor Hooks watch item is untouched and still open. #### Packaging decisions (settled 2026-08-08; §6.5 closed here) Adoption item 1's first piece — prebuilt binaries and a one-command install — is **built and not yet fired**. `.goreleaser.yaml` and `.github/workflows/release.yml` exist; no tag does. §6.5 deferred distribution naming "to packaging time", and this is it, so the rulings are here rather than in the open-questions list. **Tag day is one command.** `git tag vX.Y.Z && git push origin vX.Y.Z`. A `v*` tag is the release workflow's only trigger; merging to main releases nothing. The workflow then runs the repo's own gate, builds, and stages a **draft** release. The runbook lives in [packaging/README.md](../packaging/README.md) — this section is the argument, that file is the procedure, and neither restates the other. 1. **Targets, and each one's label.** Four: `windows/amd64`, `darwin/amd64`, `darwin/arm64`, `linux/amd64`, CGO off. The labels above are binding on the release notes and are printed per download in the release body: Windows **continuously verified**, `darwin/amd64` **smoke-verified on Intel macOS** (point-in-time, `052a9d6` — the Mac that ran it is Intel, and the label says which arch was under the smoke rather than "macOS" flat), `linux/amd64` **built, not verified**. `darwin/arm64` is the one addition to the ADR-005 list and it takes the Linux label verbatim: **built, not verified**. Shipping it was weighed against withholding it — most Macs are Apple Silicon, so refusing to build it replaces a labelled binary with no binary at all, which serves nobody and teaches nobody anything. The label is the whole claim. `windows/arm64` and `linux/arm64` are not built: no verification story and no known user, and an unlabelled binary nobody has run is the packaging form of a rendered guess. 2. **Names: `telltale` everywhere.** §6.5 named `telltale-hud` as the fallback if the bare name collided. It does not — checked at packaging time against the scoop `Main` and `Extras` buckets and against `microsoft/winget-pkgs`, both clean — so the fallback stays unused and the scoop app is `telltale`. winget needs a publisher-qualified id and gets `sanlee-ys.telltale`, the GitHub-handle convention `junegunn.fzf` and `ajeetdsouza.zoxide` already use. npm remains skipped, not renamed: the bare name there IS taken (an unrelated option parser), and §6.5 already ruled npm optional for a Go binary. 3. **The scoop bucket is in this repo, `bucket/`.** Rejected: a second repo, `sanlee-ys/scoop-telltale`. The deciding cost is a credential rather than convenience — goreleaser pushing a manifest into a *different* repo needs a cross-repo PAT held as a release secret, while pushing into its own needs only the workflow's built-in `GITHUB_TOKEN`, which the release already holds to upload artifacts. A project whose stated posture is that it reads no credentials should not mint a long-lived write token for one JSON file. `scoop bucket add` accepts any git repo carrying a `bucket/` directory, so the user's command is the same length under either choice, and the second repo buys nothing but a second thing to keep alive. 4. **The release is a DRAFT and stays one.** goreleaser's job ends at "artifacts and notes are staged"; publishing is outward-facing and is a human action. The honest consequence, recorded rather than glossed: the scoop manifest is committed in the same run, so between that commit and the publish click, `scoop install telltale` points at a URL that 404s. It fails cleanly and installs nothing — but the window is real, and the answer is to publish promptly rather than to pretend it isn't there. `skip_upload: auto` keeps snapshots and prereleases out of the bucket entirely. 5. **The release runs the repo's own gate, called and not copied.** `release.yml`'s first job is `uses: ./.github/workflows/ci.yml` — vet, the suite, the build and the three binary-level smokes, on `windows-latest`. A release that skips the gate is a false green; a release running a hand-maintained second transcription of it is a subtler one, and that is the only change `ci.yml` took (a `workflow_call` trigger). goreleaser itself runs on `ubuntu-latest`: with CGO off the cross-compile is host-independent, so the build host is a cost question, and what makes the Windows artifact trustworthy is the Windows gate that already passed. 6. **Changelog: plain `git log`, ascending, ungrouped** — not goreleaser's default Conventional-Commits grouping, and explicitly not `use: github-native`. This repo's commit voice is a lowercase sentence describing the behavior change (CLAUDE.md), not a `feat(x):` label, so CC grouping would file every commit under "Others" beneath three empty headings; squash-merge already makes those subjects the PR titles, which is the changelog anyone would write by hand. `github-native` is rejected for a harder reason: it replaces the entire release body, which would delete the platform-label table — the part of the release that has to be true. 7. **Not packaged, and why.** No Homebrew tap: macOS is smoke-verified on Intel only, and a tap resolves just as happily on Apple Silicon, which would launder "built, not verified" into "supported". No `.deb`/`.rpm`: a distro package is a support claim Linux has not earned here, and the tarball carries the label a package would drop. No npm, per (2). winget is packaged but **not automated** — submission is a pull request against a Microsoft-owned repository, and a bot opening PRs on someone else's repo every tag is a different thing from cutting a release. The manifest draft (schema 1.12.0, `zip` + nested `portable`, `windows_amd64` only) and the submission flow are in [packaging/winget/](../packaging/winget/). **Amended 2026-09-02: the tap ships.** What the refusal above was waiting for was a measurement on Apple Silicon, and `ci.yml`'s `darwin` job is one: it builds the binary on `macos-latest` (arm64) on every commit and runs the suite, `doctor`, the statusline smokes and `council ls` there, with the Windows job's honesty assertions. The tap lives in `Formula/` of this repository, on (3)'s credential argument for the scoop bucket, and goreleaser rewrites the formula at each tag. It is a **formula and not a cask**: Homebrew fetches a formula's archive with curl and writes no `com.apple.quarantine`, so the unsigned binary runs as installed, while a cask arrives quarantined and goreleaser's cask documentation answers that with an `xattr` post-install hook it labels a Gatekeeper bypass. goreleaser v2.17.1 deprecates `brews` in favour of casks, and `goreleaser check` now exits 2 naming that one key; the pin is exact, deprecated keys are removed only at a major version, and `.goreleaser.yaml` carries the measurement. Still not done, each with its reason: **homebrew-core**, whose acceptance rules ask for a notable project (75 GitHub stars among the signals) and nothing here measures that number; **winget**, per this item; **signing and notarization**, per (8). packaging/README.md is the runbook. **Amended 2026-09-11: winget carries `0.2.0`.** The manifest drafted above was submitted by hand, as this item required, and merged into `microsoft/winget-pkgs` on 2026-09-10 as #417671 (`sanlee-ys.telltale`, the publisher-qualified id from (2)). Nothing about "not automated" moves: `v0.3.0` was published 2026-09-09 and has no manifest there, because each version is its own pull request against a Microsoft-owned repository and a session does not open those. So the winget channel lags the release by exactly the submissions the owner has made, and the README says which version it delivers. No `winget install` has been run against the community source; the first one is a PARITY.md entry, the same debt the tap carries. 8. **Not signed, and the decision is the owner's** (recorded 2026-08-16). No artifact carries a signature. `.goreleaser.yaml` declares no `signs` block and `release.yml` holds no signing secret, so the claim is checkable from the two files rather than asserted. This is a **gap that is stated, not a gap that is closed**, and it takes the same treatment as an unverified platform: the label is the whole claim, and it now appears in `SECURITY.md` and in the README install section. Windows ships without an Authenticode signature, and scoop and winget deliver that same binary. The macOS archives are unsigned and not notarized, so Gatekeeper refuses one that a browser marked with `com.apple.quarantine` — **that path is unmeasured here**, because the macOS smoke ran a binary built on the Mac itself rather than a downloaded archive, and the sentence says so wherever it appears. `checksums.txt` stays the verification this release can honestly offer: it proves the archive is the one the workflow produced, and it proves nothing about who produced it. **Why it is not built:** signing needs a code-signing certificate, or an Apple Developer account and a notarization credential, held by the owner and stored as long-lived release secrets. The credential is the deciding cost here, the same way it decided the in-repo scoop bucket in (3) — except that argument cannot be won by a config choice this time, because no arrangement of the workflow produces a signature without an owner-held secret. So this is an owner decision about spend and about identity, not a contributor task, and no contributor should build the pipeline speculatively. **Posture surface (added 2026-08-16).** `SECURITY.md` carries the private reporting route, the trust model in the terms of the read/write boundary, and the signing statement above. `.github/dependabot.yml` watches `gomod` and `github-actions` weekly; it is also the watch on the TUI line, because `ultraviolet` is pinned to a pseudo-version that never moves on its own. #### The automatable remainder (added 2026-08-16) The posture paragraph above shipped the two pieces that need no automation. This subsection records the rest. It changes nothing in item 8: signing and notarization stay owner decisions, and no contributor builds that pipeline. **What now runs, and when.** | Check | Fires on | Runner | Fails on | |---|---|---|---| | `ci.yml` | push to main, pull request, release | windows-latest, ubuntu-latest, macos-latest (Apple Silicon, added 2026-09-02) | vet, the suite, the build, the binary smokes, the schema gate, the install-script gate (added 2026-08-18) | | `govulncheck.yml` | push to main, pull request, Monday 07:00 UTC | windows-latest | a reachable known vulnerability | | `codeql.yml` | push to main, pull request, Monday 07:30 UTC | ubuntu-latest | a default-suite alert | | `dependabot.yml` | weekly | none | nothing. It opens a pull request | | SBOM, through syft | a `v*` tag only | ubuntu-latest | a syft failure | | Provenance attestation | a `v*` tag only | ubuntu-latest | an attestation failure | **govulncheck is a monitor, and it is deliberately not a step in the gate.** The gate answers whether a change works, and the tree decides that answer. This scan answers whether the shipped code is vulnerable today, and the Go vulnerability database decides that answer. Two costs follow. The scan reads vuln.go.dev on each run, and the gate makes no network call except the module download, so an outage at that host inside the test job would fail a build that nothing broke. A new standard-library CVE also fails an unchanged tree, and it fails every open pull request at the same time, and no author can correct it inside their own change. A gate that fails for a reason outside the change teaches contributors to ignore it. The scan therefore runs beside the gate under its own name. It still fails on a finding, because a monitor that reports green over a vulnerable binary is the false green ADR-001 refuses. **The scheduled run is the reason this exists.** A scan on a pull request cannot find a vulnerability that becomes public after the merge. The pull request trigger scans a new dependency before it lands, and it also proves the job in the pull request that adds the job. **What the first scan measured, 2026-08-16.** govulncheck v1.7.0, against a database updated 2026-08-14, reported three reachable standard-library vulnerabilities: GO-2026-6090 in `crypto/tls`, GO-2026-6089 in `net/http`, and GO-2026-5972 in `encoding/asn1`. It reported five more that this code does not call. The traces reach `internal/grokotel` and `internal/eventsink`, which are the two packages that run a server. **No module that this repository requires caused any finding.** All eight findings had one cause: the toolchain. go.mod declared `go 1.26`, `actions/setup-go` resolved that to the 1.26.5 in the runner image, and every fix version is go1.26.6. Run 31981243687 on main confirms that CI built with go1.26.5, so this was a property of the shipped binary and not of one workstation. **The fix is a version bump, and go.mod is where the version lives.** The `go` directive now reads `go 1.26.6`. That clears all eight findings, measured both ways: the same scan under a 1.26.6 toolchain reports `No vulnerabilities found`, and the local build then reports `go1.26.6` under `go version -m`. The directive is the single source every job reads, so one line moved the race job, the gate, the release and both new scans together. `check-latest: true` on setup-go was the rejected alternative. It would float the toolchain to whatever patch exists on run day, which is the same objection this document already made to a goreleaser version range. A committed directive moves when a person moves it, and the monitor is what asks for the move. **The SBOM and the attestation take effect at the next tag, and at no earlier point.** `release.yml` triggers on a `v*` tag only, so no merge to main produces either one, and neither is added to a release that already exists. syft writes one SPDX-JSON document for each archive. `actions/attest-build-provenance` then attests the four archives, and a user verifies one with `gh attestation verify --repo sanlee-ys/telltale`. **Provenance is not a signature, and the difference is item 8's difference.** The attestation proves that this repository's release workflow built this archive, from a named commit, on a runner GitHub hosts. It does not prove who the owner is, and it does not prove that the owner vouches for the content. It needs no owner secret, because GitHub mints a short-lived OIDC token for each run. That is exactly why it is buildable here while signing is not: item 8's blocker is a long-lived credential, and this mechanism needs none. **Signing and notarization stay owner-held and unbuilt**, on item 8's terms and for item 8's reasons. **What this subsection did not verify.** The pull request that added these files proved the two scans by running them, and the run log names each job. It could not prove the release path, because a tag is that workflow's only trigger. The evidence for the release path is a `goreleaser release --snapshot --clean` rehearsal on the reference workstation, 2026-08-16, against the pinned goreleaser v2.17.1. It built all four targets, wrote all four archives, reached the SBOM stage, named the document `telltale__windows_amd64.zip.sbom.json`, called `syft`, and stopped with `exec: "syft": executable file not found`. That failure is the wanted one: the stage is reached and configured, it fails loudly rather than passing in silence, and the missing tool is the one thing `release.yml` installs and the rehearsal could not. `goreleaser check` also passes against v2.17.1. The attestation step ran nowhere. **The next tag is the first real proof of both**, and item 8's existing advice applies: rehearse it with an `-rc` tag. `.github/CODEOWNERS` names the owner for every path. SECURITY.md already tells a reporter that one person maintains this project; that file states the same fact where GitHub can act on it. What this does **not** discharge: the README hero visual and the zero-config first frame are the other two pieces of adoption item 1 and are untouched here, and the positioning line still lands with the slice, not ahead of it. *(Ledger, 2026-08-16: the zero-config first frame is DELIVERED — measured on a clean profile first, then built narrow; §7.7's "measured 2026-08-15, then narrowed" subsection is the record. The positioning line landed room-led with the identity rewrite of 2026-08-15. Of adoption item 1, the hero visual alone remains, and the recording chain below is what will produce it.)* #### The recording chain: PowerSession-rs + agg (measured 2026-08-16) The demo tape records with **PowerSession-rs 0.1.16 and agg 1.9.0**, both installed from winget at user scope with no administrator rights (`Watfaq.PowerSession`, `asciinema.agg`). VHS is rejected rather than deferred: it cannot record on this Windows build (charmbracelet/vhs #631, dead since 2025-06) and it renders xterm.js, which is not the Windows Terminal this product targets under ADR-002 — a tape of the wrong terminal is the packaging form of a rendered guess. The chain was proven against the real binary in ascending difficulty, and the hard case passed: `telltale hud` recorded its alternate screen (`ESC[?1049h` and `ESC[?1049l` each captured once), its ANSI palette colour (47 cyan, 19 green, 7 bright-black foreground codes) and its restore. The restore was tested against real scrollback rather than an escape count — a marker line, the TUI, a second marker line — and the rendered final frame carries both markers and no TUI residue. Two limits are recorded with it, because both can produce a tape that looks successful and is not. `NO_COLOR` in the recording shell silently strips every hue, which is how the first capture came back monochrome. And agg's `--rows` re-runs the byte stream at a new size, so it recovers `telltale doctor`'s scrolled-off report (76 lines from a 30-row cast) but does nothing for a TUI that drew to the size it read — the HUD at `--rows 50` leaves twenty empty rows. The tape's geometry is therefore chosen before the recording starts, not after. [packaging/tape/README.md](../packaging/tape/README.md) is the runbook and carries the full measurement; this paragraph is the decision. **No cast or GIF enters this repository**: both captures were inspected, and they carry live session names, workspace paths and the absolute path of every vendor binary on the machine. The tape stays a personal artifact and the repository holds the script that makes it. What remains is not a tooling item — the owner drives the eight beats, because a scripted race would be an invented recording. #### The one-paste Windows install (added 2026-08-18) `packaging/install.ps1` is the third Windows route, and it is the only one that needs nothing installed first: ``` irm https://raw.githubusercontent.com/sanlee-ys/telltale/main/packaging/install.ps1 | iex ``` **It exists because scoop is a prerequisite and winget is not submitted.** Item 1's "one-command install with scoop/winget first" shipped both of those, and both assume the reader already has the package manager. A reader who has neither had two choices before this: unpack an archive by hand, or build from source. A competitor sweep on 2026-08-17 read the same gap the other way round: every lane leader collapses README-read to first-run into one paste, and abtop's README carries this PowerShell shape. That is a reading of their documents, not a measurement of their installers, and it is cited as such. **What it verifies, and what it refuses to claim.** The script downloads the archive and `checksums.txt`, compares the SHA-256 **before** it unpacks anything, and deletes the download on a mismatch. It then prints, in its own output rather than only in a document nobody reads at install time, that the binary carries no Authenticode signature and that the checksum proves what the workflow built and not who built it. That sentence is item 8 restated at the one moment the reader can act on it. The script signs nothing and prepares no signing pipeline: item 8 stands unchanged. **Three refusals are in the script rather than in a note.** A machine reporting `PROCESSOR_ARCHITECTURE` other than `AMD64` is refused by name, because the release builds no `windows/arm64` binary and installing the amd64 one there would be the packaging form of a rendered guess. A tag with no published release fails with the URL that 404'd, not with a bare status code. A `checksums.txt` that names no entry for the archive stops the install rather than skipping the check. **Measured 2026-08-18, against the published `v0.2.0` release.** Windows 11, two shells: PowerShell 7.6.5 and Windows PowerShell 5.1.26100.9168. Both installed `telltale_0.2.0_windows_amd64.zip`, both computed `7a2401aa…33772528`, and that value equals the digest GitHub reports for the asset. The installed binary answers `telltale 0.2.0`, so the release ldflags survive the route. The `irm | iex` shape was exercised as `Get-Content -Raw | Invoke-Expression`, and the calling shell survived it: the script throws and never calls `exit`, because `exit` inside a piped script ends the user's session. The `PATH` branch was driven once with the real user variable and restored byte for byte afterwards: the install directory reached the persisted user `PATH` and the running shell's own `$env:Path`. That trial also measured the one surprise in this route, and it is recorded rather than smoothed over: the directory is APPENDED, so a `telltale.exe` already earlier on `PATH` — a `go install` build, in the measured case — goes on winning. `Get-Command telltale` names the one that runs. Prepending was the rejected alternative, because a script that quietly outranks a binary the operator put there is doing something the operator did not ask for. Two refusals ran end to end and installed nothing: the arm64 refusal, and a `TELLTALE_VERSION=v0.1.0` run against the tag that has no release. **What had no end-to-end live trial was the mismatch refusal**, because driving it needs a host that serves a corrupted archive. Its comparison was measured live instead: the real `checksums.txt` was parsed, a byte was appended to the real archive, and the two hashes differed. The branch that acts on that comparison is three lines below it and was unexercised until the payment below. **Paid 2026-08-19, on the MBP (Intel, macOS, PowerShell 7.6.4): two mismatch trials and one control.** The missing instrument was the corrupt host, so the trial built one. A `python3 -m http.server` on 127.0.0.1 served the published `v0.2.0` `checksums.txt` beside a copy of the real archive with one byte appended (`7a2401aa…33772528` became `8ea9111b…6eb47b8d3c`). The script under test was the shipped file with one recorded change: the `$base` line pointed at the local host. The OS and arch refusals read plain environment variables, so `OS=Windows_NT` and `PROCESSOR_ARCHITECTURE=AMD64` let the real function run on this machine; that shortcut is named below. Both trials refused, 2/2. The thrown message named both hashes and said the download was deleted and nothing was installed. No install directory was created, and no `telltale-install-*` work directory survived. The message itself is the proof that the three lines ran: it is the `throw` line's own text, and it carries the two hashes the lines above it computed. The control reversed the one appended byte and changed nothing else. The same server and the same patched script then installed cleanly: `sha256 ok (7a2401aa…)`, and a placed `telltale.exe`. Stated honestly: this trial redirected `$base` and satisfied the two host guards from the environment, so it measured the mismatch branch and not the GitHub download path or the Windows host — the 2026-08-18 block above already paid those two on the real host. STATE.md no longer carries the gap. **One footgun is recorded because it cost a parse, and the gate holds it.** The file is ASCII only. Windows PowerShell 5.1 reads a BOM-less file as ANSI, so one em dash inside a `throw` produced four parser errors under 5.1 and none under PowerShell 7. The `irm | iex` path decodes UTF-8 correctly and would have hidden this; the download-then-run path would not. `ci.yml` now parses the file under Windows PowerShell 5.1 on every push, rejects any byte at or above 0x80, and rejects an `exit` statement. It never executes the script, because executing it would download a release on every push. The ASCII arm was measured non-vacuous the way the schema gate's mutations are: one em dash appended to a copy, and the gate reported three bytes and failed. **No other channel is reshaped by this.** No Homebrew tap, no npm, no winget automation. Items 2 and 7 rule each one, and a one-paste installer is not an argument to revisit any of them. (The tap arrived on 2026-09-02 by item 7's own amendment, on a measurement rather than on this installer.) macOS and Linux keep the measured `curl` and `shasum` walk in the README, which is the same verification without a script. #### The listing and launch cadence, recorded and not executed (added 2026-08-18) This subsection records strategy that no contributor may execute. It exists because this repository rejects unrecorded strategy, and because the pieces below are owner actions on surfaces outside it. 1. **Directory listings.** `awesome-claude-code` and the neighbouring lists are the lane's standing distribution channel, and an inclusion is a pull request to somebody else's repository. That is the same class of act as the winget submission in item 7, and it takes the same ruling: **a human action, never automated, and never opened by a contributor session.** What lands in this repository is the badge slot in `README.md` and this paragraph. The badge goes in only after the listing merges. 2. **One Show HN per versioned feature, with the maintainer working the thread.** Recorded as the cadence, with one binding limit: item 2 pins the launch to ONE hypothesis, cross-harness visibility of the room, and a serial cadence must not quietly widen that claim. A second post about a second feature tests that feature's own question and is read as such. The first post is chain link 3 in `STATE.md`, and it is sequenced behind links 1 and 2, which are paid. 3. **Publish the run-evidence bar, the method, and the result.** Item 2 already defines the signal that answers the launch hypothesis: a version-bearing bug report, a real-session screenshot, a pull request grounded in running it, package-manager feedback, or an unsolicited statement of use. Nobody in this lane publishes that bar. Publishing it is the launch story only an honest-gauge product can tell, and it costs nothing to tell, because the bar is written down already. **The threshold is an owner decision and is NOT taken here.** The candidate sweep proposed "10 runs in 30 days". That number has no measurement behind it and no ruling, so adopting it would be the invented figure ADR-001 refuses. What is settled is the KIND of evidence, quoted above. What the owner names before the post: the count, the window, and what the result means if it is missed. Whatever is published then cites measured evidence, and it never cites a star count or an install count, because telltale measures neither. **`README.md` carries two slots for this work and no adoption content.** The badge slot holds one badge, the CI result, which is GitHub rendering GitHub's own run and therefore needs no third-party host and cannot go stale. A directory-inclusion badge may join it after that listing merges. A star count, a download count, an install count or a "used by" figure never may: telltale measures none of them, and a third-party render of an unmeasured number is the badge form of a rendered guess. The hero slot is for the animated capture, which stays owner-driven under the recording chain above and under the 2026-08-17 per-frame review ruling. The still SVG hero and the positioning line already landed; neither moves here. Neither track discharges what verification already owes: §3.4's remaining passive-tail items stay open (§3.7's first live Gemini pass ran and passed 2026-08-03), and adoption work does not buy an exemption from them. ### v1.1 — the flagship trio (**BUILT**) 1. **Detail pane** (inspiration: abtop / CASS drill-ins) — **BUILT, §7.11.** Select a row, get an expanded view: quota windows, extras (branch, CLI version, ctx tokens), session id, and — crucially — the session's Diagnostics and degraded-field marks, which v1 carried with no surface. The honesty machinery becomes visible product. Shipped with one thing the spec did not ask for and the design demanded: a `not sourced` line, which makes §4a.1's "can't know" versus "absent now" legible for the first time. 2. **Burn-rate forecast** (inspiration: Claude-Code-Usage-Monitor / codeburn) — **BUILT, §7.12.** The HUD samples the account window's `used_percentage` over its own runtime; the slope is a telltale-measured value, rendered with the `~` derived marker AND its sampling window (`~13:27 · 18m basis`). Never extrapolated from a guessed budget, and below the minimum basis it renders nothing at all. Shipped with two refusals beyond the spec's "minimum sample count/age": no projection past the window's own reset, and none beyond 24 h, because the render is a wall clock with no date on it. 3. **Sub-agent chips** (inspiration: claude-hud's active-subagent display) — **BUILT, §7.13.** Counts recently-written transcripts in a session's `subagents/` sidecar tree (already discovered and excluded from rows in §3.1): a `⑂~2` chip on rows running fan-outs. Corrected against the spec: the roadmap called it "pure sourced data" and it is not — the files are counted exactly, but the recency boundary is an inference, so the field is `CapDerived` and the chip carries the estimate marker. Also in v1.1: **`/` type-to-filter** on name/path substring (CASS's kernel without the embedding search) — **BUILT, §7.14.** New schema field in v1.1: `subagents` (§4a.2), declared `CapDerived` by the Claude adapter and `CapNone` by Codex. `model.Validate` gained a non-negative check for it; nothing else in the schema moved. ### v1.2 — the Windows-native leap - **`telltale notify`**: a third mode on the same binary, fed by Claude Code hook events on stdin, raising a Windows toast when a session needs input or ends a long turn. (agenttray had the idea; nobody has executed it well on Windows.) Read-only posture holds: notify consumes hook payloads, sends nothing anywhere but the OS notification API. - **Statusline context-breakdown bar** (two-line statusline): stacked mini-bar of `current_usage` components (input / cache-read / cache-creation / output) — fully sourced from stdin. ### Later / unscheduled - Themes + segment config file (ccstatusline's adoption driver). - `telltale snap`: one-shot frame render to stdout (pipeable, screenshot-able; already prototyped as a throwaway during v1 verification). - ~~Gemini CLI adapter, once its seam is verified live (§4a.7 example becomes real)~~ — **landed 2026-08-02** (`internal/adapter/gemini`, §3.7; the §4a.7 example is real, with its original sketch kept as the postscript's evidence). ### Deliberately rejected - Cross-device pairing/sync (codeburn): network egress breaks "telltale never writes". - Plan-budget "% of plan" spend meters: the budget is a guess — the exact fabrication this product exists to refuse. - On-disk cost estimation via price tables: inventing dollars from token counts. - A plugin runtime for third-party adapters: it would run a stranger's code inside the process whose whole contract is "reads, never writes, no network, no credentials", so every guarantee here would become a guarantee about someone else's plugin. The drop-file relay ([§7.23](#s7-23)) is the middle path taken instead. ## 9. Council (ADR-008) `telltale council` is the dispatch room: one brief typed once, routed to Claude's control plane by default or to the explicitly `@mentioned` seats, replies streaming side by side. `@all` convenes every seated vendor when independent answers are the point. It exists because the alternative is four terminals and a clipboard. [council.md](council.md) is the room's user-facing guide — the badge vocabulary, the routing grammar, the reading keys, every flag. This section is the record underneath it: what was measured on each vendor, what each finding cost, and what is still unverified. It is the one subcommand that is not a gauge, and the boundary is worth stating precisely rather than hand-waving. §7.8's invariant — no keybinding may mutate vendor state or send anything to a running agent — is **unchanged and unweakened**; it is a rule about the observation surfaces, and council is not one of them. Nothing in the HUD reaches council. The only way in is typing the subcommand. What moved is the *scope* of the sentence in `README.md`, from "telltale never writes" to "the gauges never write", because the old phrasing had become false the moment ADR-008 was accepted and an accepted decision the docs contradict is worse than either option on its own. ### 9.1 What v1 seats, and what it does not Five columns: **Claude Code**, **Codex**, **Antigravity**, **Cursor**, **Grok** (the fifth seat arrived 2026-08-09, §9.39 — this line said "Four columns" for six days after it landed, which is its own small lesson about headline counts in a doc whose sections are the record). Cursor was originally written up here as deliberately absent, on the grounds that `cursor-agent` was not installed. **That was not a judgement call, it was untrue** — the binary was at `%LOCALAPPDATA%\cursor-agent\cursor-agent.cmd` and had been for a month, and a test pinned the false claim in place so the day the world changed underneath it, it went on passing. It is seated now (ADR-008, fifth amendment). What remains true from the original paragraph: the `cursor` binary on PATH is the editor launcher, and council never drives it. The seat is driven on Windows too — the `.cmd` shim turned out to be a nine-line wrapper around a bundled `node.exe`, which detection resolves directly, so no prompt ever meets a shell and §9.3's refusal has nothing left to fire on (ADR-008, tenth amendment). Its posture badge stays `ro:requested`, now contradicted-by-capture rather than merely unobserved: under `--mode plan` the agent was seen dispatching shell tool calls, and `--sandbox enabled` kills the turn outright on Windows, so it is not passed there. This is the inverse of the HUD, where Cursor *is* a built-in adapter (ADR-007) because its seam is on disk rather than behind a CLI. #### The vendors this room does not seat, and the class of evidence behind each (2026-08-15) This list existed only in the owner's private notes, which meant it bound nothing and was one lost file away from being re-derived vendor by vendor. It is here so a future session can tell **"we looked and said no"** from **"nobody has looked"** — two states this project has already ruled must never render alike (§4a.1), applied to a decision rather than to a value. Every entry names the CLASS of its evidence, because they are not the same strength. An unmeasured rejection stated as a measured one is the honest-gauge defect committed in prose, and this doc's own §9.2 is the record of what that costs: the first draft of ADR-008 claimed enforcement for four seats when one had a mechanism. - **Warp** — **REJECTED, measured.** There is no blocking hook anywhere in the product: no surface at which a tool call is held while something else decides. So the fleet's guard requirement cannot be satisfied here, and that requirement is not council's own — council gates the seats it spawns (§9.8), while agent-ops ADR-012 requires each vendor to be guard-wired mechanically, per vendor. A vendor with nothing to wire cannot be brought inside it, and routing work away from an unwired vendor is explicitly not how that gap closes. - **Amazon Q** — **REJECTED on vendor status**, which is a weaker class than a measurement and is named as one: the product is dead / end-of-life. There is nothing left here to measure. - **Amp** — **REJECTED on platform.** WSL-only on Windows, and its threads live server-side. Windows is this product's primary target (ADR-002), so a WSL-only binary is not a seat here; and a vendor whose conversation state sits on someone else's server gives §9.4's native resume nothing on this machine to resume from. - **Goose** — **REJECTED on platform.** No Windows path at all. The same ADR-002 reason as Amp, without the second problem. - **Pi** (`earendil-works/pi`) — **the named instance of the re-host class below**, recorded by name because its size makes it the one people ask about. A Pi seat re-hosts model families already seated here. The HUD adapter is a different question with a different answer: the format is surveyed (§3.9b), live-verified 2026-08-16, and `internal/adapter/pi` observes sessions. Council still does not seat it. - **Every BYO-harness re-host** — **REJECTED as a class**, rather than one vendor at a time. Each of them re-hosts a model family that already has a seat here. §9.4's whole argument for turn 1 being blind is that the answers are *independent*; a harness wrapped around a model already in the room buys a column that agrees with an existing one for the reason a copy agrees with its original. That is a surface, not an opinion, and the room is priced per seat (§9.21). - **Kimi Code 0.34.0** — **INSTALLED, WAITLISTED, UNVERIFIED. This is not a rejection**, and it must not be read as one. The binary is on the reference box; its hook surface and its session surface have never been observed, so there is no measurement to seat it on and none to refuse it on either. One trap is recorded because it is the obvious wrong shortcut: the legacy `kimi-cli` documentation is **VOID** for this binary — it describes a different program — and any claim read out of it would be exactly the "read off `--help`" evidence §9.2 refuses. - **Qwen Code** — **RUNNER-UP: measured, and it fails on one specific thing.** It has a real statusline hook, which is more than several seated vendors offer. What its payload does not carry is any rate-limit window, so the quota relay (§7.15) would have nothing to write and the header could speak for this vendor only by inventing the numbers it renders. Under §4a.1 that is `CapNone` rather than a plausible fill, which leaves a seat whose gauge is permanently blank. This is the one entry worth re-checking when a vendor payload changes; the rejections above do not move until the vendor does. ### 9.2 Two claims the room refuses to leave implicit Every column header carries its own **sandbox badge** and its own **streaming granularity**, because vendors across the 4-vendor fleet differ on both and the first draft of ADR-008 got this wrong in the direction that matters — it claimed "enforced read-only sandboxing" for all seated vendors when only one had a mechanism named. | | mechanism | badge | |---|---|---| | Claude Code | `--disallowedTools ` + `--strict-mcp-config` | `ro:tools` | | Codex (macOS/Linux) | `-s read-only`, enforced by the OS sandbox | `ro:enforced` | | Codex (Windows) | `-s read-only`, enforced since codex-cli 0.149.1 — before that, **no sandbox**: `-s danger-full-access` was the only mode that could spawn a process there (the 2026-08-29 amendment below) | `ro:enforced` | | Antigravity | **no restriction flag at all**: `--mode plan --sandbox` were measured not to restrict writes, and were dropped once their only observed effect turned out to be a dead turn (§9.6b). Since the 2026-09-03 amendment council passes `--add-dir ` in every posture and `--mode accept-edits` in the write postures, because the cwd is no workspace to this vendor and a write inside a named workspace is auto-denied without the edit grant (§9.59) | `unsandboxed` | There is no level that renders as an unqualified "read-only", and after the live spike there is one that renders as the opposite. Antigravity was asked to write a file under both of its read-only flags and wrote it — file confirmed on disk, reported permission mode and tool list byte-identical to a run without the flags. That is refuted, not unverified, so it gets a fourth level badged `unsandboxed`. Deliberately not `ro:none`: every other badge opens with `ro:`, a reader scanning column headers takes in the prefix before the qualifier, and a vendor that can edit your working tree must not read as read-only at a glance. Council **stopped passing those two flags** on 2026-08-04 (ADR-008, seventeenth amendment). The badge is unchanged and so is every word behind it — it never rested on the flags being sent, it rested on the write having landed — and what moved is one clause of the detail, which used to end *"the flags are still passed; they do not restrict it"* and could not go on saying so. This is the second seat where a posture flag came off because it was doing nothing useful and one measured harm; the Codex row above is the first, and the two were decided on the same ledger. §9.6b carries the argument. **Codex is the seat where the OS changes the answer, and until 2026-08-29 it wore that same `unsandboxed` badge on Windows.** Re-probed 2026-08-04 against codex-cli 0.146.0: `-s read-only` *and* `-s workspace-write` both failed every process spawn there with `CreateProcessAsUserW failed: 5 (Access is denied.)`, including a control asked merely to list a directory. So "read-only" on Windows was not a read/write distinction — it was a seat that could not read, which is exactly how this surfaced: a live council turn answered a "thoughts on this repo" brief with *"I could not inspect the repository."* Council passed `-s danger-full-access` on Windows in **both** postures for as long as that held, because it was the only mode that ran, and the badge told the truth about that rather than keeping a comfortable word. The containment is the workspace, not the flag (ADR-008, third and twelfth amendments). **Amended 2026-08-29 — the Windows sandbox re-measured at codex-cli 0.149.1, and the read posture earned its badge back.** A chip re-probed the twelfth amendment's finding against the build now installed, in throwaway directories, one turn per probe, files checked on disk rather than read out of the reply: - **`-s read-only` enforces.** The pwsh spawn still fails with the same `CreateProcessAsUserW failed: 5` line — but only pwsh: the model retried through `C:\WINDOWS\system32\cmd.exe`, which spawned *inside* the sandbox and obeyed it. A read turn listed the directory and read a marker file, exit 0. A write turn (`cmd /c echo probe> wrote-ro2.txt`) came back `Access is denied.`, exit 1, no file on disk. - **`-s workspace-write` enforces and contains.** A write inside the workspace landed; a write outside it (and outside the temp roots the mode allows by design) was denied with no file on disk. - **The `.git` carve-out holds on Windows and CANNOT be bought back.** The `sandbox_workspace_write.writable_roots=["/.git"]` override that §9.6's macOS measurement showed unlocking `.git` was passed in both the forward-slash spelling `gitWritableOverride` emits and a backslash spelling — the `.git` write stayed denied in both, while the same override named an ordinary outside directory and unlocked it. The deny outranks the override at this build. - **The resume override enforces.** A turn resumed with `-c sandbox_mode="read-only"` had its shell write denied, no file on disk — the effect §9.7 recorded as owed and unobservable is now observed. So the postures split on Windows instead of collapsing: **read passes `-s read-only` and is badged `ro:enforced`** — with the stated residual that the pwsh spawn failure is routed around by the model's own retry, so a read turn can still fail to inspect when the model stops at the spawn error (a liveness caveat, not a safety one; its cause is unmeasured) — while **write keeps `-s danger-full-access`**, because a workspace-write seat that cannot write `.git` is the exact defect the `.git` widening exists to prevent: it edits files all session and never lands one. `vendors/codex.go` carries the capture on its constants; `TestCodexPostureIsPerOS` pins both halves of the split, and `TestNoVendorClaimsUnverifiedEnforcement` now requires the Windows claim to cite the build it was measured on. macOS and Linux are untouched. **Amended 2026-09-01 — the installed build moved to codex-cli 0.151.0, and the sandbox claims above were NOT re-measured on it.** A chip traced the seat's `failed (exit 1)` ([§9.58](#s9-58)) and re-ran the three seat argv shapes at 0.151.0 from a scratch directory. All three parsed and produced `thread.started`, so the flags this section rests on are still accepted. The four sandbox probes could not run: every turn that day died at the account's usage limit before a tool was asked for. So the table above is pinned at 0.149.1 on a machine that runs 0.151.0. [STATE.md](../STATE.md) carries the owed re-measurement. `TestSandboxBadgesAreNeverBlanket` fails the build if a bare claim reappears, and asserts the badges stay distinct — convergence on one string is how a per-vendor claim quietly becomes a blanket one again. What these badges *say* has not changed since; how loudly they say it has. §9.11 gives the two that mean "this seat can change your files" weight and the warning hue, because the room had been drawing `unsandboxed` at exactly the volume it drew `ro:tools` beside it. The words still carry the whole distinction — that is why they break the `ro:` prefix — so nothing here depends on colour; the weight only makes the word findable in a frame with four columns of prose in it. **The Claude row cost three attempts to get right, and the failure mode is worth recording.** The original ADR claimed enforcement with no mechanism named. The first correction named `--allowedTools "Read,Glob,Grep"`, which *sounds* exactly right and is not: it pre-approves tools for permission prompts, it does not remove them from the session. Running the real invocation and reading the `system/init` event's own `tools` array showed `Edit`, `Write` and `Bash` still there. Every test in the package passed at that point, because every test asserted the **flag** and none asserted the **effect** — which is this repo's own False Green failure, committed inside the feature whose entire premise is refusing it. What works is `--disallowedTools` plus `--strict-mcp-config`, and two parts of that are easy to miss. Deny **PowerShell**, not just Bash — denying only Bash leaves a working shell on the platform this product targets. And drop MCP servers, because without `--strict-mcp-config` the session inherits whatever the user has connected; the verification run surfaced Gmail write tools in a session with every built-in write tool denied, and no fixed deny list can name those in advance. The residual limitation is stated in the badge's own detail text rather than hidden: a deny list cannot cover a tool that does not exist yet. The claim is *these named tools are absent, verified*, not *this session cannot write*. The general rule this leaves behind: **a flag's name is not evidence of its effect**, and the check that matters is what the session reports about itself afterwards. **Amended 2026-08-17 — the seat's own `capabilities` array, measured and deliberately not read.** The rule above ends on *what the session reports about itself afterwards*, and Claude Code's `system/init` carries a field that looks like exactly that: a `capabilities` array. It is not the same thing. A tool list reports what the session **holds**, which is a fact about the session. A capability token reports what the vendor **can do**, which is a claim about behavior. `claude.go`'s Interrupt precedent already refused to rest on one, and this is the case that precedent was written for. Measured at **Claude Code 2.1.233** on Windows, from two live headless runs in a throwaway directory: a plain `-p --output-format stream-json --verbose` invocation, and the read posture's exact argv. Both init frames carried the same array, in the same order: ``` ["interrupt_receipt_v1","interrupt_cancel_queued_v1","msg_lifecycle_v1"] ``` Three findings, and none of them supports a gate. The array does not move with our flags, so it says nothing about the posture council asked for. The tokens do not date a build — `interrupt_cancel_queued_v1` is already in the 2.1.226 bundle, so a version check resting on their presence passes on three installed versions at once. And the one token council might plausibly want, `interrupt_receipt_v1`, guards a capability a live run already verified, so reading it can only weaken a claim that already stands. `streamLine.Capabilities` therefore parses the field, and nothing branches on it. `TestInitCapabilitiesAreParsedAndGateNothing` pins the three states apart — absent, `[]`, and the measured array — and pins that every one of them still produces the same `KindSession` event. Modelling a field nothing reads needs a reason (§7.16b). The reason here is that a modelled field is a checkable record of a measurement, and a comment is not. **`telltale doctor` cannot carry this line, and its charter is the reason.** The package doc is explicit that doctor does not start a turn, because a turn costs real quota (ADR-008 §6). `capabilities` arrives only on a `system/init`, and a `system/init` arrives only when a turn starts. So no cheap local read of this field exists, and the check doctor could offer must spend exactly what doctor exists not to spend. Recorded here instead of built. Granularity is the same discipline applied to streaming, and the spike made the answer worse than the guess. Claude streams token-level deltas, verified live. The other two were provisionally labelled `events`, on the reasoning that a coarse stream is still a stream; neither streams at all. Codex emits one `item.completed` per complete agent message and has no message-delta feature even under development. Antigravity delivers an entire response as a single `text_delta` — a one-word reply left its column blank for 73 seconds and then painted at once. Both are `GranFinalOnly`. A vendor that emits nothing until it finishes renders `PhaseWaiting` — a first-class phase, named as such in its column header — rather than an empty column that looks like slow streaming. `TestWaitingIsNotStreaming` asserts the two never render alike. That distinction was added on the theory that some vendor might not stream; it turns out to describe two thirds of the room. This is §4a.1's rule (a dropped column and an em dash must not read the same) applied to a surface where the ambiguity would otherwise be invisible. The card's WORDING is not what carries it, and §9.14 is why that matters: the body used to recite *"this vendor reports no incremental output, so nothing appears until the turn finishes"* on every waiting turn, which is council's plumbing described in council's vocabulary in the space a user came to read an answer. The distinction is carried by the header's own phase word and the granularity badge beside it; the body says only `working — the reply arrives whole.` ### 9.3 Execution: argv, never a shell Prompts are arbitrary text — quotes, ampersands, whatever was typed — so no prompt is ever interpolated into a command string. Specs are `{Binary, Args []string, StdinPrompt, Dir}` through `exec.CommandContext`. This is load-bearing on Windows specifically. `LookPath("codex")` resolves to `codex.cmd`, an npm shim, and Go's `os/exec` runs `.cmd` and `.bat` through `cmd.exe`, whose argument parsing cannot be safely quoted for arbitrary text. So `detect.go` classifies every resolved path as `KindNative` or `KindShim`, Codex and Claude take their prompt on **stdin** (Codex via its verified `-` sentinel), and a vendor that is *both* a shim and argv-only is marked `AvailUnusable` and not driven at all. The refusal is the feature; the card tells the user which env override fixes it. ### 9.4 Multi-turn is native resume, not transcript re-send Turn 1 is blind: no vendor sees another's answer, which is what makes opinions across the 4-vendor fleet independent rather than anchored. Later turns ride each vendor's own session-resume (`claude --resume`, `codex exec resume`, `agy --conversation`). Re-sending the transcript would grow input quadratically against metered quotas and flatten native turn structure into quoted prose; resume sends only the new turn and makes the blind-round guarantee *structural*, since each session holds only its own history. Cross-agent rebuttal is an explicit opt-in toggle that quotes the previous turn's finals as labelled untrusted material. ### 9.5 Layout and testing Same contract as §7.9, for the same reason: `Render` is pure over `State`, tests construct state by hand, goldens live in `internal/council/testdata/golden/*.txt` and render with `PlainStyles()`. Three columns at ≥96 cells; below that, or when a column would fall under 24 cells, the tier drops to a tab bar rather than shredding prose into unreadable ribbons. Width is measured with `lipgloss.Width`, never `len()`. Two of the frame's rows are no longer constants, and the ordering that makes that safe is worth stating: the **tier is settled before any row is budgeted**. The tab bar costs a row, and the fallback from columns to tabs is a width test — budgeting first and dropping the tier afterwards worked only while the footer was a fixed three rows, and a taller composer would have overflowed the terminal by exactly the tab bar. `resolveLayoutIn` therefore finishes deciding the tier, then spends rows: header, footer chrome, tab bar if any, collapsed-seat notice if any, then the composer up to its ceiling, and **the composer yields before the body does** — at the minimum height a six-row draft would leave the columns nothing, and a room you can type in but not read is not the trade anyone asked for. A tab bar holding a single tab is not drawn at all: it selects nothing and repeats the column header underneath it, which stopped being a rarity the moment dead seats began folding away (§9.9). One trap worth recording, because the golden tests could not have caught it: `padRight` truncates rune by rune, so on text that already carries ANSI escapes it cuts through an escape sequence and counts escape bytes as content. Goldens render with the identity style set by design, so they are blind to it. Anywhere a line is assembled from differently-styled pieces — the tab bar, the help body — padding goes through `fit`, which is ANSI-aware. `TestFitIsANSIAware` is the regression guard. ### 9.6 Invocation traps, one per vendor Each adapter hit a failure that is silent rather than loud, which is the kind worth writing down. - **Claude**: `--allowedTools` pre-approves, it does not restrict (§9.2). Enforcement is `--disallowedTools` + `--strict-mcp-config`. - **Codex**: `codex exec` and `codex exec resume` **do not take the same flags**. `-s` and `--cd` are rejected by `resume` with an argument-parsing error, and a parse error means *empty stdout* — a naive resume would blank the column on every follow-up turn with no card able to explain it. Resume carries the posture as `-c sandbox_mode=""`, derived from the same function as the spawn path so the two cannot drift, and takes its workspace from `Spec.Dir` alone. The session id is **positional**, not a flag value. - **Antigravity**: `-p` is a **string flag whose value is the prompt**, not a boolean. Written in the natural order, `agy -p --output-format stream-json ""` exits 0 and cheerfully answers a question about the flag it just swallowed. `-p` must be last, brief immediately after, every other flag before it. agy also rejects a prompt on stdin, so its brief goes in argv and is bounded by the ~32K Windows command-line limit — a real ceiling on a long brief, with no workaround short of upstream support. - **Cursor**: the prompt is a variadic positional and needs a bare `--` in front of it, or a brief that happens to open with `-` is read as an unknown option and the turn dies. And the stream sends every passage twice — deltas, then the whole message again — which §9.6c covers, because it is a parsing trap rather than an invocation one and it took two captures to state correctly. The shared shape: all four failures produce a *plausible* result rather than an error. That is why each one is pinned by a test asserting the argv this repo actually builds. ### 9.6a The activity trace carries outcomes — and says when it cannot The trace answers *what did this agent do*. Until this landed it could not answer *did it work*, which made it the same half-built gauge §4a.1 exists to forbid: `⚙ Bash: go test ./...` renders identically whether the suite passed or the build never compiled. The results were not missing, either — they were arriving in the same stream the commands came from and being dropped on the floor. A room that discards knowledge it has is the mirror image of one that invents knowledge it lacks, and both are the same failure. Four statuses, and **Unknown is the one that earns the type**. Pending renders as the bare entry, OK as `✓`, Failed as `✗` plus the vendor's own first line about why, Unknown as `?`. ASCII gets `+`, `x`, `?` — chosen around everything already spoken for, since `*` is the activity prefix, `>` the ellipsis, `]` focus and `#` the HUD's gauge fill. Every distinction is a glyph before it is a colour, so all four survive `--ascii` and a monochrome terminal. **Where the outcome comes from, per vendor, and how strong the claim is.** | | signal | verified | |---|---|---| | Claude Code | `user` messages carrying `tool_result` blocks with `tool_use_id` + `is_error` | **live**, 2026-08-04, Claude Code 2.1.220 | | Codex | `item.started` → pending; `item.completed` with `exit_code` / `status` | captured fixtures; `exit_code` and `status:"failed"` observed, `status:"completed"` **never** | | Antigravity | `step_update` ACTIVE → pending, DONE → **Unknown** | live capture shows DONE carries no success signal at all | Three things the Claude probe settled that a docs-first parser would have got wrong. Field **order** differs between captured lines, so nothing may be read from position. `is_error` is **absent** on some successes — a `Read` result carried only `tool_use_id`, `type` and `content` — so absence is success rather than unknown; Claude Code marks failure and stays quiet about the rest. And the results came back **out of order**, the second call's failure landing ahead of the first call's success, on the very first probe. That last one is why correlation is by id and never by arrival order: a trace zipped by position would have blamed the wrong command on its first real run. Antigravity is the case the Unknown status exists for. Its steps flip ACTIVE then DONE, and every captured DONE line carries `duration_seconds`, sometimes a `tool_info` with the call's parameters, and nothing whatsoever about whether the step achieved anything. agy reports success or failure exactly once per turn, in the final `result` event, and that verdict is about the *turn*. So a finished agy step renders `?` — not `✓`, and the code comment says why. Reusing the success mark would be council inventing a result on a vendor's behalf, which is the `--allowedTools` mistake (§9.2) wearing different clothes. Codex carries the same discipline in a smaller way. `exit_code` is a **pointer**, because codex spells "still running" as `"exit_code":null` and a plain int would flatten that to 0 — the spelling of success, and the most expensive confusion available on that field. And an item that completes with neither an exit code nor `status:"failed"` resolves **Unknown**, not OK: no captured line has ever carried `status:"completed"`, and guessing the success spelling from the observed failure one would be a success claim built on a string nobody has seen. That deliberately weak mapping tightens the moment a live run shows the spelling. `TestActOutcomesRenderDistinctly` fails the build if any two statuses ever render alike; `TestOverlappingToolCallsResolveToTheRightEntries` replays the real out-of-order probe. **A failed entry is a card now, and it was the one card §9.11 missed.** That pass gave every card in a column one grammar — a title with its body hanging under it — and cited the trace as somewhere the room *already* did it, on the strength of the failure detail's indent. The entry itself never got it, and a live room showed why that matters: `run_command: pwsh -Command "Get-ChildItem"` does not fit 37 cells, so the command wrapped to a continuation starting hard against the column edge, reading as a second nameless entry with the outcome mark stranded on it. It now hangs under its own `⚙`, which costs no rows and makes one call look like one call. Two things then had to change with it, and both are the kind of detail that only shows up against a real capture: - **The reason indents FOUR, not two.** Once the command hangs at two, a reason at two lands in the same column as the tail of the command it explains — telling them apart by colour alone, which this product does not do. Goldens render with `PlainStyles`, so that golden is exactly the artifact that proves it. - **The reason is flattened and bounded.** `sanitize` preserves newlines on purpose, because a vendor's prose reply is prose; a tool failure's detail is not prose, and multi-line stderr pushed through the wrapper arrived as ragged fragments at random widths. It is now collapsed to one flowing line and capped at three rows with the room's own ellipsis, so a clipped reason can never read as a complete one. The clip has an answer — `f` expands the column to the full frame, where the same reason typically survives whole — and a refusal behind it: the trace answers *what did this agent do and did it work*, not *show me the log*, and the turn-level failure still arrives in the column's note carrying the vendor's own sentence. ### 9.6b The agy trace was showing its message-passing and hiding its work Driven live, the Antigravity column's trace read `user_input ?`, `system_message ?`, `checkpoint ?`, `unknown ?` — and, for every real thing the agent did, a bare `tool ?`. Three separate defects wearing one symptom, all found by reading captured stdout (agy 1.1.10, Windows, 2026-08-04) rather than the adapter. **1. The plumbing is suppressed, and the line that decides what counts as plumbing is not "noisy".** Hiding a vendor's ACTIONS would be a false gauge — a quiet column for an agent busy editing the workspace, which is §4a.1's failure with the sign flipped. Hiding its PLUMBING is noise reduction. So the suppression is an allowlist defended per kind against a captured line, never a filter on what looks like chatter, and it lives in the adapter (`ParseEvent` returns `false`) rather than in the view, because `Render` is pure over `State` and a step that is not an action must never become one. | kind | why it is plumbing, from the capture | |---|---| | `user_input` | step 0 of every turn, `DONE`, nothing else on the line — the brief council itself just sent, echoed back | | `system_message` | same empty shape; agy placing its own message into the conversation | | `checkpoint` | `duration_seconds` and a ~120-token usage block, nothing else — a thread bookmark, never the workspace | | `error_message` | an empty marker on a failing turn: no message, no error field, no duration | | `unknown` | one per turn at a fixed preamble slot (step 1, right after `user_input`), 0.0005s and 0.0045s across two turns, no tool name, no parameters | `error_message` needed the most care, because dropping the only visible sign that a turn went wrong is the opposite mistake. It is safe for a checked reason: both captured failing turns end `result` with `status:"ERROR"` and `error:"Agent execution terminated due to error."`, and that path already produces a `KindError` carrying the vendor's sentence. The turn-level failure IS reported, with words. A rendered `error_message ?` is strictly *less* than that — an ominous name with a shrug attached. The result path now prefers `result.error` over the composed status line precisely so that argument keeps holding. `unknown` had to be argued rather than listed, since suppressing a step whose type the adapter merely does not RECOGNISE is the same class of mistake as inventing an outcome for it. The capture says this is agy's own label and not our ignorance: fixed position, half a millisecond, and no tool name — while every step in every capture that did something carried one. What would reverse the decision is written as code, not as a promise: an `unknown` step that names a tool is **not** suppressed and renders under that name, so if agy ever starts acting through this label the trace shows it that same turn. **2. agy's real tool names were on the wire the whole time and were not being read.** A tool step carries `tool_name` at the top level *and* `tool_info.name` with `tool_info.parameters` beside it. The adapter rendered `step_update.step_type` — the literal string `"tool"` — so every call, whatever it was, produced one indistinguishable entry. This is ADR-008's tenth amendment repeating itself in a second costume: Cursor's `tool_call.tool.case` lookup matched nothing because the oneof arrives flattened to a key on the wire, and every Cursor trace entry read `tool call`. Same cause both times — the fields the vendor sends were never compared against the fields the parser reads — and the same fix, which is to **parse what arrives**. Observed names: `list_dir`, `run_command`, `write_to_file`, `list_permissions`. The entry now follows the grammar the other three adapters already use — `Glob: **/*.go`, `Bash: go test ./...` — so `⚙ tool ?` becomes `⚙ list_dir: C:\Users\…\antigravity-cli\scratch ?`. The argument rule is deliberately small, because agy's parameter keys are vendor-specific (`DirectoryPath`, `CommandLine`, `TargetFile`) and an arbitrary object is not a trace line: only string values are candidates, one such value renders (which is every captured shape), several resolve to the lowest key name by byte order, none degrades to the bare tool name. Rule three is not a claim about which key matters — it is a refusal to let Go's randomised map iteration reach a rendered line or a golden, pinned by `TestAgyToolArgIsDeterministic`. **3. A failed agy tool call used to render as permanently pending.** There is a fifth state, `ERROR`, and the switch handled `ACTIVE` and `DONE` only, so the line matched nothing and the entry its `ACTIVE` twin had opened stayed pending for the rest of the room's life — the trace claiming a command was running after the vendor had given up on it. It carries its own reason in `tool_info.error.message`, so it maps to Failed with the vendor's own first line, exactly as §9.6a specifies. Failed and not Denied: `ActDenied` is council's first-hand record of its own gate keystroke, and a refusal read off someone's stream is not that. The `DONE → Unknown` rule above is unchanged and narrows in one direction only — agy does report per-step failure, and it still reports no per-step success. **The resume note was a misdiagnosis, and the fix is to the claim rather than to the mechanism.** A seat whose restored thread failed its first turn used to say *"the saved thread was refused — this seat's history is gone."* **Measured**, single trial, 2026-08-04: `agy --conversation ` **does** resume in 1.1.10. The same `conversation_id` came back, `step_index` **continued** (10 → 11) rather than restarting at 0, and `result.num_turns` was 2. That demonstrably live thread's turn nevertheless ended `status:"ERROR"` / `"Agent execution terminated due to error."`, and a separate attempt died before any thread was involved at all — a bare `result` with an **empty** `conversation_id` and *"Eligibility check failed: UNAVAILABLE (code 503): The service is currently unavailable."* So agy turns fail transiently for reasons that have nothing to do with the conversation, and "the history is gone" is a claim the evidence does not support. The behaviour was deliberately untouched at the time: one failed turn still dropped the id, for the reasons ADR-008's ninth amendment gives at length, and no new signal was invented to tell the two cases apart because none had been observed. Only the sentence was narrowed, to the three things known — the first turn on the restored thread failed, the seat has let the saved thread go, and the next brief starts a new session with the brief re-applied. **That is now out of date, and the evidence above is what dated it** (ADR-008, sixteenth amendment). The paragraph declined to change the rule on the ground that a record is not a fix, which was right — and left a measurement sitting beside a rule it contradicted. The rule's default is unchanged: a restored id whose first turn fails is dropped, because a seat retrying a genuinely dead id rebuilds the same doomed invocation on every turn for the life of the room. What changed is that a failure which is **identifiably transient** is now treated exactly as a cancellation already is — nothing was learned about the thread, so nothing is forfeited, and the seat stays on probation so the next unclassified failure still costs it the id. Two classes qualify, and both are positive evidence that the vendor never reached the conversation. **Pre-flight**: the failures `failureNote` already classifies off captured stderr — not signed in, an untrusted workspace, a sandbox the vendor's own config demands and its own help refuses, a binary that vanished — each documented at its case as exiting before any model call; and, one step earlier, a dispatch that never started a process at all. **Vendor-reported outage**: the 503 quoted above, matched on agy's own sentence, with the capture's empty `conversation_id` as the corroboration that it died before a thread was involved. Everything else still drops, and the asymmetry is deliberate — a lost conversation costs one conversation, a wedged seat costs every turn of the room. Claude and Codex get nothing vendor-specific here: neither has a measured transient signal, only measured *dead-thread* strings pointing the other way, and their behaviour is unchanged. agy's commonest failure sentence — *"Agent execution terminated due to error."* — is deliberately **not** classified, because it was captured on a demonstrably live thread and is also what a dead one would plausibly produce, and a string on both sides of a distinction is evidence for neither. The classification is produced where the evidence is (the runner's stderr classifier, the adapters' result parsers) and travels on the event as a small enum. It is never re-derived by matching the rendered note: that note is prose written for a narrow column, and keying a mechanism off it would make every wording change a silent behaviour change. It lives on `Model` and never on `State` — a decision input the renderer has no business reaching. **The card that says this changed shape too.** One ⚠ plus a single sentence carrying an outcome and a mechanism wraps to three lines of uniform weight in a 37-cell column, and three of those side by side reads as a room on fire over a seat that will simply start a new session. It is now §9.11's card grammar: a short title — *thread not restored — starting fresh* — with the mechanics hanging under it, quieter, and **no warning mark**. That is the same fact `reattachCard` already states calmly at idle when no thread came back, learned a turn later; the ⚠ has to go on meaning *something went wrong* for the notes where something did. The words carry the card in every glyph set, so `--ascii` loses nothing. **A side measurement, separately labelled, with its confound stated.** Under `--mode plan --sandbox`, agy's `run_command` was refused with *"granting access to C:\: Access is denied."*, the agent gave up, and the whole turn died `status:"ERROR"` with an empty response. The control run with both flags **dropped** ran a shell command and returned `status:"SUCCESS"`. ADR-008 and the `baseArgs` comment previously said `--sandbox`'s effect on the shell "was NOT tested and is not claimed"; this is the first evidence on it, and that comment no longer says so. **It is one trial per arm with an uncontrolled difference: the two turns issued different command lines (`pwsh -Command "Get-Location; Get-ChildItem"` versus `Get-ChildItem`), so it is not a clean A/B**, and the refusal's mention of `C:\` may be about a drive root rather than about the flag. What it does establish is the flag's observed cost — a dead turn, with nothing rendered. The posture flags were **not** changed on the strength of it; that is a decision to make deliberately and separately, and this was a record, not a fix. **That decision is now made: the flags come off, in both postures** (ADR-008, seventeenth amendment). The open question this section carried — *should council keep asking agy for `--mode plan --sandbox` when both are measured to do nothing?* — is closed the way the evidence points, and the ledger is one-sided rather than a close call: - On the **write** side, the flags were measured restricting nothing. Asked to write a file under both, agy wrote it; reported permission mode and tool list were byte-identical to a run without them, and `write_to_file` was still in the list. Refuted, not unproven. - On the **shell** side, the one and only effect either flag has ever been observed to have is the paragraph above: a refused `run_command`, an agent that gave up, and a turn that ended `status:"ERROR"` with an empty response. The user sees a blank column. - So the flags bought **no restriction that was ever observed**, at the price of turns that die with nothing rendered. That is not caution; it is the appearance of caution paid for in the vendor's actual answers. The confound above is unresolved and does not need to be: it concerns *why* the turn died, and the decision only needs *that* it did, set against a benefit measured at zero. **No honesty claim moves with them, and that separation is the point of having waited.** §9.13 deliberately changed what the room *says* about this posture and nothing about the posture, because a documentation pass that quietly retunes a safety flag is exactly what this file exists to prevent. This is the other half, made on its own, by the owner. The badge stays `unsandboxed`; the detail loses the clause claiming the flags are passed, because they are not, and a detail describing council's own behaviour inaccurately is the one class of false claim this repo has no excuse for. The containment was never these flags — it is the workspace (§9.2, ADR-008 third and twelfth amendments), and agent-ops ADR-012 rules the same way independently. Deliberately **not** part of this: `--dangerously-skip-permissions`. Dropping a flag that restricted nothing and adding one that approves everything are different acts, and the second stays refused on both seats that offer it. **Amended 2026-09-03: two flags come ON, and neither is a restriction.** `--add-dir ` in every posture and `--mode accept-edits` in the write postures. The first names the workspace, because the cwd alone is none to this vendor. The second is the vendor's edit-only grant, and without it print mode auto-denies every write inside a named workspace. [§9.59](#s9-59) carries the measurement. `--dangerously-skip-permissions` stays refused. ### 9.6c The Cursor stream says everything twice, and the second time does not always look alike cursor-agent under `--stream-partial-output` sends a model call's text deltas and then that call's **complete message** as one more assistant event. Appending both renders the passage twice, which is the whole of this defect in both of its appearances. The first capture (2026-08-04, a turn asked to reply `PONG`) showed deltas `"P"`, `"ONG"` each carrying `timestamp_ms` and the repeat `"PONG"` carrying none, so the adapter dropped the event whose `timestamp_ms` was absent. That rule was derived from turns with no tool call in them, and **every such turn is one model call** — so it was a rule about the end of a *turn* being used as a rule about the end of a *message*. A turn that runs a tool is several model calls, each ending in a repeat of its own segment, and those mid-turn repeats carry `timestamp_ms` like any delta. The column rendered the segment, the segment again, then the next one — `X X Y` — which is what the owner saw on a long Cursor reply and what replaying the captured turn through the old parser reproduces exactly. What separates them is **`model_call_id`**: present on every whole-message repeat that ends a mid-turn model call, absent from every one of 108 captured deltas, and carrying the vendor's own per-segment numbering (`…-0-x7su`, `…-1-15l2`) that also appears on the `tool_call` events between the segments. The adapter now drops an assistant event when `model_call_id` is present **or** `timestamp_ms` is absent — the second is kept because the *turn-final* repeat still carries neither, so dropping it would trade this bug for the first one. Presence rather than absence is the point, and it generalises past this vendor: a missing field cannot distinguish "the vendor is telling me this is a complete message" from "the vendor stopped sending that field". `internal/council/vendors/testdata/cursor-segmented-turn.jsonl` is the whole turn, redacted, replayed by `TestCursorSegmentedTurnRendersEachPassageOnce`, which asserts the streamed body equals the reply the vendor itself put in its `result` event. That `result` remains the safety net if both fields ever go: the room uses it whenever a column streamed nothing, so the failure mode is a column that fills at the end, never one that is wrong. ADR-008's twentieth amendment carries the argument. ### 9.7 Status The room opens, detects the four seats, renders both layouts and every degraded state, takes a brief, and dispatches it. Claude streams incrementally; Codex and Antigravity render the waiting card and fill at once. Quitting the room kills every child, including the persistent one — and no longer strands the conversation: a bare `telltale council` reopens the one saved room by default, `--fresh` starts over, and `/cd ` typed in the composer moves the room to another workspace between turns, with the persistent Claude seat following by respawn on its own session id (ADR-008, ninth and eleventh amendments). Multi-turn is one live process for Claude (§9.8) and for Cursor (§9.36), and native resume for the two seats that are still batch programs. Cross-agent rebuttal (§9.4) and per-column scrollback are **built and shipped**, and the scrollback now spans the whole conversation rather than one turn (§9.9): the room keeps a per-column transcript, echoes the brief that produced each turn, composes in up to six rows, and folds the seats it cannot drive out of the grid. Not built: per-vendor cancel — `ctrl+c` still ends the whole turn. Known gaps, stated rather than buried. Codex's non-shell write path is untested — asked to create a file with its own patch tool it declined, but that was a model choice and says nothing about enforcement. Neither vendor was observed producing a failure event on stdout, so the error branches are modelled on exit code plus stderr rather than an observed schema. Antigravity's `--print-timeout` is left at its 5-minute default, which is a hard ceiling on a long council turn and a policy choice worth making deliberately later. Unverified and scheduled as a live spike before the Codex and Antigravity columns ship: the Codex `--json` event schema and delta granularity, whether `codex -s read-only` engages on Windows, and Antigravity's stream-json schema, conversation-id location, stdin support and `--sandbox` semantics. Those columns render honest *requested* badges until it says otherwise. **That spike ran, and the Windows sandbox question is closed the other way.** `-s read-only` does not engage on Windows in any useful sense: it fails every process spawn, reads included, and so does `-s workspace-write`. Codex on Windows is invoked `danger-full-access` and badged `unsandboxed` — see the §9.2 table above and ADR-008's twelfth amendment. What remains open on this seat is narrower: whether the `-c sandbox_mode=` override actually changes behaviour on the *resume* path. The key is accepted; its effect has never been separately observed, and until this change every mode failed identically so there was nothing to observe. *(Both halves of that paragraph moved on 2026-08-29, at codex-cli 0.149.1: the Windows sandbox now enforces and the read posture is `-s read-only` again, and the resume override's effect was observed — a resumed read-only turn had its shell write denied. §9.2's 2026-08-29 amendment carries the measurements. The paragraph above is kept as the record of what was true at 0.146.0.)* One claim in this section is looser than its measurement and is flagged rather than quietly corrected, because fixing it is a separate change to a separate surface. The Claude column's granularity word is `tokens`. Measured over a 250-word reply the deltas are **~80 characters each, about three a second** — genuinely incremental, and not tokens. Measured identically under the persistent invocation and under a spawn-per-turn control, so it is a pre-existing overstatement rather than something §9.8 introduced. ### 9.8 One live process, and the gate it makes possible Every Claude turn used to be a fresh `claude -p --resume`. A one-word "gm" cost about 25 seconds and $0.23, nearly all of it session init, paid again on every turn. That was the visible cost. The structural one is what forced the change: **a batch process cannot ask permission.** Its stdin is written and closed before the first token arrives, so it has no channel to ask on and none for an answer to come back on. `--input-format stream-json` keeps one process alive taking one JSONL message per turn on an open stdin. Verified live against Claude Code 2.1.220 rather than read from documentation: two turns down one stdin came back under the same `session_id` with the same pid, `system/init` is re-emitted at the *start of every turn* (a parser reading it as "a new session" would reset the seat once per turn), and one `result` per turn is the only end-of-turn signal there is, because there is no exit to infer one from. Cancelling a turn now **interrupts** rather than kills — `{"subtype":"interrupt"}` on the control channel, measured to end the turn and leave the process answering a further one. Killing would also work, and would throw away the session init the room just paid for, so cancelling one turn would quietly make the next one expensive. **The reported cost changed meaning and the badge changed with it.** `total_cost_usd` is a running total for the process: $0.1061493 → $0.1177296 across two turns while the per-turn `usage` block stayed at 2 input tokens both times. That cell has meant "this turn" everywhere else in the room, so it now reads `$0.1177 session`. Rendering a session total unlabelled would be a false reading of a true number; subtracting to recover the turn would be council inventing a figure, which §8 rejects. #### The gate With `--write`, the seat that can ask **does** ask. Every tool call raises an approval card in its column — the tool and its argument line, formatted exactly as the activity trace formats it — and the room enters a gate state: `y` approves, `n` denies, and the vendor is stopped until one of them is pressed. Blocking was measured, not assumed: the answer was withheld for twenty seconds and nothing else arrived on stdout in that window. Three flags turn it on, none is optional, and **two of them do nothing alone**: | flag | what it does | what happens without it | |---|---|---| | `--permission-prompt-tool stdio` | routes the request onto the stream | **absent from `--help`** and real; alone, the session runs in auto mode, no request is ever emitted and the file is written | | `--permission-mode manual` | makes the call ask rather than assume | alone, there is nobody to ask, so the call short-circuits to *"you haven't granted it yet"* and the vendor gives up | | `--setting-sources ""` | stops the user's own permission rules pre-approving the call | measured on a machine allowing `Bash(mkdir:*)`: `mkdir zzz` **ran ungated** and the directory was created | **The third flag became a FALLBACK on 2026-08-12** and the table row is the record of why it was ever needed. Council injects its own `PreToolUse` hook instead, which runs at step one and beats an allow rule at step five, so the operator's settings stay loaded. The dated block at the end of this section carries the measurements and the build. The third is the honesty of the whole feature, and it is the one nobody would have thought to test. Permission *allow rules* in settings files are consulted **before** the callback, so a call they cover never reaches the gate at all. Without that flag, "nothing writes without your keystroke" is simply false — and false quietly, on a machine whose owner wrote those rules years ago for a different purpose. **Amended 2026-08-11, at the end of this section.** That paragraph is still true and it is still the record of 2026-08-04. It reads one step of the evaluation, not the step above it: an `ask` rule is consulted BEFORE an `allow` rule, and it reaches the callback the allow rule would have skipped. So the third flag is no longer the only way to be honest. The measurement and the build it implies are below. **One limit is stated on the badge rather than buried here.** Shell commands the CLI itself classifies as read-only are approved without asking — `git status` was ungated under both setting-source configurations, and so is `echo` — so the claim is about calls that *change* things and is worded that way everywhere. **RETIRED 2026-08-12, and kept because it is the record of a hole that existed for eight days.** Everything in the next four paragraphs describes council copying the USER's hooks into the ephemeral file. The seat no longer drops their settings, so there is nothing to copy and a copy would run every one of their hooks twice. The ephemeral file survives, built the same way and for the same reason, carrying council's own gate hook instead. **The second limit was a hole, and it is now closed.** Dropping the setting sources also dropped the user's own hooks and user-level commands from that seat. Half of that is the feature working: the allow rules are what the gate replaces. The other half was collateral — a `PreToolUse` hook is a screen the user built, nothing was replacing it, and the calls it covered are disproportionately the ones the gate never sees. Measured: in the gated posture, `echo ` raised no request and simply ran. `--settings ` composes with `--setting-sources ""` — the sources stay dropped and the named file is still read — so council copies the user's `hooks` section into an ephemeral file of its own and points the gated seat at it. Two properties of that file are load-bearing: - **It is built by naming one key, never by deleting others.** The same spike showed a `permissions` block inside a `--settings` file re-admits the allow rules: an allowlisted `mkdir` ran with no request and the directory landed on disk. An allowlist of exactly `hooks` cannot rot as Claude Code adds settings keys; a denylist would. - **The badge is derived from whether the file exists**, not from whether council tried. An unreadable settings file, an empty hooks section and a temp directory that could not be created all end in the same place, and the column says the guard is absent rather than claiming one. The file is absolute (a relative `--settings` path resolves against the *child's* working directory, which is the workspace, and fails), 0600, removed on teardown, never logged and never rendered — the same privacy discipline `--brief` carries, for the same reason: only a boolean crosses onto `State`. The honest residual: hooks fire as that file described them at spawn time. Editing the real settings mid-session does not propagate until the next room. **A denial is not a failure, and the difference took a fifth outcome value to keep.** The vendor reports a refusal as an `is_error` tool_result carrying council's own refusal text back — so read off the stream alone it is indistinguishable from a tool that broke, and the trace would say the command *failed* when what happened is that it was *not allowed to run*. `ActDenied` is recorded from the keystroke, before the echo arrives, and the echo cannot overwrite it. It renders `✗ denied by you`: the words carry the distinction, colour only seconds it, and it is the one line in the trace that is not a reading of a vendor's words. **The gate is Claude-only, and that is a fact about the other CLIs.** `codex exec` and `agy -p` are batch programs — read a prompt, answer, exit. Neither has a channel a question could arrive on. Their columns keep `WRITES`; only the seat that asks carries `gated`. Giving all four the same badge would be the blanket claim §9.2 exists to refuse, one level up. **Amended by §9.36, and the amendment is narrower than it looks.** The Cursor seat is now a live ACP process and it *can* be asked: `session/request_permission` blocks it until answered, measured on both branches. It still does not carry `gated`, because it does not ask about EDITS — measured twice, it wrote a file and raised nothing. So the last sentence above holds with one word changed: only the seat that asks about **everything that changes anything** carries `gated`. Council answers Cursor's requests all the same, because an unanswered one blocks the vendor forever. `--write --auto` restores the old behaviour for the times nobody is watching: `acceptEdits`, the `WRITES` badge, the user's settings left alone — and therefore no injected hooks file either, since a room that loads those settings natively would otherwise run every hook twice. Gating is the default because the room the user opened is the one they are looking at; unattended is the exception and has to be typed. #### The gate can keep the user's settings, measured 2026-08-11 — and the build is not authorised **Superseded by the block below it, 2026-08-12, and kept whole.** Nothing changed in the product on this date; this is the record of the probe and of the ruling it waited for. The ruling came the next day — measure the two open unknowns, build only if both support it — and both did. **The claim.** Claude Code evaluates a tool call in six steps, in this order: hooks, deny rules, ask rules, permission mode, allow rules, then the `canUseTool` callback. That is the live documentation of 2026-08-11 (`code.claude.com/docs/en/agent-sdk/permissions`), not memory. `ask` is step three and `allow` is step five, so an `ask` rule reaches the callback that an `allow` rule would have skipped. The same docs say a `PreToolUse` hook may return `permissionDecision` `"allow"`, `"deny"`, `"ask"` or `"defer"`, and that `ask` beats `allow` when both apply. If either holds in the binary, the gate can keep `--setting-sources` and gate anyway — and keeping them carries the user's own hooks, deny rules and user-level commands back in natively, which is the whole of what this seat gives up today. **The rig, because this repo does not read a claim off a doc.** A probe replicated this seat's own argv — `baseArgs` plus `Session`, gated posture, `--model haiku` added to keep the turns cheap — spawned it against a throwaway directory, wrote one turn on stdin as `Turn()` builds it, and answered any `can_use_tool` request with `behavior: "deny"` as `Decide()` builds it. `--setting-sources ""` was dropped on every arm except the one that reproduces what council ships. **The decisive observable is the filesystem, never the stream**: the brief asked for one command that creates a marker, and the arm is read by whether the marker is on disk. Claude Code 2.1.226, Windows 11, `claude-haiku-4-5` on every turn, two trials each. | arm | what it changed | requests | marker created | trials | |---|---|---|---|---| | **A** adopter | zero-rule `CLAUDE_CONFIG_DIR`, no `--setting-sources` | — | — | **blocked** | | **A2** | user settings live, `touch probe-marker` | 0 | **yes** | 2/2 | | **A3** | user settings live, `install -d probe-marker` | 1 | no | 2/2 | | **B1** | user settings live, `mkdir probe-marker` | 0 | **yes** | 2/2 | | **B2** | B1 plus an injected `ask` rule for `Bash(mkdir:*)` | 1 | no | 2/2 | | **C** | user settings live, injected `PreToolUse` hook returns `"ask"` | 1 | no | 2/2 | | **C control** | same hook wiring, hook returns no decision | 0 | **yes** | 2/2 | | **shipped** | council's argv today, `--setting-sources ""` | 1 | no | 2/2 | **B1 says the 2026-08-04 finding still reproduces.** An allow rule covers `mkdir`, the call ran, and the directory is on disk. Nothing here retires the original measurement. **B2 is the decisive arm.** One rule was added — `{"permissions":{"ask":["Bash(mkdir:*)"]}}` in a `--settings` file — over settings that already allow the same shape. The call raised a request, the denial was honoured, and nothing was created. The request named its own cause: `"decision_reason_type":"rule"`. **C says a hook can do the same job, and says more on the way through.** A hooks-only `--settings` file whose `PreToolUse` hook returns `permissionDecision: "ask"` gated the same allow-covered call, and the request arrived carrying `"decision_reason_type":"hook"` with the hook's own sentence in `decision_reason`. The hook wrote a breadcrumb on every trial, so its run is provable off the stream. **The control matters as much as the arm**: C changed two things at once, a `--settings` file and an "ask" behind it, so the same file was run again with a hook that returns nothing. The call went ungated and the directory landed. The decision causes the gate, not the file. **Two facts fell out that were not the question.** `--settings` composes with the user's settings rather than replacing them — every sources-live arm ran the user's own `SessionEnd` hooks, including the arms passing `--settings`, and the shipped arm ran none. And a write shape no rule covers already reaches the gate with sources live (A3), so today's flag is not what makes the gate fire; it is what makes it fire *uniformly*. **What build this implies, and it is the hook rather than the rule.** A2 is why. `touch` creates a file and it ran ungated on both trials, so the user's rules cover more shapes than anyone would enumerate, and an `ask` list built shape by shape leaks exactly the way an allow list leaks. That is the same defect `hookset.go` (now `gatehook.go`) already refuses by naming one key instead of deleting many. A matcherless `PreToolUse` hook has no list to leak: the documentation's own advice for a check that must run on every tool call is a hook, for this reason. So the shape to build is council injecting its own `PreToolUse` hook that answers `"ask"`, into the same ephemeral `--settings` file it already writes, and **dropping `--setting-sources ""`** — which returns the user's deny rules, their user-level commands and their hooks to the gated seat, and retires the hooks copy in `hookset.go` (the file became `gatehook.go`) along with the badge that reports whether it worked. **What is NOT settled, and each item is a reason the build waits.** - **The adopter arm did not run.** A temporary `CLAUDE_CONFIG_DIR` holds no credentials, so the turn died at `"apiKeySource":"none"` and `Not logged in · Please run /login` before any tool call. Copying a credential store into a probe directory is a redline, so the arm stays unrun. The shipped arm is the nearest evidence for the same question: with no rules in force at all, the call reached the prompt on both trials. - **A matcherless hook was never measured.** This rig measured a `Bash` matcher. The claim that one hook sees every tool call is documentation, and documentation is what this section exists to distrust. - **Composition with the user's own `PreToolUse` hooks is unmeasured.** The docs rank `deny` over `defer` over `ask` over `allow` when several apply. Council would be adding a second hook to a file the user also populates, and the credential guard is exactly the hook that must not be weakened by the addition. - **A hook is a process per tool call.** The seat that was re-founded to stop paying process cost per turn (§9.33, §9.36) would take on a spawn per call. Nothing here timed it. - **The badge's sentence would have to change.** Today it claims a guard because a hooks file exists. Under this build the guard IS the gate, and "the user's hooks are carried" stops being a separate claim — a badge that kept saying it would be reporting a file that no longer does that job. #### The two deciding measurements, and the build, 2026-08-12 The owner ruled: measure the two items above that decide the build, then build only if both measurements support it. Both did, and the build is in. Claude Code **2.1.228** (the box moved on from 2.1.226 between the two dates), Windows 11, `claude-haiku-4-5` on every turn, two trials per arm, throwaway directories, the same rig as the block above — a probe replicating this seat's own argv, one turn written on stdin as `Turn()` builds it, `can_use_tool` answered as `Decide()` builds it, and **the decisive observable is the filesystem, never the stream**. The adopter arm stays unrun for the same reason: copying a credential store into a probe directory is a redline. **(a) A matcherless hook fires for every tool shape, and the ask reaches the card.** The arm kept the operator's settings live — no `--setting-sources ""` — and injected one `PreToolUse` entry with no `matcher` field, returning `permissionDecision: "ask"`. | arm | `mkdir gate-a` (an allow rule covers it) | `install -d gate-b` (no rule covers it) | `Write gate-c.txt` (not a shell command) | on disk | trials | |---|---|---|---|---|---| | **M** matcherless hook returns `"ask"` | request, `hook` | request, `hook` | request, `hook` | nothing | 2/2 | | **M control** same file, hook returns no decision | **no request, directory created** | request | request | `gate-a` | 2/2 | | **shipped binary** the file council now writes, `telltale hook gate` | request, `hook` | request, `hook` | request, `hook` | nothing | 2/2 | Every request carried `"decision_reason_type":"hook"` and the hook's own sentence in `decision_reason`, forwarded verbatim to the `can_use_tool` card. **The control is what makes this a finding**: the same file, the same hook process running — its breadcrumbs prove it — and only the decision removed. `mkdir` went ungated and the directory landed, which also re-reproduces the 2026-08-04 bypass on 2.1.228. The decision causes the gate, not the file. **The matcher forms were measured against each other** in one turn, four entries side by side writing to four breadcrumb files: **matcherless, `"*"` and `""` each saw both the `Bash` call and the `Write` call; `"Bash"` saw only the `Bash` call.** That one hook sees every tool call was documentation until this turn. The absent field is what ships, of the three equivalent forms, because it is the only one that cannot later be read as a pattern somebody should widen. **A Windows trap, and it is the worst failure this feature has.** Claude Code hands the hook command to **`/usr/bin/bash`** — Git Bash, on the platform this product primarily targets. The first three arms measured nothing because bash ate every backslash of a native Windows path: ``` /usr/bin/bash: line 1: C:UserssanleAppDataLocalTempclaudeC--…askhook.exe: command not found ``` `exit_code: 127`, `outcome: "error"` — and **a hook that fails to run makes no decision, so every call ran ungated while the badge went on claiming a gate**. It was found only because a `SessionStart` hook was planted in the same file to prove the file was read at all, which costs no model turn. Council quotes the command and swaps the separators; `TestTheHookCommandSurvivesGitBash` pins both. **The first fix for it was wrong, and the Linux CI job is what said so.** `filepath.ToSlash` is a **no-op on Linux**, where a backslash is a legal filename character, so it made the conversion depend on the host Go compiled for. The string is not read by the host — it is read by bash, on every platform, where a backslash is the escape character and cannot survive as itself. The swap is now unconditional, and the test feeds a Windows path on every runner. **(b) The hook costs tens of milliseconds, and the operator's own settings cost more.** The measure is the interval between the assistant's `tool_use` block landing on stdout and the `can_use_tool` request landing — the window Claude Code evaluates permissions and runs hooks in. | arm | per-call gap, all samples (ms) | median | trials | |---|---|---|---| | **shipped today** — `--setting-sources ""`, no hooks at all | 23.0, 12.3, 21.2, 10.5, 6.8 | **12.3** | 2 | | operator's settings live, **no** council hook | 284.6, 280.5, 291.0, 234.6 | **282** | 2 | | live + council's hook, a 3 MB probe binary | 358.9, 351.7, 249.5, 348.4, 298.4, 258.4 | **323** | 2 | | live + council's hook, **the shipped 14 MB `telltale.exe`** | 523.6, 491.2, 461.2, 451.8, 439.7, 362.7 | **456** | 2 | Read the rows against each other rather than against zero, because **most of the delta is not the hook**. Loading the operator's settings at all costs ~270 ms per call — that is their own `PreToolUse` hooks running, and it is the thing this build BUYS, not a price it adds. Council's own hook adds **~41 ms** as a small binary and **~174 ms** as the shipped one. Measured directly, outside the CLI, 20 spawns through the same Git Bash: the small binary is **36.2 ms median**, `telltale.exe` is **54.4 ms median** — the 14 MB single binary links the TUI framework on a path that runs once per tool call, which is ADR-002's statusline argument arriving at a second door. Against a warm Claude turn of 6.4 s (`STATE.md`'s traced `@all`), three gated calls add ~0.5 s. That is the owner's "tens of milliseconds is fine" band at the binary's own cost and inside it at the process's; it is nowhere near the "a second per call" that fails. **What shipped.** Council writes an ephemeral `--settings` file containing exactly one key, `hooks`, holding one matcherless `PreToolUse` entry that runs `telltale hook gate` — a new mode beside `telltale hook cursor`, which drains stdin and prints one decision object and nothing else. `--setting-sources ""` is **dropped**, so the gated seat loads the operator's deny rules, their user-level commands and their own hooks again. `--permission-mode manual` is **kept**: the documentation makes it an alias for `default` on 2.1.200+, which would make it decoration, but every arm of both probes carried it and nothing has measured the seat without it — the same rule that kept `--permission-prompt-tool stdio` when it was absent from `--help`. **The fallback is the old build, not a hole.** A room whose hook file cannot be written — no temp directory, a binary that cannot locate itself — passes `--setting-sources ""` and gates the 2026-08-04 way. It gives up the operator's settings and the column says so. Weaker in what it keeps, never weaker at the gate. **One cost was not on the list of five, and it changed the room.** The hook asks about EVERYTHING, which is the point — and `Read`, `Glob` and `git status` raised **no request at all** under the old flag, because Claude Code approves what it classifies read-only before the callback. Under the hook all three raise one. Shipping only the hook would have tripled the cards, and this room already knows what that costs: the first session with the gate carded the user thirty-four times, which is why `autoApproveRoutine` exists. So council answers them itself — `autoApproveRoutine` for shell commands, and a new positive list of tool names that change nothing (`Read`, `Glob`, `Grep`, `NotebookRead`) for the calls that are not shell commands. Positive, so a tool Claude Code adds next month draws a card rather than being waved through; `TodoWrite` is deliberately absent, because nothing here measured what it writes. **The badge's sentence changed, as the fifth item predicted.** It no longer claims the operator's permission rules are dropped, because they are not. The wired branch says their settings stay loaded and that council's own hook asks first; the fallback branch says the settings were dropped and why. Both branches still say nothing runs until you answer. **Of the five unsettled items, two are settled, two are retired by the build, and one remains.** The matcherless hook and the per-call cost are measured above. The badge sentence and the `--setting-sources ""` question are decided by what shipped. **Composition with the operator's own `PreToolUse` hooks is still unmeasured** — council now adds a second hook to a file the operator also populates, the docs rank `deny` over `defer` over `ask` over `allow` when several apply, and the credential guard is exactly the hook that must not be weakened by the addition. The ranking makes a weakening unlikely (a `deny` from their hook beats council's `ask`), and "unlikely by documentation" is the standard of evidence this section exists to distrust. **The per-call cost tolerance is HELD — owner ruling, 2026-08-15.** The measurement above stands: the shipped 14 MB binary costs ~54 ms per gated call because the one binary links the TUI framework, and a sibling hook binary would save ~133 ms per gated call. The owner ruled the saving does not buy an ADR-002 amendment: roughly half a second across three gated calls on a 6.4 s turn sits inside the "tens of milliseconds is fine" band the build was accepted under, and the split's real price is not the binary — it is the packaging (the scoop manifest ships one exe), `hookCommand`'s self-location, room/hook version skew, and a missing-sibling fallback whose only honest shape is `--setting-sources ""`, which trades away the operator's settings to save milliseconds. The one-binary shape stands unamended. Reopen this only with a measurement that changes the arithmetic: more gated calls per turn than the ~3 assumed, or a vendor change that raises the per-call floor. #### The operator's deny beside council's ask, measured 2026-08-15 — the last item closes **The last unsettled item of the five is now measured, and the operator's guard does not come back weaker.** The item said this: council adds a `PreToolUse` hook to a file the operator also populates, the documentation ranks `deny` over `defer` over `ask` over `allow`, and the credential guard is the hook that must not be weakened. The rank held live. The deny also does more than win — it stops the call before council's card is drawn at all. **The environment, and the two claims it makes into hypotheses.** Claude Code **2.1.228** (`claude --version`, recorded before any arm), Windows 11, `--model haiku`, two trials per arm, throwaway directories. The floor for this measurement is 2.1.221: that release fixed auto mode overriding a hook's `ask`, and 2.1.222 fixed exit code 2 not blocking. Both are changelog claims, so both are hypotheses here, and both are confirmed by the arms below. A measurement taken before 2.1.221 would be void. **The rig is the shape of the two blocks above.** A probe replicates this seat's own argv (`baseArgs` plus `gateArgs` plus `--input-format stream-json` plus `--settings`), writes one turn on stdin as `Turn()` builds it, answers `can_use_tool` as `Decide()` builds it, and **reads each arm off the FILESYSTEM, never off the stream**. The brief asks for one command, `mkdir deny-probe-marker`, which the operator's own allow rules already cover. No credential store is copied anywhere: `CLAUDE_CONFIG_DIR` is untouched, the operator's real login is used, and `~/.claude/settings.json` is never edited. **The deny hook has the credential guard's SHAPE and none of its content.** It is a `PreToolUse` entry with matcher `"*"` that writes a reason to stderr and exits 2 — the mechanism the real guard uses (`Exit 0 = allow, exit 2 = block`). It is installed in the THROWAWAY workspace's own `.claude/settings.json`, which is a real setting source and belongs to nobody. It writes a breadcrumb on every run, so its execution is provable off the disk. **A wiring probe costs no model turn, and it is what makes the arms readable.** The seat was started and its stdin closed without a turn. A `SessionStart` breadcrumb from the workspace's own settings arrived while `--settings` carried council's gate file, so the two surfaces demonstrably compose. That is the same `SessionStart` trick that found the Git Bash trap above. | arm | hooks in front of the call | callback answer | requests | marker on disk | trials | |---|---|---|---|---|---| | **(a) control** | operator `deny` alone | allow | 0 | no | 2/2 | | **(b) the question** | operator `deny` + council's gate hook | **allow** | 0 | **no** | 2/2 | | **(b) again**, `--include-hook-events` | same | allow | 0 | no | 2/2 | | **(b control)** | operator hook exits 0 + council's gate hook | allow | 1 | **yes** | 2/2 | | **(c) anchor** | council's gate hook alone | deny | 1 | no | 1/1 | **The callback answers ALLOW in arm (b), and that is the whole design of the arm.** A denial pressed at the card would leave nothing on disk whatever the hooks did, so it cannot tell a holding deny from a broken one. Answering `allow` inverts the test: if council's `ask` had displaced the operator's `deny`, the marker lands. It did not land, on either trial. **The (b) control is what makes (b) a finding.** One thing changed — the operator's hook exits 0 instead of 2 — and the marker landed on both trials. So the pipeline can create it, council's `ask` still fires when nothing denies, and arm (b)'s empty directory is the deny. **Both hooks ran, concurrently, and the stream says so.** `--include-hook-events` is observability only and is NOT part of the seat's argv; it was added to one arm to see who answered: ``` hook_started PreToolUse:Bash hook_started PreToolUse:Bash hook_response PreToolUse:Bash exit 0 outcome success stdout {"hookSpecificOutput":{"hookEventName":"PreToolUse","permissionDecision":"ask","permissionDecisionReason":"telltale council gates every call in this room"}} hook_response PreToolUse:Bash exit 2 outcome error stderr BLOCKED by the operator credential-guard-shaped probe hook: this command matches a protected pattern. ``` **The card never appeared, and the model reads the operator's sentence rather than council's.** No `can_use_tool` request was emitted on any (a) or (b) trial. The turn received: ``` PreToolUse:Bash hook error: [python3 "…/deny_hook.py"]: BLOCKED by the operator credential-guard-shaped probe hook: this command matches a protected pattern. ``` That is stronger than the ranking predicted. The documentation says `deny` beats `ask`; what runs is that a `deny` ends the evaluation, so council never gets a question to draw. **The room's gate cannot weaken the operator's guard, because the guard answers first and the gate is never asked.** **Both changelog claims are confirmed at 2.1.228.** Exit code 2 blocked on every arm that used it, six trials in total, which is 2.1.222's fix behaving as described. The hook's `ask` reached the card rather than being auto-answered in arms (b control) and (c), which is 2.1.221's. Neither is read off the changelog now. **One stream trap re-appeared and is worth the line.** On a turn where nothing was created, the CLI's own `post_turn_summary` said `"status_detail":"mkdir deny-probe-marker executed"`. A rig that read the stream would have recorded the opposite of what happened, which is why every arm in this section is read off the filesystem. **What this did NOT measure, itemized so nobody reads it as wider than it is.** An operator hook that denies by printing `permissionDecision: "deny"` as JSON, rather than by exit code 2 — the real credential guard uses exit 2, so the rig measured the shape that ships on this machine. `defer`, which no hook here returns. And the adopter arm, which stays unrun for the reason it has always stayed unrun: copying a credential store into a probe directory is a redline. **The cursor seat was measured the same night and the finding is not about deny beating ask.** It settles whether that seat's ACP path carries hooks at all, which three records in this repository disagreed about. §7.16's amendment of the same date carries it. **The agy arm was SKIPPED, and a skipped arm stated is better than an unmeasured claim.** The proposal was to read `agy -p "/hooks"`'s registration table if that probe is zero-cost. It could not be confirmed zero-cost: `agy --help` lists no local hooks subcommand, `-p` is print mode, and `--disable-slash-commands` exists precisely because slash commands are expanded into the prompt. So the probe would spend a turn from a constrained pool and would return a model's description rather than a registration table. That arm proves wiring and order only; it could never speak to deny-beats-ask semantics. ### 9.9 The room remembers — a conversation, not a ticker Everything above builds a very good way to **send one turn**. What it did not build is somewhere to have a conversation, and the gap was structural rather than cosmetic: dispatching turn N cleared turn N-1's body, trace, clock and cost off the screen, and the user's own words were never rendered anywhere at all. So the room could show you four answers to a question it could not show you, and then throw them away when you asked the next one. This was PR 4 of council's original plan, ratified and then skipped. **The transcript is per column and per turn.** A finished turn is *pushed* to `Column.History` rather than erased: `TurnRecord` carries the brief, the reply, the trace, the note, the elapsed and the cost, and `columnText` renders the whole list oldest-first, each turn opened by a separator naming it. Three consequences fall out of it being one flat list of lines: - **The scrollback needed no second mechanism.** The window, the overflow markers, the tail and `MaxScroll` were already the code that moved through a column's lines; a transcript is just more of them. `g` reaches the first thing that seat was ever asked. - **Each past turn carries its own numbers.** The column header and the badge line are chrome describing the turn *in flight*, so a turn scrolled back to would otherwise sit under someone else's clock. A turn that ended badly names its phase on its separator; a turn that ended normally does not, because "done" on every one of them is noise on the common case and makes the two that matter harder to find. A running total keeps its `session` word (§9.8) — losing it in history would turn a true figure into a false one. - **A seat that sat out a turn records nothing for it.** Routing means turn 4 can go to Claude alone, and Codex's transcript then skips from 3 to 5. Filling that gap would be the room inventing a conversation. **The echo is the principal's words, and that is a boundary rather than a phrasing.** What a seat literally receives is not what is echoed: a first turn is sent with the `--brief` file prepended, whose content is deliberately held off `State` (§9.8's privacy discipline), and a rebuttal turn is sent with the other seats' answers fenced in front of it, which are other vendors' words. Echoing "exactly what was sent" would have put a private file on screen and labelled another model's output as the user's. So the brief is echoed, marked with the same `›` the composer uses — the glyph carries it before the colour does — and what rode along with it is *reported* on its own muted line. It goes through `sanitize` like everything else that reaches `State` and is **not** redacted: this is the user's own typing echoed to the user on the user's own screen, so covering it would hide a secret from the one person who already has it, do nothing about the copy just sent to four vendors, and make the echo disagree with what was dispatched — which is the one thing the line exists to show. **Memory is capped at 50 turns per column**, oldest out first, and nothing is written to disk. The room file (`~/.telltale/council/room.json` — one global room, the workspace a mutable field inside it; ADR-008, ninth and eleventh amendments) stays keys-only: it holds session ids and no content, and scrollback is not state worth persisting. **The composer grows to six rows.** One elided line was not somewhere anyone could compose a brief worth sending to four agents. `ctrl+j` inserts a newline; `enter` still dispatches, and the mode line says both. The newline goes into the draft **raw**, deliberately bypassing `sanitizeKeepingSpace` — that filter exists so a *pasted* newline cannot tear the footer apart, and it still does exactly that; what it must not do is flatten the one the user asked for by name. A deliberate newline survives to every transport this repo drives: Claude and Codex take the prompt on stdin, Claude's persistent turn is JSON-marshalled so the newline is escaped in the envelope, and agy takes it as a single argv element on a native binary with no shell anywhere in the path (§9.3) — Go quotes an argument containing a newline, and the ~32K Windows command-line ceiling is unchanged. A draft taller than the ceiling keeps its tail, where the cursor is, and spends one row saying how much is above it, in the same words the column overflow markers use. **A dead seat stops eating the width.** A column whose availability is `NotInstalled` or `Unusable` held a quarter of the terminal for the whole session to display one card that never changes — on the reference machine that is Cursor, permanently. Those seats now fold out of the grid and the survivors take the width. What must not fold away is the *fact*: one muted line under the header names each collapsed seat and which failure it is, keeping §4a.1's distinction between "not installed" and "installed but not drivable" intact at one line instead of one column. A seat nobody can see is one a user has no reason to go looking for, which makes silent collapse worse than the column it replaced. `--vendor` is the explicit control, and it mirrors the HUD's flag while doing more than filter: `all` keeps every detected seat on screen, and a comma list seats exactly those — drawn **and** dispatched to, since drawing a seat you cannot see while spending its quota is the same class of hidden state this product exists to refuse. Naming a seat forces it on screen even when it is absent, because a user who asked for it is owed the card explaining why it is not there. It parses the `@mention` vocabulary rather than a second one, so `--vendor agy` and `@agy` are the same word, and `Seated()` counts only seats that are both drivable and in the room so the header's `3/4 seated` keeps meaning what it says. ### 9.10 A mode that could not scroll, and the mouse it did not get The room shipped with a per-column scrollback, page keys, `g`, `G`, overflow markers that count what is hidden, and a full-width expand — and was reported as having **"no way to scroll up or down if the output that each agent provides is long."** The report was correct about the experience and wrong about the cause, which is the interesting part: none of that machinery was missing, and all of it was unreachable. `turnColumnFinished` puts the room in **compose** when the last column lands, so the mode the user is in at the moment four long answers arrive is the mode that reads keys as text. `composeKey` forwarded the six keys it recognised and dropped everything else into a branch that appends `msg.Text` — and an arrow key carries no text, so every scroll key did nothing at all, silently, from the one moment they were wanted until the user guessed at `esc`. The keys were not absent; they were being swallowed by a rule written for letters. **The rule is now the test rather than a list.** A key that carries no text cannot *be* composer text, so it keeps the meaning it has in view mode: `↑`, `↓`, `pgup`, `pgdown`, `tab` and `shift+tab` are shared between the modes through one function, which is what lets the mode line promise them without a second implementation to keep in step. The letter aliases stay view-only, because in compose `j`, `k`, `g`, `G` and space are the letters j, k, g, G and a space — the same rule that keeps `q` the letter q here. `tab` had to come with the scroll keys rather than after them. They address the *focused* column, so a mode that can scroll and cannot change which column it scrolls can only ever read whichever seat happened to be focused when the turn ended. `left`, `right`, `home` and `end` are deliberately still dead in compose. They are where an in-draft cursor goes if the composer ever grows one, and spending them on focus now would make that a change to muscle memory rather than an addition. **The overflow marker names its keys, on the focused column only.** `↑ 53 more above` told a reader that something was hidden and nothing about how to see it. It now carries the keys that would move it — but only on the column those keys address, because the same hint on the three seats beside it would be three false claims. The hint is mode-aware: `f expand` is dropped while composing, where `f` is the letter f. A marker that advertised a key the current mode does not have is precisely the dishonesty §7.8's always-on mode line exists to prevent. It is also dropped, `f` first, when the cell cannot hold both — the count is never traded for the hint, and at a three-seat room's 37 cells the short form is what fits. The same honesty now runs the other way as well, in §9.11: `f` and `tab` are dropped from the mode line *and* the marker in a room with a single seat on screen, because expanding the only column to the width it already has and cycling focus around one seat are both nothing happening. A key that does nothing is as much a false promise as a key that goes unnamed. #### Mouse wheel scrolling: rejected, with the measurement **Measured**, against the compiled `charm.land/bubbletea/v2` v2.0.8 by running a program per mode and reading the bytes it wrote: | `View.MouseMode` | emitted on enter | |---|---| | `MouseModeNone` | nothing | | `MouseModeCellMotion` | `ESC[?1002h` `ESC[?1006h` | | `MouseModeAllMotion` | `ESC[?1003h` `ESC[?1006h` | Those three are the whole enum. **There is no wheel-only mode**, and there is no DEC mode that would provide one: under 1000, 1002 and 1003 alike the wheel is reported *as buttons 4 and 5 inside button reporting*, so a program cannot ask for the wheel without also claiming the left button. 1002 is button-event tracking — press, release, and motion while a button is held. **Inferred** from that, and from Windows Terminal's documented behaviour rather than from a run: while 1002 is set, a left-press and drag belongs to the application, so the terminal's own text selection is suppressed unless the user holds the bypass modifier (shift). That is the trade, and it is a bad one **for this room specifically**. Council exists to put four vendors' answers side by side so they can be read and taken away; making the answers harder to select with a mouse in order to make them easier to scroll with a mouse spends the product's output to buy a convenience for its input. The keyboard path is complete — it is what the rest of this section fixed — so the wheel would add no capability at all. §7.8 already records "deliberately absent: mouse support" for the gauges; council is a different surface and got the question asked again on its own terms, and the answer came back the same. Recorded rather than left as a gap, because "nobody tried" and "it was measured and refused" are different facts, and this repo does not let them render alike. ### 9.11 The room was correct and it was flat §9.10 fixed the last thing that did not *work*, and the room was driven live the same day. The report back was three words — *"where are the UI updates?"* — and it was right. Every sentence in this section so far is about what the room says; none of them is about how it looks, and the accumulated answer was a surface with one typographic level in it. A seat's name, a safety claim, a key you can press and four hundred lines of vendor prose all arrived at the eye with exactly the same emphasis, separated by nothing but two horizontal rules three rows apart. Everything was true and nothing was findable. The rule this section is written under is §7.1's second: **every distinction is carried by a glyph, a word or a number FIRST, and colour only reinforces it.** That is what makes `--ascii` and `NO_COLOR` correct by construction, and it is also, read the other way, a budget: a surface that may not lean on hue has to earn hierarchy from shape, position, weight and air. Council had spent almost none of that budget. What follows is the pass that spent it, and the constraint every item was checked against. **No colour was added, and that was not a close call.** The palette is still exactly `Text / Muted / Identity / SevOK / SevWarn / SevCrit` from `internal/theme` (§7.5), and `internal/theme` was not touched — it is shared with the stdlib-only statusline binary (ADR-002), so a token added for a TUI would be a coupling paid for by a binary that cannot use it. What *was* added is **weight**, which is an attribute rather than a hue: `Strong` is Identity at full weight and `Alert` is SevWarn at full weight, and `PlainStyles` renders both as the identity function. That last property is the whole reason weight is safe here — it changes no cell's width and no line's content, so every layout golden is blind to it and nothing it marks is the sole carrier of anything. **The state a seat is in is now a shape.** `done` / `failed` / `cancelled` / `idle` / `unavailable` were five words at the far right of a 37-cell column, told apart by the word and by a colour behind it, in a room that holds three of them side by side. Each now leads with a mark, and the vocabulary is deliberately a **reuse of meanings this room already owns** rather than a second alphabet: | phase | mark | ascii | where the meaning comes from | |---|---|---|---| | idle | `○` | `.` | the only new glyph — the HUD's own weakest-state dot, with the HUD's own ASCII form (§7.5) | | waiting / streaming | spinner | `-\|/` | already sat in this slot; it is now the in-flight member of one vocabulary rather than a special case | | done | `✓` | `+` | the trace's own "the vendor reported this worked", said about the whole turn | | failed | `✗` | `x` | the trace's own "the vendor reported this broke" | | cancelled | `⚠` | `!` | what a note and the unavailable card already open with: *this did not complete normally* | | unavailable | `⚠` | `!` | same claim, and the **word** is what separates it from cancelled | Two of those share a mark, and that is the design rather than a collision. glyphs.go argues at length that a character already spoken for is not a mark — but that argument is about a character meaning two *different* things, and here it means the same thing twice. The distinction between "cancelled" and "unavailable" is carried by the word, which always renders, in both glyph sets and with colour off. `TestPhasesRenderAsDistinctMarks` fails the build if any two states ever render alike; `TestPhaseMarksSurviveASCII` fails it if the one new character collides with anything already claimed. **Rule 4 is untouched**: the spinner is still the room's only moving cell, because none of the other four marks move. **The column header is one line instead of two labels.** `▸Claude Code` at the far left and `idle` at the far right with twenty-five dead cells between them reads as two unrelated things, which is what it was. The name now takes full weight — it is the anchor a reader scans for — the state leads with its mark, and the gap between them is filled with a rule. The rule is not decoration: it is **this room's existing grammar for "a label and the numbers that belong to it"**, which is exactly what `turnRule` has always drawn for every turn in the transcript underneath. The live turn's header and a finished turn's separator are now the same line form, so a reader learns one shape rather than two, and the header loses the ability to read as two things at once. It degrades in the right order: below the width where a rule fits, the **name** is truncated and the state is kept, because a clipped seat name is still recognisable and a clipped state word is not. Two cells of air each side of that rule, not one, and the reason is `--ascii`. The ASCII rule is `-` and the ASCII spinner's first frame is also `-`, so a streaming column at one cell rendered `------------ - streaming` and the mark disappeared into the rule pointing at it. Two cells is also what this product already puts around the `│` that separates zones in the header and the mode line, so the fix and the convention are the same number. **One rule per column instead of two three rows apart.** The full-width rule under the room header stays — it is the HUD's anatomy (§7.2) and council is meant to be the same product — and the per-column rule under the badges is gone. It was the weaker of the two and it was doing almost nothing: by the time the eye reached it, it had been told nothing the two lines above had not already said. The header now carries a rule of its own, which separates the seat from its content in the same gesture that binds its name to its state. **The row was not reclaimed for the body; it was spent on a blank one**, because what the reading area needed was air between chrome and content, and a blank line separates two blocks more quietly than a second horizontal line does. That row is *reserved* even for a seat with no posture to state, so the bodies of three columns start on the same screen row. A grid whose rows do not line up is a worse trade than one empty claim slot — and `MaxScroll` no longer subtracts a literal 3 for the chrome. It measures the chrome by drawing it, from the same function the renderer uses, which fixes an off-by-one that was already live on a column with no badges. **The badge line looks like a claim instead of like debug output.** It is unchanged in what it says — `TestSandboxBadgesAreNeverBlanket` still guards that, and §9.2's argument is untouched — and changed in three ways in how it is shaped. It is indented to the seat name above it, so it reads as a property of that seat rather than as the first line of the reply; an unindented row of bare lowercase tokens at the top of a column is precisely what debug output looks like. Its cost is right-anchored, giving the two chrome rows one shape twice over: label on the left, value on the right. And the posture badge takes weight and the warning hue **when, and only when, it says this seat can change your files** — `WRITES` and `unsandboxed` are loud, `ro:*` stays chrome, `gated` takes the weight without the severity because a gated seat is the room working rather than a risk. That last one is the pass's only change with a safety argument behind it. §9.2 is emphatic that a claim you cannot see is not a claim, and then the room drew `unsandboxed` at exactly the volume it drew `ro:tools` beside it. Colour is still redundant — the words break the `ro:` prefix on purpose and are what actually carry it, which is why the plain style set renders every badge as its own bare word — but a claim a hurried reader skims past is doing half its job. `TestAWriteCapableBadgeDoesNotRenderLikeAReadOnlyOne` pins it. **The degraded columns are cards.** `⚠ Codex is not seated`, a blank, a reason paragraph at the same indent, a blank, and a closing sentence at the same indent is three fragments floating in a column: nothing on screen said the reason belonged to the title, so a three-seat room with one dead seat read as unrelated paragraphs. Every card in a column now has one grammar — **a title at weight, its body hanging under it** — and it costs no rows. The room was already doing this in the two places it needed it least (the prompt echo indents under its `›`, a failed call's detail indents under the call) and in none of the places it needed it most: the notes, the unavailable card and the approval card, where a wrapped second line started hard against the column edge and read as a new statement. On notes the mark carries the hue and the words stay plain, which is the same split the trace's outcome marks make. **The transcript reads as a conversation.** A brief and the answer to it arrived as consecutive lines at the same indent, told apart only by a glyph at the start of one of them — a distinction you have to *read*, on a surface built for comparing four answers at a glance. There is now a blank row between them, and the echoed brief takes full weight, because in a column of vendor prose the user's own words are the thing you scroll looking for. **The row is a swap, not a cost**: it came from between the turns, where a labelled full-width rule was already doing the separating. The transcript is exactly as tall as it was. What that leaves is three boundaries with three strengths, ranked: a labelled rule where the turn changes, a blank row where the speaker changes, a blank row where the kind of content changes (what the seat *did*, then what it *said*). **The footer has a figure and a ground.** Six items of identical weight separated by identical bars is a wall the eye slides off — which is, concretely, how a room with working scroll keys, page keys, `g`, `G` and a full-width expand got reported as having no way to scroll at all (§9.10). The key renders at full intensity and its label recedes, the same figure/ground split the column header makes between a seat and its state, and it costs no cells. Two items are dropped outright in a room with one seat on screen: `f` expands the only column to the width it already has and `tab` cycles focus around a single seat, and a mode line that promises a key which does nothing is §7.8's surprise pointing the other way. The gate line got the same treatment, where it matters most — the call about to be run was being drawn at the same faint volume as the keys that answer for it. Two smaller repairs fell out of the same reading. The room header now separates its own name from the workspace with the ` │ ` the HUD's header uses, instead of a bare space that made `council ~/code/telltale` read as one run-on label — and fixing that surfaced an off-by-one in the header's gap arithmetic, which had been overrunning the frame by exactly one cell and having `fit` quietly eat it. And the collapsed-seat notice is truncated with an ellipsis rather than handed to `fit`, which cuts silently: at 120 columns on the reference machine it had been losing the last word of its own remedy and looking like a sentence that stopped. **What was declined.** Hoisting the badges into the column header's gap at the single-column tiers, which would have bought a body row and killed the widest dead gulf: it makes the chrome height depend on content, so the body would grow a row at the moment a cost arrived mid-turn — a layout jump, which §7.1 rule 4 does not budget for. Giving `cancelled` a glyph of its own: nothing unclaimed in the ASCII set reads as "stopped", and the rule that a distinction may be carried by a **word** is there precisely so a glyph does not have to be invented for every case. A "role" line under each seat naming what that vendor is for (review, IDE, tiebreak): council has no such field, ADR-010's allocation is a fleet fact rather than something this room measured, and a room that stated it would be asserting something no adapter sourced. ### 9.12 The scroll keys worked; which column they moved was the thing nobody could see §9.10 fixed a room whose scroll keys were dead in the mode a finished turn drops you into. The room was then driven again, and reported as unable to scroll a **second** time: > "scrolling works for your window. i tried scrolling up/down in agy and cursor. could not." Every word of that is accurate, and none of it is a bug. The keys address the **focused** column, they have always addressed the focused column, `tab` moves focus in both modes since §9.10, and the second and third seats scroll exactly as the first does once the keys are pointed at them. `TestFocusThenScrollMovesThatColumn` says so in the product's own terms — two tabs, one `↑`, the third column leaves its tail and the first does not move — and it is kept as a test precisely so the changes below are never mistaken for a mechanism fix. **What failed is the affordance, and it failed in three places at once.** Each of them is individually defensible, which is why reading the code did not surface it: - **Three columns each said they were hiding something; one of them said how to look.** `↑ 36 more above` appeared verbatim on every column with content off screen, and the key hint rode only on the focused one — correctly, since naming `↑↓ scroll` on a seat those keys do not move would be three false claims (§9.10). The result is that the *unfocused* markers were the ones a reader was most likely to be staring at, and they named nothing. Pressing `↑` then moves a column the user is not looking at, and a scroll key that visibly does nothing is a scroll key that does not work. - **The focus marker was competing with three identical anchors.** §9.11 gave every seat name full weight, on the correct argument that a name is what a reader scans for. The cost only shows up live: with all four names at the loudest level this surface has, the entire distinction between the column the keys move and the three they do not was one `▸` glyph in a frame carrying four columns of prose. - **The compose mode line named the arrows and not the key that aims them.** §9.10 wired `tab` into compose *because* the scroll keys address one column — it says so in as many words — and then listed `↑↓ scroll` on that line without `tab focus` beside it. The one moment the user is certain to want both is the moment four long answers land, which is exactly when this line is on screen. **Three fixes, all of them words and weight, no new colour and no new key.** A marker on a column the keys do not move now names the key that would move them there: `↑ 36 more above │ tab to focus`, against the focused column's `↑ 51 more above │ ↑↓ scroll │ f expand`. This is the same rule the existing hint follows — *a marker states the key for THIS column and never a neighbour's* — applied to the case that had been left blank rather than a new rule bolted beside it. It needs none of `f`'s mode-awareness, because `tab` really does move focus in both modes; and it is empty in a room with one seat on screen, where there is nothing to tab to, for the same reason the mode line drops `f` there. The seat name's weight now says **which column the keys move**. Unfocused names keep the identity hue and give up the weight; they are still names and still legible, and they have stopped competing with the one fact that varies across the row. This needed a small type rather than a bool: `seatFocus` separates *is this column marked* from *do the keys move it*, because the two agree in the side-by-side tier and part company in the tabbed and expanded ones, where the tab bar above already carries a marker and the column beneath it is still the one being scrolled. Conflating them is what made a single call site pass `focused=false` for a column that had the keys. `tab focus` joins the compose mode line, immediately after the arrows it aims. It is offered whenever more than one seat is on screen and deliberately **not** gated on whether some column currently overflows: a hint that appeared the moment a reply grew past its column would be a footer cell that changes while output arrives, which §7.1 rule 4 does not budget for — and this line's promise is about what the mode can do, not about what the vendors happen to have said this turn. **What did not change**, because the rules that produced the original design are still the right ones. The count is never traded for a hint, in either form. No badge, no help row and no keybinding moved — the help panel already documented `tab … in compose too`, and its 17-row budget (§9.11, `TestHelpFitsTheSmallestRoom`) is untouched. Every distinction added here is a word or an attribute: `PlainStyles` renders the focused and unfocused headers identically, so every layout golden is blind to the weight, and `tab to focus` is the same string under `--ascii`. The general lesson, in this file's own terms: §9.10 recorded a mechanism that was complete and unreachable. This is the same shape one level up — a mechanism that was complete, reachable, and **unattributed**. The room said *something is hidden here* three times and *here is how to see it* once, and a user reading the two-thirds of the room that named no key concluded, reasonably, that the feature was missing. Nothing measurable was wrong; what was wrong is that the honest thing and the actionable thing were on different columns. ### 9.13 The badges were honest and nobody knew what they meant §9.2 argues that every column states its own posture, §9.11 gave the two that mean "this seat can change your files" weight and the warning hue, and twelve amendments of ADR-008 made each word behind them defensible. The room was then driven, and the report back was a question: > *"why do i care codex and agy are 'unsandboxed'? what does this mean, why are they > sandboxed, and must they remain that way? i'm really confused here."* Every previous complaint in this section was about something being wrong. This one is not. The badges are correct, they are the most carefully-argued strings in the product, and to their primary user they were **three lowercase tokens with no reachable explanation.** `unsandboxed` reads as jargon-with-a-negation, which invites exactly the two wrong readings the question contains: that the sandbox is something council switched off, and that a sandbox is what was keeping the room safe. **Two things were missing, and only one of them is a legend.** The first is the vocabulary. There was no plain-English gloss of the badge words anywhere a user could reach without reading an ADR. What existed was four muted lines at the bottom of the help panel, below the fold at a 24-row terminal, saying that each column states its own posture — a sentence about the *policy* rather than about any of the words. The second is worse and was found by grep. **`SandboxClaim.Detail` rendered nowhere at all.** It is the full argument behind each badge — what was passed, what was measured, what is therefore claimed — it is written per vendor per OS, it is asserted by tests, it is quoted into ADR-008, and no surface read it. The field's own doc comment said it was "shown in the degraded/help text". It was not shown anywhere. §9.2's rule is that a claim you cannot see is not a claim; **the argument for a claim is under the same rule**, and this one had been invisible since the badges landed. **The fix is a second help page, and the split is by kind rather than by length.** `?` now cycles: keys, postures, closed. Both pages spend the same hard 17-row budget (§9.11), both end with the `?` line that leaves them, and three presses always return the room from anywhere — the panel's one non-negotiable property, since `?` is the only documented way out of it. Page one's closing paragraph became the pointer to page two, which is what makes a second page a feature rather than a place. Page two is a legend of **every** badge this product can render, not only the ones the current room shows. A user who has never typed `--write` should be able to find out what `WRITES` means before they type it, and a room-specific legend could only ever explain the room you are already in. Each entry renders its badge through `Styles.ForSandbox`, the same function the column header uses, so the legend cannot teach one weight and the room show another; `TestEveryBadgeIsExplained` walks every `SandboxLevel` and fails the build when a badge exists with nothing here to say what it means. **Nothing was softened, and that is asserted rather than promised.** These are glosses on the badge words, never replacements for them. `TestThePostureLegendDoesNotSoftenAnyClaim` pins the load-bearing phrases — `unsandboxed` still says *nothing restricts*, *measured*, *change your files*; `ro:requested` still says *never observed* — and forbids "read-only", "safe" and "cannot write" from appearing in the gloss for any level that can write. The badges break the `ro:` prefix on purpose; a legend that put the word back would undo that in the one place a reader goes to have it explained. Below the legend, and below the fold at the 24-row floor, is this room's own seats with each one's `Detail` in full — the first time that field has rendered. The ordering is deliberate: the detail is unreadable without the vocabulary, and the vocabulary fits the budget where four paragraphs of measured prose never could. It is the same trade page one already makes with its closing paragraph, and the residual is stated rather than discovered: at the shortest terminal this room will draw in, the per-seat half is scrolled past rather than absent. **Every `Detail` was reordered so its first clause answers "so what?".** Not one factual clause was removed, weakened or added; what changed is which end of the sentence the consequence sits at. `"named write/exec tools denied and MCP servers dropped; verified against..."` opens on a mechanism a user has to decode before they learn anything, and now opens *"this seat has no write or shell tools in its session, so it cannot edit your files"* with the verification and the deny-list residual behind it. Codex on Windows opens on *"nothing at the OS level stops this column reading or writing here"* rather than on the flag that produced it. This is §7.1's rule about glyph-word-number ordering applied to prose: the distinction goes first and the evidence reinforces it. **The question's third clause got an answer too, and it is the one that mattered.** *Must they remain that way?* The badge table answers it where a first-time reader is, and the answer is not about flags: **no badge is what keeps this room out of your files — the workspace is.** `unsandboxed` on Codex is not a setting anyone chose to leave off; both sandboxed modes were measured failing every process spawn there, so read-only was a seat that could not read (ADR-008, twelfth amendment; true when written — §9.2's 2026-08-29 amendment later moved that seat's read posture to `ro:enforced` on a re-measurement, and the sentence about the workspace is unchanged by it). The control that holds is `--cd` into a throwaway worktree, and the fleet contract rules the same way independently: `agent-ops` ADR-012 rules capability parity, with guard wiring rather than lane shape as the control. A column that looked read-only because of a broken sandbox was never a safety property; it was a defect wearing one's clothes. **No posture flag moved.** Whether council should keep asking agy for `--mode plan --sandbox` when both are measured to do nothing is an open decision (§9.6b) and belongs to the owner, not to a documentation pass. This section changed what the room *says* about the posture and nothing about the posture. *(That decision was made the same day, separately and by the owner: the flags come off. §9.6b carries the ruling and ADR-008's seventeenth amendment records it. The split held — the pass that changed the words and the ruling that changed the behaviour are two changes with two arguments, which is what let each be judged on its own.)* The general lesson, in this file's own terms: §9.10 found a mechanism that was complete and unreachable, §9.12 found one that was complete, reachable and unattributed. This one is a claim that was complete, visible, attributed — and **untranslated**. Twelve amendments of adversarial care went into making three words defensible to a reviewer, and none of them asked whether the person the words are *for* could read them. Honesty that only survives an expert audit is a claim made to the wrong audience. ### 9.14 the honest sentence was in the wrong room §9.2 rules that `PhaseWaiting` must never be mistaken for streaming, and it is right. The card that enforced it was reported from a live room in the bluntest terms this project has had yet: > *"'working. this vendor reports no incremental output..' looks ugly as fuck. i get why you > put it there but yuck — you can hide the wiring underneath the floor of our council room."* **Both halves of that are correct, and they are about different things.** The distinction is load-bearing and stays. What does not belong in the body of every waiting turn is the *argument* for it. Read it as a user rather than as its author: "this vendor reports no incremental output" is a sentence about council's own plumbing, in council's own vocabulary, occupying the space someone opened this room to read an answer in. And because two thirds of the seats are `GranFinalOnly`, it was not an occasional card — on an ordinary turn it was most of what was on screen, three columns wide, until the vendors came back. **What carries the distinction now was already there, and had been all along.** The column header names the phase: `waiting` against `streaming`, on every frame, in both glyph sets, above the scroll where it cannot be read past. Beside it the granularity badge says why — `final only`, or a deliberate blank. That word is the claim; the body sentence was never the claim, it was a *paraphrase of the badge*, printed where the badge could already be seen. So the body is one line. Three of them, because they are three different claims and collapsing them would be the failure §9.2 exists to prevent one level down: | when | line | why not the others | |---|---|---| | final-only, nothing yet | `working — the reply arrives whole.` | states what to expect, from a measurement two vendors earned | | granularity never established | `working — nothing has arrived yet.` | must NOT borrow the sentence above — the fifth amendment's rule that an unestablished claim may not wear a measured one's words | | it has acted but not spoken | `working — the steps above are what it has done so far.` | there IS something on screen; pointing at it beats describing the seat | None of them uses a word about incremental output, deltas or granularity. `TestWaitingIsNotStreaming` now asserts that in both directions: the body says what to expect, the frame carries the word `waiting`, a streaming frame does not, and the vendor-internals vocabulary is **absent** — which is the assertion that stops the explanation creeping back in one clause at a time. **The wiring went under the floor, and the floor is the help panel's posture page.** That page already exists (§9.13) and already had the shape for this: a claim on the column, its argument somewhere it can be read properly. What it did not have is any gloss of the granularity word at all — §9.13 gave the sandbox badges a legend and left the badge beside them undefined, which was survivable only because the waiting card was reciting the explanation in the reading area. Taking that out is what turned the gap into a debt. **The gloss sits inside each seat's own block rather than in a room-independent legend, and that is a deliberate departure from how the sandbox words above it are presented.** §9.13's argument for a legend covering badges this room does not show is that a user who has never typed `--write` should learn what `WRITES` means *before* they type it. There is no equivalent here: **nobody chooses a granularity.** It is a property of whichever vendors are installed, so the only granularity words a reader can ever meet are the ones their own room is already displaying — and a sentence beside the word it defines beats making someone match two lists. It goes under the posture rather than beside it, because the two answer different questions about one seat and only one of them has consequences. `TestEveryGranularityIsExplained` walks the type and fails the build for a value that can render on a column with nothing to say what it means — the guard `TestEveryBadgeIsExplained` gives the sandbox levels, for the same reason. `GranUnknown` gets an entry precisely because it prints no word: the blank is the claim, and it is the one case a reader cannot decode by reading the header. **The residual, stated rather than discovered.** Each seat's block grew a line, so at the 24-row floor the last seat's paragraph is cut a little earlier than it was. That is §9.13's own stated trade, one line deeper — the per-seat half is what a taller terminal gets, and nothing above the fold moved. The panel's hard 17-row budget is untouched and `TestHelpFitsTheSmallestRoom` still holds it. The general lesson, in this file's own terms: §9.13 found a claim that was true and untranslated, and translated it. This is the same audit run once more on the *result* — because the translation was correct, and it was put in the wrong room. A sentence can be honest, legible, and still wrong to print, if the place it prints is the place someone came to read something else. Every earlier section here asked whether the room says the truth. This one is the first to ask **how much of the room the truth is allowed to take up.** ### 9.15 getting an answer out of the room Council exists to put several vendors' answers where they can be compared. What it never had is a way to take one *away*. §9.10 already noticed the gap from the other side and refused the obvious fix: mouse support was rejected partly because enabling the wheel claims the left button too, which would cost the native click-drag selection this room's output depends on. That refusal protected a workaround. It did not build a feature. > *"go with your yank key suggestion"* **`y` copies the focused column's reply. `Y` copies the whole turn.** Two keys rather than one with a modifier meaning, because they produce different documents — and `shift` for the wider version of a motion is what this room already does with `g` and `G`. **What `y` takes is the sanitized `Body` the renderer is showing, and the three things it is not are each a rule this file already holds.** Not the raw stream: everything on `State` has been through the redaction and sanitize choke point, and a clipboard is a *worse* place for a credential than a screen because it outlives the room. Not the trace: what a seat did and what it said are different kinds of claim (§4a.1), and that does not stop being true because the destination is a document. Not a neighbour's: it addresses the **focused** column, the same column every scroll key addresses, because a copy key that took from somewhere other than where the eye is would be §9.12's failure with a clipboard attached. It falls back to the newest finished turn when the current one has produced nothing yet — "the last answer" is what a user means by this key, and the notice names the turn the text actually came from rather than the one on screen. **`Y`'s format has one job: be readable a week later.** Seat headings, and the brief at the top, because four answers to a question the file does not contain are unreadable. The brief is the user's own words, which §9.9 already echoes un-redacted on the user's own screen for the same reason — and what rode *along* with it does not go in, which is the same boundary §9.9 draws: a first turn carries the `--brief` file whose content is deliberately kept off `State`, and a rebuttal turn carries other vendors' words. Only seats that took **this** turn are included; a seat that sat out still holds an older reply, and filing that under this turn's heading would be the room inventing a conversation into a document, where it outlives every chance to notice. **The key collision is the interesting part, and it was already resolved.** `y` approves a tool call a vendor is blocked on (§9.8). `key()` routes a pending gate to `gateKey` first and `gateKey` answers `y` itself rather than falling through, so the approve key keeps the letter it has always had and yank does not exist while a vendor is stopped. That was true before this landed; what is new is that it is now **asserted**, because losing that race would mean a keystroke the user believes approved a write quietly copying text instead — and their next move would be to press it again. In compose mode `y` is the letter y, the same rule that keeps `q` the letter q there (§9.10). **The mechanism is OSC 52, and its limit is stated rather than glossed.** Verified by reading the installed module rather than the internet, because v1 answers for this are wrong: `charm.land/bubbletea/v2@v2.0.8`'s `clipboard.go` returns a `Cmd` whose message `tea.go` turns into `ansi.SetSystemClipboard`, emitting `ESC ] 52 ; c ; BEL` **unconditionally** — no capability probe, no terminal query, nothing that can decline in a way this program could observe. | | claim | strength | |---|---|---| | the key produces the command, carrying the right text | asserted by a test that calls the `Cmd` | measured | | the sequence reaches the terminal | bubbletea writes it on the next message pump | read from the module source | | the terminal honours it | Windows Terminal accepts OSC 52 writes in current builds | ~~**INFERRED** from documented behaviour, not run~~ — **FALSIFIED on macOS, 2026-08-10** | That last row cannot be closed from inside this repo: the only observer that could settle it is the terminal, and it sends nothing back. So the notice claims what council **did** — "copied claude code's turn-3 reply" — and never what the machine now holds, and the honest check is a person pressing `y` and then `ctrl+v`. The notice is not decoration for the same reason: with no acknowledgement available, it is the *only* feedback that the key did anything, and a silent copy would be indistinguishable from a terminal that ignored the sequence — the ambiguity §4a.1 forbids everywhere else here. **Amended 2026-08-10: the inference was wrong, and it took a second machine to see it.** On the macOS box `y` reported "copied …" and the clipboard was untouched, in the same build where the key works on the Windows one. Terminal.app does not implement OSC 52 clipboard writes at all; iTerm2 does but ships the permission off. Nothing was broken in council — the gauge was reporting an action it had structurally no way to observe, which is the failure §4a.1 exists to prevent, wearing the costume of a limitation everyone had already agreed to. **A native helper is now tried FIRST wherever one exists** (`clipboard.go`): `pbcopy` on darwin, `wl-copy` then `xclip` on linux, and nothing on Windows, where OSC 52 is measured working and a process spawn per keystroke would buy nothing. The reason is not that the helper is better plumbing — it is that it is **checkable**. `pbcopy`'s exit status is a fact about the clipboard; the escape sequence has none, and never will. OSC 52 stays as the fallback because it is the only mechanism that survives SSH. The two paths never both fire. Sending the sequence as well would put the text on the clipboard twice where both work — harmless — and would also hide the failure of one behind the success of the other, which is not. The test story changed with it, and this is the part worth carrying: the old test asserted the `Cmd` was produced, and **passed for two days while the key did nothing on macOS**. It was asserting the artifact rather than the effect, the mistake this file records four other times. The native path is now round-tripped through the OS (`pbcopy` in, `pbpaste` out) and the fallback is driven by a stub, so the mechanism a machine happens to have no longer decides which half of the feature its suite covers. **An empty yank issues no command at all.** Writing `""` through OSC 52 is the documented way to *clear* a clipboard, so a copy key that found nothing to copy would silently destroy whatever the user had. "Nothing happened" and "your clipboard is now empty" are different outcomes and this key must not spell them the same way. **The help row is a merge, and the merge is the honest shape rather than a saving.** The panel's budget is hard (17 rows, §9.11) and a copy key documented below the fold is a copy key nobody finds — but the reason `y`, `Y` and the gate's `y`/`n` share one line is that they *collide*, and the one place a reader could learn that is the line naming all of them. Splitting them would have spent a row to make the collision harder to see. **Declined: writing the turn to a file as a fallback.** A `~/.telltale/council/last-turn.md` would work in any terminal and needs no escape sequence at all, which makes refusing it worth an argument rather than a sentence. ADR-008's ninth amendment ratified council writing exactly one file and ruled what may be in it: **keys, not content** — session ids and no transcript, because each vendor already stores its own history and anything copied there would be a second copy of a private conversation in a location the user never chose. A file of four vendors' answers in the state directory is precisely that, and it would break the rule in the same release that a mechanism needing **no disk at all** was available. The terminal-support residual is real and is paid in a notice, not in a contract. The general lesson, in this file's own terms: §9.10 measured a fix, refused it for a good reason, and recorded the refusal — which is the process working. What it did not do is ask what the user had actually been trying to *do* when they reached for the mouse. The answer was never "scroll"; it was "take this answer with me", and that want went unnamed for two sections because the request arrived wearing the costume of a mechanism. ### 9.16 `/flow`: the hop holds the authority, and it has to say so out loud A `/flow` chain is `@seat verb [task] [write:]`, arrows between hops, dispatched **one hop at a time to exactly one seat**. Ordinary dispatch can address Claude by default, one or more explicitly mentioned seats, or the whole committee via `@all`; a flow hop has no such choice. It is an instruction to its one named seat, and a chain that fanned out would hand each hop's authority to seats the chain never mentioned. Nothing becomes a flow without the literal `/flow` prefix. A bare `->` in prose stays prose — "compare approach A -> approach B" is a question — because the alternative is ordinary briefs silently acquiring orchestration semantics, write gates and all. **Write authority is declared, never inferred.** The parser as shipped decided a hop could mutate the workspace by checking whether its last token contained `.`, `/` or `\`. Spelled out, that is: a sentence that ended in a period was a write hop; so was a task naming a file it was only meant to *read*; so was a Windows path quoted inside a question. English does not grant permissions and neither does punctuation. **Only `write:` does.** The verb is a label — `publish` confers nothing. The target is checked at **parse** time, which is the load-bearing part rather than a tidiness preference: once parsed, the seat is spawned holding authority and pointed at the path, so parse is the last moment the answer is still free. Refused there — absolute paths in *either* platform's spelling (`filepath.IsAbs` alone does not consider `/etc/shadow` absolute on Windows), `..` in any segment under either separator, an empty target, two targets on one hop, and a `write:` token occupying the verb slot. The last is refused loudly rather than read as a read hop: silently demoting a declared write is the same class of lie as silently promoting a read, and it is the more dangerous one, because the user believes they authorized something and the log agrees. **Posture belongs to the step, and it only ever moves down.** | the hop | the room | what happens | |---|---|---| | no `write:` target | `--write` | **read** posture. The room's authority is not the hop's. | | no `write:` target | read-only | read posture, unchanged | | `write:` | `--write` | gate first (`y`), **then** spawn, at the room's write posture | | `write:` | read-only | **blocked**, and nothing is spawned | The bottom row is where the two dishonest options live and both are refused. Running it read-only would let the seat return, the hop report `returned`, and the chain advance past a publish that never happened — §4a.1's ambiguity with a receipt attached. Upgrading the room would mean a program granting itself authority the person who started it withheld. So there is no `y`/`n` card here at all: a gate implies a keystroke exists that makes the action legal, and none does. The notice says the room is read-only and names **`/write`**, the control that would change it. That last sentence used to name the *flag*, and §9.17 quotes the older version of it as the tell that a control was trapped at launch. It stayed true for exactly as long as a relaunch was the only remedy; `/read` and `/write` made it false and nothing was asserting on it, so the room went on telling a user with a chain half-typed to quit and start over. Corrected, and pinned by `TestAWriteHopIntoAReadRoomNamesTheControlNotARelaunch` — see §9.17's closing note on why a fix that lands everywhere except the sentence that motivated it is the shape to look for. **The gate is before the spawn, and that is now pinned by a test that counts processes.** On a write hop, zero vendor processes exist until `y`; `n` cancels with zero. A gate drawn after the spawn is a notification. **The persistent seat forced a real choice.** Claude's seat is a long-lived process (§9.8) and its posture is argv — fixed at spawn, with nothing in the stream-json envelope able to change it mid-session. This is the same measured constraint as `cwd`, which is why `/cd` respawns rather than redirects. So a hop needing a posture the live process was not launched with can either be sent anyway or trigger a respawn. **It respawns**, on `/cd`'s own `--resume` composition and under the same one-attempt probation, and the column says so. Sending it would have been the silent downgrade this whole section exists to forbid, and the tell would have been invisible in exactly the way that matters: the badge would read READ while the process still held write flags. **Retention was running backwards.** The artifact prune sorted filenames as strings, so `turn-10-*.md` sorted before `turn-2-*.md` and the cap deleted the *newest* ten and kept the oldest — reachable only at turn 10, which is the first moment retention does anything at all. It sorts by the parsed turn number now. A name that does not parse sorts first and is pruned first, because an unrecognised file in that directory is not a receipt this store wrote and must never displace one that is; and nothing there panics, since it reads a real directory that can hold anything. **How these are tested, and why it is worth a paragraph.** Every one of the six security properties is asserted on something observable — the number of processes spawned, the exact argv handed to the spawn, or the chain's state — and never on a helper returning `true`. This repo's recorded failure mode is a test that checks the flag instead of the effect, and it would be perfectly at home here: a flow that computed "read posture" correctly and then spawned a write invocation passes any test that asks the posture function what it thinks. The posture assertion witnesses `@cursor`'s argv rather than `@codex`'s for a measured reason — on Windows, codex's read and write sandbox flags collapse to the same value, so codex's command line cannot testify to a posture on this machine. ### 9.17 a control you need mid-session cannot live in a flag **The rule: state that changes while the room is open is reachable from inside the room. A flag is for what is true at launch and stays true.** A control that only exists as a flag answers, at the one moment you cannot yet have the question, something you will only learn by working — so the remedy for noticing it is to quit the room and lose the conversation. This is not a new principle here. It is the one council has already applied twice, case by case, without ever writing it down. - **The workspace stopped being an invocation input.** `--cd` still exists, but `/cd ` moves the room between turns and every seat follows on its next dispatch (§9.16, `roomcmd.go`). - **Posture stopped being opt-in.** The room writes by default and `--read` is the opt-out, because once the gated seat could raise an approval card, "all the flag still did was make a room you opened to get work done unable to do any until you remembered a word" — and that demotion cites the workspace one as its precedent. Two demotions with the same argument is a rule. `roomcmd.go` even states the scope it was decided under — "the workspace is **the one piece** of room state the P0 demands be movable from inside" — and that claim is what has now failed. It was true when the room could not run long enough for anything else to drift. A room used as a daily driver drifts in several places at once. **The tell is a refusal that names a flag.** §9.16 has one already: a `/flow` write hop into a read-only room is blocked, and "the notice says the room is read-only and names the flag that would change it." That sentence is the defect in miniature — the room knows exactly what you want, knows exactly what would grant it, and can only tell you to quit and start over. Any notice whose remedy is a relaunch is this bug. #### The sweep Every council control, classified. This is a **source read**, not a live run — the claims below are about where a control is reachable from, which argv and `roomcmd.go` settle, not about vendor behaviour, which would need measuring. | control | verdict | why | |---|---|---| | workspace (`--cd` / `/cd`) | **compliant** | the launch flag has an inside-the-room twin; the flag's own help says so | | `--fresh` | **violates** | a conversation fills up *by being used*. The only reset is room-wide and launch-time, so clearing one seat costs the other three their threads | | `--trace` | **retired by `/trace`** | its own doc said it answers "a question that is asked on the days a turn is inexplicably slow" — a day you identify from inside a slow turn. The flag remains for a run you already intend to measure; see below for what the sweep found underneath it | | `--read` (posture) | **retired by `/read` and `/write`** | see the refusal above. Note what is *not* an objection: posture is deliberately never restored from the saved room, because "a posture that can arrive from a file is not one anyone typed." A posture typed into the composer is typed. The flag stays: opening a room that only talks is a real thing to want at launch | | `--auto` | **retired by `a`** | whether the gated seat asks before each tool call is a preference you form partway through a batch, not before it — so the surface is a third key on the card, not a room command. The flag stays: opening a room you already know you will not be watching is a real thing to want at launch | | `--vendor` | **retired by `/seat`** | who is *seated* was launch-only. `-@seat` routes one turn and is explicitly "a different control from an @mention" — routing is not reseating. The flag stays: opening a room with a chosen set is a real thing to want at launch | | `--brief` | **arguable, not filed** | it is defined as first-turn context, so re-briefing is a different feature rather than a missing surface for this one. Left out deliberately; do not fold it in without deciding that question on its own | | `--ascii`, `--no-title` | **legitimately launch-only** | properties of the terminal, not of the room. They do not change while it is open | | `--resume`, `--write` | **vestigial** | accepted and ignored; kept for muscle memory | #### What satisfying the rule costs Not every control can simply be flipped in place. Posture and `cwd` are **argv** — fixed at spawn, with nothing in the stream-json envelope able to change them mid-session — which is why `/cd` respawns the persistent seat rather than redirecting it. That is the pattern, not an obstacle: respawn lazily on the next dispatch, under `--resume` composition and the same one-attempt probation, and let the column say so. A mid-session control may cost a respawn; it may not cost the room. #### The surface, and why it is not a slash command **Ruled: a key on the focused seat.** Focus already ships (`▸` + `Strong`, `hierarchy_test.go`), so the seat is already named without anyone typing its name. The alternative was a room command with the mention grammar (`/clear @codex`), and the argument against it is vocabulary. `roomcmd.go` intercepts only a draft that *is* the command, "so no vocabulary is quietly stolen from the conversation" — but two words are already spoken for, `/cd` and `/flow`, and `/clear` is a word people mean for a vendor, since it is a real Claude Code command. A key takes nothing from the composer. **It must not be automatic.** The obvious version — the room notices a seat is near its ceiling and clears it — reads `context_pct`, which for the codex adapter is declared `Derived`, not `Reported`: Codex ships a denominator, telltale computes the percentage, and the HUD marks it with a leading `~` "rather than passing it off as a vendor figure." ADR-005 settled this class for the fleet — status is advisory, never a gate, and nothing irreversible branches on it. A dropped thread is irreversible. #### What `c` does, and the two things it gets right by construction The first control built to this rule. `c` in view mode arms a confirmation for the focused seat; `y` drops that seat's thread, `n` keeps it, and **any other key cancels**. The flow gate falls through to `viewKey` and this one does not, because they ask different questions: that one blocks a chain already running, so reading the columns is part of deciding, while this one interrupts nothing and the safe reading of a key nobody meant to press is to put the thread back out of reach. It refuses while a turn is in flight — `/cd`'s rule, for `/cd`'s reason — and a seat with nothing to clear is told so rather than handed a card whose `y` does nothing. **The ordering is load-bearing and it fails silently.** `seatProcess` re-arms `resumeIDs` from `m.sessions` whenever it replaces a live process; that is what carries a thread across a `/cd`. So the deletes come *before* the kill. Reversed, the id is handed straight back, the next brief resumes the conversation the user just ended, and every word on screen still says cleared — which is why it is a named test (`TestClearSeatKillsThePersistentProcessAndDoesNotRearmTheThread`) rather than a comment. **The drop is saved immediately, not at the next dispatch.** The room file is what a reattach reads, so a clear held only in memory would be undone by quitting: the user ends a thread and finds it waiting for them, which is the failure the control was built to remove. **`Cleared` is its own field, not `!Restored`.** "This seat never had a thread" and "you ended this seat's thread" reach the same next brief, and collapsing them is zero-vs-absent (§4a.1) applied to a conversation. The marker is a labelled rule in the transcript's own grammar, drawn last because that is when it happened — and **the turns above it stay**. What was cleared is the thread the next brief would have continued, not the record of what was said; blanking the reading surface to report a vendor-side change would be the room destroying the thing it exists to show. It retires in `startTurn`: once the brief is sent the seat has a thread again, and a marker outliving that would describe a break the room has already healed. #### `/trace`, and the thing the sweep found underneath it **The clock was always running.** `runner/clock.go` measures every turn unconditionally — `newClock`, `begin` and `end` sit on the ordinary path — and `--trace` only decided whether `emitTurnClock` had a sink to hand the record to. So every slow turn anyone ever watched *was* measured, at full spawn/wait/stream resolution, and the numbers were dropped on the floor because nobody had predicted that turn before the room opened. That reframes the fix. A `/trace` that merely installed the sink from here on would move the prediction from launch to the previous turn rather than retiring it — you would still be waiting for the slow turn to happen *again*. So the sink is installed for the life of the room and keeps the last **200** records (`maxTraceRing`: `maxHistory`'s 50 turns at four seats, so the trace reaches back exactly as far as the transcript that made you want it). Turning the trace on is opening a *file*, and the first thing that file receives is what the room already held. `/trace ` enables and reports how many held turns it wrote; `/trace off` stops; bare `/trace` reports where it is going, or how many turns are held if it is off. The ring keeps filling while the trace is off, so stopping costs nothing and starting again reaches back over the gap. **One line carries a sixth field, and only a racer's does.** A turn dispatched by `/arena` appends `race=arena/t` after `total=`. The runner cannot infer that from a `Spec`, so the room hands it over (`runner.Spec.Race`); §9.37's dated block of 2026-08-16 carries the seam, the probe behind it, and why an ordinary turn appends nothing rather than a dash. **Three deliberate refusals, each with a reason that is not "consistency":** - **No council-chosen path.** Bare `/trace` reports and never enables. A no-argument form that picked a file would make council write a second file on its own initiative, and the sentence in `README.md` and `CLAUDE.md` — the only mode that writes anything to disk, one file, `room.json` — would become false. A `--trace`/`/trace` path is one the user named. - **It does not refuse mid-turn**, unlike `/cd` and `c`. Those change state the seats are actively using; this opens a file on the room's side and changes nothing any vendor can observe. The turn you cannot explain is usually the one still running, so refusing here would refuse at the only moment that matters — and because the clock emits at `end()`, a trace opened mid-turn still catches that turn. - **A relative path resolves against the ROOM's workspace**, not the process's cwd, matching `/cd`. The room is the frame of reference for everything else typed into it. **The help panel taught the sweep something too.** `/trace` first went in below the panel's hard 17-row budget, on the theory that a diagnostic can be demoted. It cannot: `helpBody` clips at the body height and does not scroll, so a row past the fold is not a cheaper row, it is **no row** — the same failure that put the posture explanation out of reach and split this panel into two pages. All three room controls now share one row inside the budget, and `TestHelpNamesEveryRoomControlAboveTheFold` pins both the fold and the controls, which until now were asserted only by a comment. #### `/read` and `/write`, and the two asymmetries in them The third control built to this rule, and the one the rule was written for: §9.16's refusal of a `/flow` write hop into a read-only room "names the flag that would change it", which is this defect stated as a feature. The room knew what was wanted, knew what would grant it, and could only say *quit and start over*. **The confirmation is asymmetric, on purpose.** `/read` applies at once; `/write` asks `y`/`n`. They are not the same act. Tightening takes authority away from four seats, and the worst case of a stray `/read` is a turn re-run. Loosening hands editing and command authority to every seat in the room — and in an `--auto` room, hands it with nothing left asking. `c` spends a keystroke on its irreversible direction for exactly this reason, and anything that is not `y` cancels here for `clearGateKey`'s reason: this gate interrupts nothing, so a key nobody meant to press must not be able to arm the room. **The card names which write you are getting.** Gated write and `--auto` write reach the same badge-bearing posture by different routes and only one of them asks first, so the confirmation says which — a card promising "claude asks before each change" in an `--auto` room is a promise that room cannot keep. §4a.1 applied to a prompt rather than to a gauge. **Neither direction is offered mid-turn**, which is `/cd`'s refusal rather than a house style. Posture is argv, fixed at spawn, so seats already running hold the flags they were launched with whatever the room now says. Landing the flip under them would put a read-only badge over a live process still holding write flags — the disagreement between claim and process that the per-step posture rule exists to forbid. Nothing is killed: `seatProcess` already respawns on a posture mismatch under the same measured `--resume` composition it uses for `/cd`, so a `/read` that is `/write`d back before anyone dispatches costs nothing at all. **The badges are rebuilt, not just the flag.** `Sandbox` is computed once in `stateWith` from `opts.Write`, so a posture that moved without `applyPosture`'s loop would leave four columns advertising authority the room had just taken away — a displayed value no longer coming from what is true. `TestPostureFlipRebuildsEveryBadge` asserts on the rendered badge rather than the field. The `WRITES` and `gated` glosses were updated in the same change for the same reason: both credited `--write` as the only way to reach them, which would send a reader looking for a relaunch out of the glossary that explains the thing. **Only the bare word is a command**, unlike `/cd` and `/trace`. Those take an argument, so `/cd ` and `/trace ` are unmistakable; these take none, and both are words a person addresses a room with. `/write a test for this` and `/read the design doc first` are ordinary briefs, and intercepting them would swallow a turn and run a setting instead — worse than stealing a word, because the user watches their brief vanish rather than being told it was a command. Posture is still never restored from the saved room. `TestReattachDoesNotRestoreWritePosture` is unchanged and still holds: "a posture that can arrive from a file is not one anyone typed" — and a posture typed into the composer is typed. #### `/seat`, and the control it is deliberately NOT `--vendor`'s twin, taking the same argument for the same reason `/cd` takes `--cd`'s: `/seat claude,codex` and `--vendor claude,codex` are one grammar, read through the same alias table `@mentions` use. Two tables would let `/seat agy` work and `@agy` not. **What it does not do is the design.** An unseated seat keeps its thread, keeps its process, and keeps every id that would resume it. Only two things change: it is not drawn, and it is not dispatched to. Killing the process to reclaim it was considered and rejected on the ruling that a returning seat picks up its own thread where it left off: - **The thread is the thing being protected.** A seat with a live process and no reported session id yet holds its whole conversation *in that process* (§9.8). Killing it there destroys a thread `seatHasThread` calls real — silently, on a command nobody reads as destructive. Dropping a thread is `c`'s job, and `c` asks first. - **Nothing is being spent.** An unseated seat is never dispatched to, so an idle process costs a process and no quota. Trading a guaranteed-correct return for a resource nobody is short of is the wrong trade. So reversibility is by construction rather than by a resume that could fail: `/seat all` puts everyone back mid-conversation with nothing to go wrong. What it buys is what the fold-out already buys an uninstalled seat — the **width** goes to the seats answering. **Sitting out is a different control and already exists.** A seat nobody addresses does not answer and is not billed; §9.19 renders a long absence as one line rather than ten. `/seat` is for the seat you want off the *screen*, not merely quiet — which is why it was worth building even though the quota problem it looks like it solves was already solved by the default route. **It warns when it unseats the default route.** Silence goes to claude, so a room without claude answers nothing until every brief is `@mentioned`. Dispatch already refuses a zero-seat route per turn; saying it once at `/seat` time is the difference between a rule learned now and one discovered on the next enter. #### `a`, and the field that was nearly a landmine The last control on the sweep, and the only one whose surface is **not** a room command. The preference forms while a card is on screen — you decide to stop being asked at the eleventh identical card, not at a shell prompt — so `a` sits beside `y` and `n`, where the question is. It approves the card in front of you as well as the ones after it: an `a` that turned asking off and left the current request pending would answer the general question and not the one on screen. **The queue is drained, not discarded.** A pending gate is a vendor *stopped* mid-call, and `queueGate`'s own rule is that nothing may quietly drop a request — a dropped queue leaves columns waiting forever with no card left to explain why. So every card behind the current one is approved, and the notice says how many. **It takes effect on the REQUEST, not on the next spawn.** A process already running keeps the gate flags it was launched with, so it goes on sending requests after `a` is pressed. If those queued, "stop asking" would keep asking until the turn ended — the promise broken at the moment it was made. `queueGate` reads the room's state per request and answers immediately; the respawn that drops the flags happens later, through `seatPosture`, on the next dispatch. **It is not a one-way door.** `a` alone in view mode turns asking back on, and the footer carries a permanent `a not asking` cell whenever the gate is off. Without that cell the room would sit ungated with the way back documented nowhere on screen — the §9.17 defect rebuilt one key later. The cell is not sheddable, for `t grid`'s reason: shedding it would drop the way out of a state rather than a convenience. **And the field is stored negated, which is the finding worth keeping.** The obvious shape is `Asking bool` — whose zero value is *does not ask*. Every `State` built as a literal would have been a silently ungated room, and the reason that was caught at all is that five existing gate tests build their State by hand and went green while asserting nothing. A safety property whose default is off is the wrong way round however carefully the constructor sets it. `GateOff bool` read through `Asking()` makes the zero value the guarded room and turning the gate off an act. `TestTheZeroStateAsks` pins it. #### Nothing left on this list Every control the §9.17 sweep found in violation now has an in-room surface: `--fresh`→`c`, `--trace`→`/trace`, `--read`→`/read`/`/write`, `--vendor`→`/seat`, `--auto`→`a`. Every flag stays, because each names a room you may genuinely want at the door; none of them is any longer the only way to get one. `--brief` remains deliberately unfiled — it is first-turn context by definition, so re-briefing is a separate feature rather than a missing surface for this one, and folding it in without deciding that question on its own is still the thing not to do. #### The sweep built the controls and left two sentences describing the old room Every control landed and two strings went on describing council as it was before them. Neither was cosmetic, and the shape they share is worth more than either fix. **The refusal that motivated the whole sweep was the last thing to be fixed by it.** §9.16's `/flow` write-hop block still said the room "was opened with `--read` — reopen it without that flag", which is the §9.17 defect quoted verbatim *as the specification of the bug* and then left running. The remedy is `/write`, it was two PRs old, and the notice sent a user with a half-typed chain out of the room to fetch a flag they no longer needed. **A refusal is the surface least likely to be re-read after the thing it refuses becomes possible**, because it is written once, by the person who knows it is correct, and then only ever seen by someone who is already stuck. It also now reports the *posture* rather than the launch argv, since `/read` reaches this state too and "opened with `--read`" would be false as well as useless. **And `/write`'s confirmation card was reading the flag instead of the room, which is the one that could cost something.** The card exists to say which write you are getting — §4a.1 applied to a prompt — and it chose its wording from `m.opts.Auto`. That field is only the *seed*: `stateWith` copies it into `GateOff` at launch and `a` has moved it independently ever since. So a room opened gated, told to stop asking, then `/read` → `/write`, offered a card promising "claude asks before each change" with nothing left to ask. The user reads a promise of a checkpoint and gets none — the failure this card was built to prevent, arriving through the control that was supposed to prevent it. The `--auto` wording went with the field, because the flag is no longer the only route into an ungated room and naming it asserts a cause that may not be there. **The rule, stated once so the next control inherits it: a flag that gains an in-room twin stops being the answer to "what is the room doing" and becomes only the seed.** `dispatch.go` already says this for the request path — "`m.st.Asking`, not `m.opts.Auto`: the flag only SEEDS this at launch" — and the two misses were both places that had not heard. Every `opts.*` read on a demoted control is now either a launch-time decision (`wantsGateHook`, the `savedPosture` record) or a bug, and the way to tell them apart is to ask whether an in-room control can move the state underneath it. Both fixes are pinned by tests rather than comments for the same reason: in every room nobody typed a control into, the flag and the state agree, so the fixtures cannot tell them apart and neither could review. ### 9.18 a strip said four fifths of a name it could have said whole in two letters Since the default route stopped being everyone, the ordinary turn narrows the frame to one seat and leaves the rest at `stripColumn` — fourteen cells. Every layout rule in §9.11 was written for a column three times that, and at fourteen the room did the opposite of what §9.11 ruled in both halves of the chrome at once. `Antigravity` rendered `Anti…`. The badge row rendered `ro:tools to` and `gated fina`, and the overflow marker rendered `↑ 12 more abov`. The ruling those violate is §9.11's own: **a clipped seat name is still recognisable and a clipped state word is not**, so identity yields first. A clipped state word that is also the prefix of another word in the same vocabulary is worse than damage — it reads as a different claim. `fina` is not a broken `final only`; it is a thing this room does not say. So at strip width the room **sheds whole words** rather than cutting them, in a fixed order that is a pure function of the width — which is what lets the frame sweep pin the whole ladder instead of a golden per state: - **Identity collapses to two letters.** `CC ✓ done`, `CX ○ idle`, `AG ⠋ streaming`. The tags are the HUD's own, character for character, because a reader who learned `CX` is Codex from the HUD's grid must not meet a second abbreviation in the room. They are *copied*, not imported: the seam between the two surfaces is the normalized session model and `internal/theme`'s numbers and nothing else, and a test asserts the strings by literal so the copy cannot drift in silence. - **The clock goes, then the focus mark, then — for `unavailable` alone — the tag itself.** `8s` is the meta on that line and `turnRule` already ranks a label above the numbers that belong to it; every finished turn still carries its elapsed on its own separator. The arithmetic behind the rest: nine cells of `streaming` plus its mark leaves exactly three, which is a two-letter tag and the space after it, and `unavailable` at eleven leaves room for a mark or a tag but not both. - **The badge row keeps the posture word and drops the cost and the granularity.** §9.2 is emphatic that a claim you cannot see is not a claim, so the safety word is the last thing on that row to go; the cost is a number the transcript records on every turn separator, and the granularity word exists to keep `waiting` from reading as a slow `streaming` — both of which are now the only thing on the header one row above. A badge too long for a strip would drop rather than clip, and stays readable at full length on the `?` postures page. - **The overflow marker sheds `more`, then `above` / `below`.** The count is never traded: how much is hidden outranks which way to press, which outranks the filler between them. **The focus mark is the one deliberate loss, and §9.12 is why it is affordable.** Two cells of `▸ ` at fourteen is the difference between a tag and no tag for every nine-letter phase word. §9.12 had already found that the glyph was the *weakest* part of that signal — "one `▸` in a frame carrying four columns of prose" — and moved the load-bearing half onto weight and onto the overflow marker's own words, `↑↓ scroll` against `tab to focus`. Both cost no cells and both survive here. A strip is by construction the seat this turn was **not** addressed to, so spending a seventh of its width marking it, at the price of its identity, inverts the priority §9.11 set. **What was declined.** Keeping the two-cell indent on a strip so the focused and unfocused forms line up: it is chrome that exists to align a *name*, and at strip width there is no name — the header starts at column zero and the badges start under it, so the strip reads as one flush-left block rather than as a column with its margins still on. Shortening a phase word to fit (`cancel` for `cancelled`, `stream` for `streaming`): a different word is a different claim, and the vocabulary is shared with the help panel and the transcript. Giving the strip a narrower vocabulary of its own — a second alphabet is exactly what §9.11's phase marks were built to avoid. ### 9.19 sitting a turn out cost a line a turn, and wore the wrong mark doing it Since the default route became one seat (#99), three columns sit out every ordinary turn. The room said so, correctly, once per turn — and a quiet seat's transcript became a column of identical warnings with the answer it actually gave scrolled off the top: > `⚠ not addressed in turn 2` / `⚠ not addressed in turn 3` / `⚠ not addressed in turn 4` / … Two things are wrong there and they are separate. One is the arithmetic. The other is the mark. **Consecutive skips coalesce, at render time only.** A run of turns this seat was not part of is one muted line — `not addressed in turns 2–7`, singular for a run of one. The run is the fact; the turns inside it are not separately interesting, and a reader who wants one has the numbers. The **data model is untouched**: nothing is written down for a turn a seat did not take, which is §9.9's rule and the reason a transcript skips from 3 to 5 in the first place. The runs are *derived* from the gaps between the turns that ARE recorded, so `[` and `]` still hop between real turns (§9.20) and no record says anything it did not say before. A run broken by a turn the seat took starts a new line in place, so the transcript still reads in order, and the LIVE turn's skip keeps a line of its own — the run above it is history, that one is the turn the user is deciding whether to act on. A run is never claimed **before the oldest record**. History is capped at fifty and drops the oldest first, so a column whose early turns were evicted would otherwise report "not addressed in turns 1–29" about turns it may well have answered. Inventing an absence is the same error as §9.9's inventing a conversation, run the other way. Underneath the rendering bug was a data one, and it is the reason `Column.Skipped` exists at all: the note is written on the LIVE column, and the live column is what `startTurn` files into history. A seat that answered turn 1 and then sat out through 7 filed turn 1's record wearing `not addressed in turn 7` — a turn that succeeded, with someone else's absence stapled under it. A skip is not a fact about any turn this column recorded, so it does not travel with one. **The mark is demoted to `○`.** `⚠` opens a note because a note reports something that did not complete normally — a cancellation, a seat that is not there. Sitting a turn out is neither, and it was a fair mark only while a narrow route was the exception. Drawn on the ordinary case it is a warning the eye learns to skip, which is the same argument `ActDenied` makes for `SevWarn` over `SevCrit` and the reattach card makes for no mark at all. `○` is what this room already spends on *nothing has been asked of this seat*, which is exactly what a skipped turn is, said about one turn instead of a session. It survives `--ascii` as `.` against the warning's `!`, so the demotion is legible with colour switched off — and the word carries it first either way. **An idle strip says where it left off.** At fourteen cells (§9.18) a backgrounded seat had a header, a posture word and a run of skips, and the one thing a reader wants from it is which turn it last took: `last: turn 8 ✓`, above the coalesced line. Every part of it is measured — the number is the turn this column recorded, the mark is that turn's own phase — and a seat with nothing behind it renders nothing rather than a placeholder, because absent is absent (§4a.1) and this room does not draw `last: —`. Strip width only: a wide column already answers the question with the turn separators themselves, and repeating it there would be the room being loudest where it has the least to add. **What was declined.** Recording a `TurnRecord` per skipped turn so the coalescing could read one list: that is the room writing down a conversation that did not happen, and it would put the skips in `[`/`]`'s path. Dropping the live skip into the coalesced run to save a row: the run is history and that line is now, and a reader deciding whether to re-address a seat should not have to read a range to find out. And giving the skip a mark of its own — `○` already means this, and a second glyph for one meaning is the collision `glyphs.go` argues against. ### 9.20 the transcript is turn-wise and the only way through it was line-wise §9.9 gave every column a real conversation and §9.10 and §9.12 made it reachable and attributed. What none of the three changed is the *unit*. The scrollback moves a line at a time, a page at a time, or all the way to either end — and the thing being scrolled through is a list of turns, each one a labelled rule, a brief, and however much prose a vendor felt like producing. So the room could tell you, honestly and precisely: > `↑ 509 more above` and nobody has ever counted lines. The number is measured, it is correct, and the only question a reader actually has — *how far back is what I asked?* — is one it cannot answer. `g` goes to the beginning and `G` goes to the end, which are the two positions in a transcript that need no help finding. **`[` and `]` walk the focused column one turn at a time.** They land the turn's separator on the viewport's top row, which is the position that makes the brief and the answer to it readable in one screen, and they take their offsets from `columnLines` — the *same* pass that produced the lines — rather than recomputing where a turn starts from `History`. A second derivation of "how tall is this turn at this width" would agree with the first until the day a card grows a row, and would then disagree silently, since both answers would still be plausible line numbers. **Backwards is the audio player's rule, and it is the one people already have in their hands.** `[` from the middle of a turn lands on *that* turn's head; only a second press reaches the one before it. That falls out of the definition rather than being a special case — "the last head strictly above where we are" produces both — and it means the key answers "start this again" and "go back one" with the same press, in the order a reader wants them. **The two ends are deliberately not symmetric.** `[` at the first turn does nothing: there is no turn 0, and a wrap would make a key pressed one time too many jump an entire conversation. `]` past the last turn restores the tail and `Follow`, because what comes after the last turn is the live output — that is `G`'s answer to the same question, not a second one. Every landing goes through `applyScroll`, so `Follow` drops exactly as it does for `↑`; a column pinned to the tail while displaying turn 3 would be lying about which of the two it is doing. **In compose they are the characters `[` and `]`.** No rule was added for that: §9.10 replaced the composer's list of exceptions with a test — a key that carries text *is* text — and brackets carry text. This is the same contract that keeps `q` the letter q there, and it is asserted rather than assumed, because a bracket that scrolled instead of typing would corrupt a draft in a way the user would only find after pressing enter. #### The marker states the coordinate, and §9.12's rules decide what it costs The overflow marker is where the count lives, so it is where the coordinate belongs: `↑ 25 more above │ turn 3 │ ↑↓ scroll │ f expand`. Three constraints from §9.12 and §9.10 bound the whole design, and each one closed a question: - **The count is never traded away**, in any form, at any width. It was §9.10's rule about the key hint and it is unchanged: how much is hidden outranks both how to reach it and what it is. - **The coordinate sheds FIRST** — below even `f expand`. It says *where you are* while the hints say *what you can do about it*, and a marker that dropped a key to keep a coordinate would be §9.10's trade run backwards. Concretely it rides only on the widest hint form, so at the three-up tier's 37 cells the keys win and the coordinate is simply absent; `f`, the reading tier, is where it appears. That is the graceful degradation, not a gap in it. - **A marker states the key for THIS column and never a neighbour's** — §9.12's rule, applied to a fact rather than a key. An unfocused column keeps `tab to focus`, the one thing a reader looking at it can act on, and gets no coordinate at all: putting the question in front of the answer is how §9.12's bug worked in the first place. **Which turn it names is the part that could have lied.** The choice was between the topmost hidden separator and *the turn the line immediately outside the fold belongs to*, and only the second is honest when a turn is half on screen: a long reply running off the top is still the turn you are reading, while the topmost hidden separator can be several screens further back and answers a question nobody asked. So `turnAt` takes "the last turn that started at or before this line", the two markers on one column name two *different* turns, and a column with no turns at all — an unavailable card, a seat never asked anything — prints nothing rather than `turn 0`, because a coordinate the room does not have is omitted and never invented (§4a.1). #### The footer learned to shed a cell instead of losing its way out `[ ] turn` joins the view mode line immediately after the arrows, offered unconditionally for §9.12's reason — the promise is about what the mode can do, not about how many turns a vendor happens to have taken, and a footer cell that appeared at the first dispatch is chrome moving while output arrives. That exposed something the line had been getting away with. At the tabbed tier the six hints fit **exactly**, and `statusLine`'s only answer to running out of width is to truncate — from the right, which is where `? help` and `q quit` live. A motion key bought with the panel's documented way out and the room's only quit key is precisely the trade §9.11's footer pass existed to refuse. So a hint may now be marked *sheddable*: when the line does not fit, the sheddable cells are dropped whole, newest-first, before the ellipsis is allowed to choose. Exactly one hint carries the mark, and it is the one this section added — this is a rule about which cell goes, not a licence to hide keys. The help panel took it inside the hard 17-row budget by merging onto the row that already holds the other jumps, the way §9.15 merged `y`/`Y` onto the gate's row: `g / G first turn or newest; [ ] step one turn at a time`. "jump to the" paid for it — the line above already says `scroll`, so the verb was never carrying anything. **What was declined.** A turn coordinate on unfocused columns, which the width would have paid for out of `tab to focus` (above). Numbering the hop in the notice line — "turn 3 of 7" is a progress bar for a conversation, and the marker already says how much is left in the unit the scroll keys use. And a `[`/`]` that moved *focus* between columns when a column has one turn: two motions on one key, resolved by content, is the kind of binding that is only ever right for the person who wrote it. ### 9.21 the room knew what the turn would cost and did not say #99 restored the cheap default: silence goes to Claude alone, and the committee is convened by typing `@all` or naming the seats. That settled *which* route is expensive and made every expensive route explicit — and it left the footer stating the route in the same words whether it reaches one vendor or four. `→ everyone` is accurate, and how much `everyone` is depends on what is installed and on what `--vendor` left in the room, which is exactly the part a user cannot read off the word. **The routing cell states the bill when the draft would reach more than one seat.** `→ everyone (3 seats)`, `→ everyone but codex (2 seats)`. The room already computes this number — `dispatch` refuses a turn that reaches nobody by counting it — and the moment it is worth knowing is the moment before `enter`, on the cell that is already answering the same question. - **One seat states no count.** `→ claude` names every seat it reaches in its own text; a cell that restates its neighbour is how this footer became the wall §9.11 had to take apart. From two upward the route names a *set*, and the size of a set is not in the word. - **It counts seated ∩ addressed**, through the same `State.SeatsIn` the dispatch gate now calls. A route may name a vendor that is not installed or that `--vendor` left out; that seat is never spawned, so billing for it would quote a price for a turn that does not happen. `Model.seatedIn` became one line delegating to it rather than a second copy — a bill derived from different arithmetic than the dispatch is a bill for a different turn. - **A refused route prices nothing.** `mixed @ and -@` addresses nobody, and the one thing that cell owes a reader mid-typing is what is wrong with the line they are still holding. - **No colour, no cell, no new glyph.** The count is the *label* half of a hint and the route is the *key* half, which is the figure/ground split every other item on this line already makes (§9.11) — so the seat names keep their intensity and the number recedes to chrome for free. It is parenthesised because that is this room's existing grammar for a qualifier on the thing in front of it (`(+2 queued)`, `(turn 1 is blind)`), and because weight is invisible under `NO_COLOR`, where `→ codex, agy 2 seats` runs the price into the list it is pricing. The rebuttal tag moved to its own cell so the count could sit against the route it prices; it kept its intensity by keeping the key half of a hint. #### The header carries the live turn's route Once `enter` is pressed the composer clears and its routing cell resets to the *next* draft's default, while the columns take anything from seconds to minutes. For that whole window the room has nowhere at all that says where this turn went — and each column's transcript does not record participation until it lands. So the header's turn cell carries it while it is live: `turn 10 → everyone`, `turn 10 → codex, agy`, reverting to plain `turn 10` when the last column finishes. **The route becomes history at that instant**, and the transcript is where history goes; a header still naming it would be describing the past in the one cell that describes the present. `State.TurnRoute` is a **pointer**, and that is §4a.1's zero-vs-absent rule rather than a style choice: `Route{}` is a real and extremely common route — it is what `@all` parses to — so a value field could not tell "this turn went to everyone" from "no turn is running". The same distinction `Column.CostUSD` draws with the same mechanism. It is set where the turn actually starts rather than beside `FrameOwners`, because everything above that line can still refuse the dispatch and a route on the header of a turn that never began would report a spend that never happened. The two have opposite lifetimes on purpose: the geometry outlives the turn so nothing reflows under a reader (§9.11), the route is retired with it. **It prints the route's own `label()`**, never a second vocabulary — what the header shows is what would have to be typed to produce it — and the arrow is the literal one the composer's cell uses rather than a `Glyphs` entry, so one fact cannot drift into two spellings. **A `/flow` hop states no route at all.** A hop is dispatched to exactly one named seat (§9.16) and the cell immediately to its right already says which, so the route would be the header saying the same thing twice — and the arrow, appended after the hop, would read as pointing at it. This is the same rule as the shedding below rather than an exception to it. #### Shedding order: a fact with a home elsewhere yields to facts that have none The header already elides the workspace path from the left, and the new cell had to be ranked against it. **The route sheds first** — before the path, before `3/4 seated`, before `briefed`. The route is on screen in the composer a keystroke earlier and in the transcript a moment later; the workspace is nowhere else at all, and it is the one fact here that changes *what the agents can see*, which is why it has been on screen at all times since this header was written. So the route is added only when it costs nothing that was already there: the path keeps its cells if it had them, and where there was no room for a path either way, the counts keep their gap. **What was declined.** A dollar figure beside the seat count: cost is reported per seat per turn where a vendor reports it, and multiplying a seat count by anything would be council deriving a number and presenting it as read — the top item on this repo's rejected list (§4a.1). Billing the *route's* vendors rather than the seated ones, which would have been one line shorter and would have priced seats that are never spawned. And a count on the one-seat case, which is a number whose only reading is "yes, one". #### Amendment, 2026-08-17: the room shipped the quota relay and never read it The cell above prices a turn in SEATS, which is the half of the bill council could count. The other half was already on disk and nobody was looking at it. `telltale statusline` has relayed every quota window it renders to `~/.telltale/quota/.json` since 2026-08-07 (§7.15) and the HUD has read it ever since — while the one surface that actually *spends* those windows, four or five accounts at a time on one keystroke, said nothing about any of them. The room could tell you a turn would reach three seats and not that one of them had nothing left to answer with. **Council reads the relay. It writes nothing.** `internal/council/quota.go` reads `quotacache` at room open and again when a turn tears down, as a `tea.Cmd` returning a `quotaMsg` — never inside `Render`, which stays pure over `State` (`TestRenderIsPure`), and never on the tick, because the file only changes when the user's own statusline fires. The read/write boundary is untouched: council's one sanctioned write is still `room.json`, and a second one would have to be argued from scratch rather than inherited from a read. **§7.17's declined "per-row quota" does not bind here, and the reason is arithmetic.** That ruling refuses a quota cell on a HUD *row* because a row is one session: five Claude sessions would each draw the same 42% and read as five separate budgets, asserting a per-session limit that does not exist (§7.1 rule 6). A council **seat** is not a row. There is exactly one seat per vendor in a room, so a seat reading is an account reading printed once against the account it describes. Nothing here is ever drawn per session, per turn, or twice for one vendor. ##### What renders on a seat A text reading on the badge row: the window's own label, the vendor's own percentage, the reset countdown while it fits, and the reading's age. ``` ro:tools tokens 5h 12% resets 1h04m 7d 6% resets 5d00h 2h ago ro:tools tokens 5h 12% 7d 6% 2h ago ro:tools tokens 5h 100% ⚠ stale 19h ago ``` The word `resets` rather than the HUD's `↻`: council's `Glyphs` has no slot for that mark, and minting one would grow a set §9.26 keeps deliberately small. The cells between parts are this room's own two spaces (`historyMeta`, the badge row itself), not the HUD's middle dot, for the same reason — one surface, one joiner, and nothing new to give an ASCII partner. - **No gauge track, and that is a ruling rather than a shortcut.** The HUD spends a bar on this because it has a header line to spend it on. Council would need a fill colour to draw one, which re-opens both the closed `isDark` question and `style.go`'s standing rule that council adds no hues of its own — a large purchase for a signal the percentage beside it already carries. The label and the digits are words and numbers, so `--ascii` and `NO_COLOR` lose nothing at all. - **The reading takes the space the row has LEFT.** A new claim does not evict an older one: the posture badge is the safety claim §9.2 refuses to let yield, the granularity word is what keeps `waiting` from reading as a slow `streaming`, and the cost is the one figure on this line the transcript also records. All three keep their cells; the reading takes what remains, sheds its countdowns, and drops **whole** rather than clipping (`stripBadges`' ruling — a clipped percentage is a different number). At the reference 120 columns a three-seat grid gives a column thirty-eight cells and only the roomiest badge row has space for a figure; the footer cell below is what a narrow room keeps instead. - **The age is the HUD's, verbatim.** `2h ago` from five minutes, escalating past five hours to `⚠ stale 19h ago` — same threshold, same word, same order (word, then glyph, then hue). `quotaAgeShown`, `quotaAgeWarn` and `quotaAgeWord` are **copied** rather than imported, on `vendorTag`'s precedent: internal/council and internal/hud share the normalized session model and internal/theme's numbers and nothing else, and `TestSeatQuotaAgeMatchesTheHUDsThresholds` pins all three by literal so the copy cannot drift in silence. A reader who learned `stale 19h ago` on the statusline must not meet a second spelling of it in the room. - **A window relayed for its reset time alone renders nothing.** It says when something will change and not what is left, which is the only question this line exists to answer. ##### What renders on the route cell One seat's name and one of that seat's own readings, when the reading says the turn may not land the way the reader expects: `⚠ claude 5h 100%`, `⚠ agy stale 19h ago`. It sits against the route it qualifies, in compose mode only — the header's live-turn route is already too late to act on. - **It computes nothing.** The refusal above declined a dollar figure beside the seat count because multiplying a count by anything is council deriving a number and presenting it as read. The same refusal binds here and is wider: no total across seats, no average, no count of how many seats are affected, no percentage arithmetic of any kind. Every character after the vendor id is copied off one window. - **Seated ∩ addressed**, the same intersection `State.SeatsIn` counts and `dispatch` loops over. Warning about a seat this turn will not reach is a warning about a turn that does not happen. **The count cell is untouched** — it keeps its present grammar and its present arithmetic. - **A hundred per cent is the only threshold, and it is the vendor's.** Ninety, or "nearly full", would be council picking a severity boundary no vendor published — the same class of guess as filling a `CapNone` field with a plausible value (§4a.1). - **Staleness outranks fullness for one seat**, which is `quotaAgeWarn`'s own argument: a reading past it may no longer be assumed to describe now, so a stale 100% is not evidence the window is full, it is evidence the room does not know. Reporting it as `100%` would be the nineteen-hour incident reproduced in a new room. - **The word "spent" is refused.** §7.17 owns it for token counts, and quota and spend are the two claims that view exists to keep apart. The reading needs no verb: `5h 100%` says it. ##### Zero, absent, and the three vendors that are absent forever `Column.Quota` is a **pointer**, the same mechanism `Column.CostUSD` and `State.TurnRoute` use and for the same reason. A vendor at 0% of its window has been measured and draws `5h 0%`. A vendor with no relayed reading draws **nothing at all** — no dash, no placeholder, and not one cell of width, which `seat-quota-absent.txt` pins by rendering the two states side by side in one frame. Collapsing them is the more dangerous direction of the zero-vs-absent bug here: an unrelayed seat would read as a fresh account and invite a dispatch the room has no evidence will land. Cursor and grok are in the absent class permanently, and so is Gemini: none of them writes quota to disk in any form a passive reader can see (§7.17's structurally-absent row), so no relay entry can ever exist for them. Codex is absent from this surface for a different reason — its quota lives in its own store, and this room reads the relay and nothing else. The room does not distinguish the three, and it does not have to: on this surface they are one fact, *this room has no reading*, and each vendor's own sentence explaining why is §7.17's job on the surface built to hold a paragraph per vendor. **Expiry is the read's, not the room's.** `quotacache` drops a window whose reset has passed and any entry over 24h old before council ever sees it (§7.15), and a read that no longer speaks for a vendor **clears** that seat. A room that kept its previous reading would be displaying a percentage §7.15 calls not stale but FALSE. ##### Limitations, recorded rather than left to be found - **Exactly one seat is named on the route cell, and a second is not counted.** Column order decides. Ranking two seats would mean ranking a stale reading against a full window, and there is no measurement behind such an order; a count would be the aggregate this cell may not compute. What carries the rest is each seat's own badge row — and at a width where the badge row shed its figure, a second affected seat is not on screen. - **The reading is as old as the last statusline render in that vendor.** Council writes no relay of its own, so the post-turn read sees a turn's cost only after that vendor's statusline fires again. This is exactly what the age suffix exists to say, and it is why the age never sheds. - **A room open past `quotaAgeWarn` with no statusline activity escalates every reading it holds.** That is correct rather than noisy — the readings really have outlived the fleet's shortest window — but it means a long idle room ends up with a warning on its footer that only the vendor's own statusline can clear. ### 9.22 four answers to one question, and no way to read them as one Council exists to put several vendors' answers side by side. Everything from §9.9 onward built the surface that does it — a real transcript, per seat, scrollable, attributed, navigable a turn at a time — and every one of those sections improved a **column**. The room therefore had a comparison surface with no way to read a comparison. To see what four seats made of one brief you scrolled Claude to turn 10, remembered it, tabbed, scrolled Codex to turn 10, remembered that, and tabbed again; §9.20's `[` and `]` made each of those one keystroke and did not change what the exercise was. **The document already existed.** §9.15's `Y` assembles exactly this: the brief once at the top, then every seat that took *this* turn, labelled, in seating order — and it was ruled, argued and tested a release ago. What it could be read in was a clipboard. So this section adds no content model at all; it renders the one `Y` already had, and the two now come from the same call (`turnEntries`), which is the point rather than a tidiness. A page and a paste that disagreed about who was in a turn would be two honest-looking documents with nothing on screen to say which was the room's answer. **`t` swaps the body between the by-seat grid and one turn's page.** One key and a toggle, because the two are one transcript read two ways rather than two places — and it opens on the turn the grid was already following, so the projection changes and the subject does not. #### What the page is, line by line, and why none of it is new - **The turn's own rule**, carrying the number, where the turn went, and how long it took: `turn 10 ──────── → claude, codex 41s`. Same `labelRule` grammar, same shedding, meta before number, as every separator since §9.11. - **The brief once**, under the composer's own `›` at full weight. Four copies is what a *grid* has to do — each seat's prompt is a fact about that seat (§9.9), since a turn can reach two seats and not a third — and it is precisely what a page must not. - **Each participating seat under its own labelled rule**, name at weight, then its activity trace, then what it said, with §9.11's boundary strengths unchanged. The only thing that differs from a column is what the strongest boundary is *about*: a turn there, a seat here. That is what swapping the projection means. - **A seat that sat the turn out does not appear.** §9.15's rule, for §9.15's reason: it still holds an older reply, and filing that under this turn's heading would be the room inventing a conversation — on the surface built to compare them, where it would be believed. - **Failed and cancelled seats keep their note cards.** A turn's page shows what actually happened; the two turns anyone scrolls back for are the ones that went wrong. **The route is read off participation, not off `State.TurnRoute`.** §9.21 retires the live route the instant the last column lands, because the header describes the present. What outlives it is the measurement — a `TurnRecord` exists for exactly the seats the brief reached — so the page states who took the turn, through `Route.label()` so what is displayed is still what would have to be typed to reproduce it. **The clock is the longest seat's own elapsed**, because a turn is over when its slowest seat lands. A sum would be the wall time of a room that dispatched serially and a mean is a duration no seat ever took; both would be council deriving a number and printing it as read, which is the top item on §4a.1's rejected list, and being in seconds does not exempt them. A turn still running carries no turn-level clock at all — how long it took is not a fact yet — while each seat's own rule carries its running one, from `State.Now`. #### The two rules that shaped it more than the layout did **Gate precedence, in both projections.** A pending approval renders on the page as *chrome* — above the scroll, like `columnChrome` — and the argument is stronger here than in the grid: a vendor is stopped, the live page follows its own tail, and a card inside the body would be pushed off screen by the output of the very call it is asking about. It also **names the seat**, which the grid's card never had to: there the card's position *is* the seat, and one page has no position left to carry it. And `y`/`n` still answer the gate before they mean anything else, because `key()` routes to `gateKey` first in either view. That was already true; §9.15 made it asserted, and it is asserted again here — a keystroke the user believes approved a write must never quietly copy text instead, since their next move is to press it again. **§7.1 rule 4 decided what the footer says.** A turn arriving while an older page is open **never moves the view**: content jumping out from under a reader because a vendor finished is the thing the bottom-anchor and the frozen-geometry rules exist to prevent. But a reader on turn 10 of a room now on turn 11 is looking at something stale, and silence about that is its own dishonesty — so the drift goes where a reader already looks to learn what the keys mean. The mode word is `TURN 10/11`. §9.20 declined "turn 3 of 7" and this is not that reversed: that was a progress bar offered in the *notice* line, describing a hop that had already happened. This is §7.8's always-on mode label answering which projection is live, which is the one thing the body has been ruled out of saying. Pressing **enter** is the exception that proves the rule: dispatching from a page lands on the turn just sent, because that move is the user's, not a vendor's, and a projection that answered a new brief by staying on turn 7 would show an old conversation while spending quota on a new one. #### What it does not get, and the two keys that say so There is **no column focus** on a page, so `tab` and `f` do nothing — and both are dropped from the mode line rather than left promising something, which is §7.8's surprise pointing the other way (§9.11's footer rule). The overflow markers follow: the focused-column form, no `tab to focus` and no turn coordinate, since every line on a page belongs to the same turn and the mode word already names it. For the same reason **`y` and `Y` produce the same document here**. A per-seat `y` needs a per-seat focus, and a projection whose whole unit is the turn deliberately has none — so the narrower key takes the wider document rather than guessing which seat was meant. `y yank` is named on this mode line and not on the grid's, because here the key takes the thing in front of the reader, which is what makes it worth a cell. `i` is the deliberate omission from that line. The six cells the page needs are its own motions and its two ways out; the composer is one `t` away in a mode line that names it, and it is the first row of the help panel. A footer short of width starts cutting into `?` and `q` (§9.20), and that is the trade this line was designed never to make. `t` joins the help panel by merging onto `f`'s row, inside the hard 17-row budget: `f gives one column the full width; t gives one turn the whole room`. Not a saving — the same category, the way §9.15 merged `y`/`Y` and §9.20 merged `[ ]` onto `g`/`G`. Both keys answer one question, *how much of the room is the reading area*, and a reader looking for either is looking for the other. Everything else is reuse rather than resemblance. The page plans as **one column at the full frame**, which is the tabs tier's own arithmetic, so the height budget, the 60-column floor, the composer's growth and the collapsed-seat notice are identical in both projections — a second layout path for a surface that *is* a column at full width would just be a second place for the frame to tear. The scroll window, the overflow markers, the tail and the clamp are §9.9's own argument applied once more: a page is a flat list of lines, and this room already knows how to move through one. `[` and `]` keep the words they have in the grid, at the same unit, so there is one motion to learn; `g` and `G` reach the same two positions — the oldest turn still in memory and the live end — in the projection's own unit. A turn the fifty-turn cap has evicted has no page, and says so rather than drawing an empty one: "nobody answered" and "the room no longer remembers" are different facts (§4a.1). #### Declined - **Cross-seat diff or agreement marks** — "these two agree", "this one dissents". A page puts the answers where a person can judge them; a mark would be council judging them, which no adapter sourced and no vendor reported (§4a.1). It is the same refusal as the "role" line §9.11 declined, with a harder consequence: a wrong agreement mark is one a reader would act on. - **Persisting the projection.** `room.json` stays keys-only (ADR-008, ninth amendment) and which turn someone was looking at is not state the next session should inherit — §9.9's argument for not persisting the scrollback, one surface up. - **Per-seat focus inside the page**, with `tab` cycling seats and `y` taking one of them. It would import the grid's whole focus apparatus — a marker, a weight, a hint on every marker — into a view whose entire claim is that the turn is the unit, and it would buy one thing the grid already does better. v1 lacks it deliberately, and `y`'s behaviour here is what falls out of that rather than a limitation worked around. #### Amendment, 2026-08-17: the act ledger — the same turn, read for what the seats DID **The gap.** The page above answers *what did the seats say about turn 10*. The other half of a turn is what they **ran** — the tool calls, the commands, the edits, the one that was refused at the gate — and the room has been parsing, redacting, retaining and rendering every one of those since §9.6a. What it renders them in is a 37-cell column, where the outcome is a single mark and a wrapped command is most of the width. So the record existed, in full, with nowhere to read it: to answer *did anything fail in turn 10, and where*, you scrolled a column at a time and read outcomes off four glyph shapes. **`T` opens the same turn's acts.** Not a third projection — a second FACE of the one the page already resolved. `TurnView` keeps deciding which turn is on screen; `TurnView.Ledger` decides which of that turn's two records is drawn. Everything else is untouched: `[`, `]`, `g` and `G` move the same coordinate, the scroll window and the overflow markers are the same code pointed at a different list, and the page's own geometry (one column at the full frame) is unchanged. A second `TurnView` would have been a second answer to "which turn is open" and a second scroll model to keep in step with it. **The key is SHIFT on `t`, and that is the only spelling that puts the third reading beside the two it belongs with.** `t` gives one turn the whole room; `T` gives that turn's acts the whole room. Every free lowercase letter left in this keymap is free *because it means nothing here*, and a projection filed under an unrelated letter is a projection a reader finds by accident. The capital is unclaimed — `Y` and `G` are the only two this room binds — and in compose it is the letter T, which needs no second list: `composeKey` routes any key carrying text into the draft, the contract `q`, `f`, `c` and `t` already keep. It **flips** rather than navigates (`toggleArenaDiff`'s shape, one scale up: `d` flips one seat's arena block between the stat and the whole patch), so a reader who walked back to turn 7 is still on turn 7 in either face. `t` keeps meaning the reading face from anywhere, including a re-open after a close — a `t` that sometimes landed on the ledger would be two keys wearing one name. **The outcome is a WORD, and that is what the width buys.** `⚙ Bash: go test ./... ✓ ok`, `✗ failed` with the vendor's own first line under it, `? outcome unknown`, `✗ denied by you`. The mark is `actMark`'s, unchanged; the word beside it is the signal it seconds, so `--ascii` and `NO_COLOR` lose the mark and lose nothing else. **An act with no reported outcome never renders as one that worked** — that is `runner.ActStatus`' whole reason for existing (antigravity's steps flip ACTIVE then DONE and no captured line has ever carried a success signal), and a surface that states an outcome on every line is exactly the shape that invites a default. An unresolved call splits once more: while the seat is waiting or streaming it is `running`, and once that seat has landed it is `no outcome reported`, because the vendor never said the step ended at all. The predicate is `turnEntry.working()`, not `turnEntry.Live` — the newest turn stays the column's *current* one long after every seat has finished, and reading `Live` alone would report a dead call as running for the rest of the session. **The header states the retention window, from the live constant.** `maxHistory` drops the oldest turn per seat, so "the acts" is a claim with a hard floor under it, and an unqualified one would be the room offering a record while silently forgetting the far end of it. It is a LINE hanging under the rule rather than meta on it, because `labelRuleIn` drops its meta whole when the width will not take it — correct for a route and a count, wrong for the sentence that bounds the claim, which would then vanish exactly where the room has least room to make it. The clipboard document carries it too, and there it matters more: on screen a reader re-checks the bound by pressing `[`; in a file pasted into an issue a week later that sentence is the only thing saying the record was ever bounded. **A seat that recorded nothing says `(no acts recorded)`, not that it did nothing.** A trace is a reading of what a vendor chose to report, so "this seat did nothing" is a claim no adapter here can source — the §4a.1 distinction between "we could not read this" and "there is nothing there", on the one surface a reader would take as the record of it. The same words carry the turn-level zero on the rule, so there is one spelling rather than two. A seat that SAT THE TURN OUT is absent entirely, which is §9.15's rule and binds harder here: an older turn's `git commit` filed under this turn's heading would be a history the room invented, in a document somebody pastes into a review. **`y` and `Y` follow the face.** Both keys already produce the page's own document, and `YankPage` is what keeps that promise true once there are two of them — a copy key that took the replies while the acts were on screen would break the one claim that earned it a footer cell here. The document is built from **the same `turnEntries` call** the screen renders from, and there is **no second sanitizer**: everything on `State` has already been through the one redact-and-sanitize choke point, so a cleaning step of the ledger's own would be a second answer to what is safe to put on a clipboard, and the two would differ the day one was updated. **The help panel merged, not grown.** `f / t / T` on the row that already holds both, inside the hard budget: a row of its own would push the `?` line off a 24-row terminal, which buys discoverability for one key by taking away the way out of the panel. "gives" paid for it twice — the verb is established by the first clause and the two after it read as the same sentence. The row lands at its 114-cell budget exactly. The **mode word** is `ACTS 10/11` against `TURN 10/11`: two documents at one coordinate would otherwise leave §7.8's always-on statement of what is on screen unable to tell them apart. The mode line's cells are unchanged, `t grid` included — it is still true from either face, and it is the way out this line may never shed. **What it deliberately does not get.** No note cards: how a seat's turn ENDED is already on that seat's own rule in `seatMeta`'s words, and a card under it would spend rows restating an outcome the reader has just read. No replies: that is the other face, one keystroke away in a mode line that names it, and a ledger carrying the prose too would be the page with extra rows. No new record of any kind — this section adds no content model, exactly as §9.22 added none. And **per-seat focus is still declined**, on the ruling above: a projection whose whole claim is that the turn is the unit does not grow a focus one face later. Verified offline. `ledger_test.go` pins the five outcome words staying five and an unrecognised status rendering none, the unresolved call splitting on `working()`, the retention sentence read off `maxHistory` rather than typed, the recorded-nothing wording, the sat-out seat's absence from both the screen and the paste, the face flip not moving the turn, `T` staying the letter T in compose, the gate still outranking `y`, and the help row's width and prose column. `act-ledger.txt` and its `--ascii` twin are the frame. No test here spawns a vendor. Nothing in this section is a claim about vendor behaviour, so no live run is owed: the acts it draws are the ones §9.6a already measured, at a width that can afford to name them. ### 9.23 the frame dashed, and the outline whispered while its entries shouted §9.11 through §9.22 spent the room's typographic budget on *columns* — a seat's name, its state, its cards, its transcript — and every one of them was measured against what a reader could find. What none of them looked at is the thing holding the columns apart. Read the repository's own goldens as pictures rather than as assertions and the frame is the first thing wrong with them. **The rails were a property of the prose, not of the grid.** The `│` between two columns was drawn per row, on the test *does any column have ink on this line*. That predicate exists for a real reason: a tall idle window used to draw four bars straight down through an empty screen to the footer, and Phase 2 removed them. But the room seats three transcripts of different lengths beside each other, and §9.11 spends a blank row as a boundary in three separate places — between a seat's chrome and its content, where the speaker changes, where the kind of content changes. So the ordinary case is that all three columns are blank on the same line several times per screen, and the frame blinked out on every one of them. `transcript.txt` broke at rows 11 and 13, `skips-coalesced.txt` at 5, 10 and 13, `unavailable.txt` at 19. The per-row rule solved the void and created a stutter, and a stutter is worse: an edge that dashes in and out at irregular intervals reads as damage, and it read as damage at precisely the rows where the design had placed air on purpose. **A row carries a rail when some column has content on it, or when it is a lone blank row with content above and below.** Two consecutive blanks end the band; the next word starts a new one. A separator is *structural* — it says these are different columns — and that claim is as true on a quiet row inside a conversation as on a loud one. **One row is the whole threshold, and it is the room's own number rather than a tuned one.** Every deliberate blank this surface draws is exactly one row, and §9.11 names all three of them. A one-row gap is therefore a boundary the design placed *between two things it means to keep together*, and drawing the rail through it is drawing what was meant. Two rows is nothing the design asked for — the bottom-anchor pad, an idle room, a column that ran out of transcript long before its neighbour did — and there a separator has nothing to separate. The **literal** reading was tried first and rejected on the evidence: rails on every row from the frame's first word to its last. It is a simpler sentence and it produces a worse room. An idle frame at 120×60 has chrome at the top and `no turn dispatched yet.` anchored at the bottom, so one span runs fifty-five rows of bar through nothing at all — exactly the shape Phase 2 removed, re-derived from a nicer-sounding rule. Contiguity is worth having up to the point where it starts asserting a grid over emptiness. `TestTheRailNeverDashes` and `TestRailsDoNotSpearAVoid` hold the two ends apart, and the older `TestRailsStopThroughEmptyBody` is kept unchanged as the third witness that this pass did not quietly trade one for the other. **The turn page's outline takes the weight its entries already had.** §9.22 gave a page two levels of heading — the turn's own rule at the top, then one labelled rule per participating seat — and drew the parent wholly `Muted` while `seatRule` gave every child `Strong`. The room's hierarchy upside down: the eye landed on four vendor names and had to hunt *upward* to find out which turn it was reading, on the one surface whose entire claim is that the turn is the unit. The turn rule now takes the same split every heading in this room takes — the label at weight, the rule and the numbers hanging off its end receding — which is the figure/ground rule the column header and the mode line already make, applied to a heading instead of to a key. **The grid's copy of that line is deliberately untouched, and the asymmetry is the argument.** Inside a column a turn separator sits under a seat name already at weight; there it is the child, and muted is its correct rank. On a page it is the root. The same line changes weight because it changed what it is the parent of, which is what swapping the projection means. `strongLabelRule` is one implementation for both callers, extracted for `labelRule`'s own reason: the thing being kept in step is the grammar, and a second copy would drift from it one narrow-terminal fix at a time. Weight costs no cells and `PlainStyles` renders it as the identity function, so this half moved no golden — `TestPageTurnRuleOutranksItsSeats` asserts it where colour is asserted (§9.5), and asserts the grid's separator did *not* move in the same breath. **One separator, spelled one way.** The collapsed-seat notice joined its remedy with `" │ "` — one cell of air — while the room header, the mode line and the column gutters all use two. §9.11 argues that number from `--ascii`, where the rule glyph and the spinner's first frame collide at one cell, and the notice was the single place in the product spelling the room's only separator a second way. It now reads from `gutter`, so it cannot drift again. **What was declined.** Making the rail's weight or hue say anything — it is chrome, and a frame that varied would be competing with the content it exists to bound. Drawing the rail through the bottom-anchor pad so every frame has one unbroken edge: that pad is the void, and it is the case Phase 2 was written about. And a per-column rail extent, so a short column's gutter stops early: the gutter belongs to the boundary between two columns rather than to either of them, and one of the two ending sooner is not a fact about the line between them. ### 9.24 the middle of the grid breathed and its edges did not §9.23 fixed the frame's continuity. This section is about the space inside it, and about a number that was never chosen — it was assumed, in about eighteen places, and the two halves of it had to agree by hand. **The pad was a literal, and so was its twin.** The margin between the terminal's edge and anything council draws was a bare `" "` in roughly ten builders — the header, the notice, the column grid, the tab bar, the single-column and turn-page bodies, five row shapes in the composer, the mode line, the help panel — with its arithmetic twin, a literal `2` meaning *pad×2*, in eight more places that subtract it back out to get a usable width. Those two families have to agree exactly. A builder that paints more than its arithmetic subtracts pushes the row past the terminal edge and `fit` eats the overflow in silence, which is precisely the off-by-one §9.11 found in the header's gap. `framePad` names it and `framePadStr` derives the string from it, so the paint cannot drift from the sums. **The extraction shipped as its own commit with the value still 1** — every frame byte-identical, not one golden moved — because a refactor that also changes behaviour is a refactor nobody can check. A `- 2` that is *not* the frame pad, like `labelRule`'s two cells of air around its rule, is deliberately left as a literal; the constant is not a licence to unify every 2 in the package. The extraction turned out to be **incomplete on the first pass**, and the value change is what found it: `header`'s `pathWidth` and its affordability test were still subtracting a literal 2. At `framePad = 1` that is indistinguishable from correct, which is exactly why it survived — the bug is invisible until the constant moves, and it surfaced as the header clipping `no brief` to `no brie` at 68 columns. That is the argument for the constant restated as evidence. **One to two, because a margin narrower than the gutters inside it is the wrong way round.** The interior of the grid gave two cells each side of every rail; the frame's own edge gave one. So the outermost boundary was the tightest thing on screen, the room read as crowded against the terminal, and the middle read loose — the inverse of what a grid wants. The screenshot pass that set `gutter` to 2 named that feeling exactly ("rigid / cramped") and fixed it in the one place it happened to be looking. `framePad` is now the same two, for the same reason, and the room has one number for *air between things* rather than two that disagree. It costs two cells of total width, and one of them landed somewhere worth recording: **at 80 columns the view-mode footer came out one cell over.** This room sheds whole cells rather than clipping words (§9.18), so `f expand` becomes the second rung of the shed ladder after `[ ]`. `f` and not `tab`: `tab` is how a reader reaches the other seats at the tabbed tier, which is the only tier this bites at, so shedding it would strand them on one column — while `f` is the cell §9.11 already ranked lowest, on the argument that it expands a column to a width it already has. Adding a second rung also made the shed *order* load-bearing for the first time, so it is now stated — **shed order is list order** — rather than left to a backwards walk that read as "newest first" and was not. **stripColumn goes 14 → 18, from an arithmetic floor to a reading width.** Fourteen was derived, and derived correctly: the widest phase word is nine cells, its mark costs two, and the remaining three are exactly a two-letter vendor tag and its space (§9.18). That answers what a strip's *header* cannot go below. It says nothing about the prose underneath, and prose is most of what a strip draws. At fourteen the prose shredded. §9.19's coalesced skip line — on most turns the **only** content a backgrounded seat has — came out three rows deep as `○ not` / `addressed in` / `turn 4`, with the phrase that carries the meaning split across two of them. `last: turn 8 ✓`, which §9.19 introduced with "room" as its stated goal, wrapped in a long room. A column whose every line breaks mid-phrase is not narrow, it is unreadable, and the entire point of keeping these seats on screen (§9.18) is that a reader takes them in at a glance. Eighteen is the smallest width that puts `○ not addressed` and `last: turn 137 ✓` each on one line. The header floor still holds — fourteen is still where the header itself would break, so eighteen clears it by four and §9.18's shedding ladder is untouched. The four cells come out of the primary column, and `weightedWidths` refuses the weighted split outright rather than ship a primary under `minColumn`, so at a frame narrow enough for four cells to matter the room falls back to equal columns instead of trading a readable strip for an unreadable seat. **The change paid for itself in rows.** `skips-coalesced.txt` is the clearest reading: with each block a row or two shorter, the same body height now holds seven more turns of transcript, and the overflow marker went from `↑ 8 more above` to `↑ 1 more above`. Wider columns showing *more* content is not the trade anyone expected from spending cells, and it is what happens when the alternative was spending three rows to say four words. **What was declined.** A width-dependent pad, so narrow terminals keep one cell and wide ones get two: the tier ladder already varies what is *said* by width, and varying the frame's own geometry as well would make two different rooms out of one resize. Trimming the footer by clipping instead of shedding, which is the trade §9.11's whole footer pass exists to refuse. And unifying every literal 2 in the package behind the new constant — `labelRule`'s air around its rule is the same number for an unrelated reason, and tying them together would mean a future change to one silently moving the other. ### 9.25 the panel that lists what the room can do was not listing it Three of the four items here are the same defect wearing different clothes: a surface that knew something and did not say it. The fourth is a surface that said something it did not know. **The help panel clipped in silence, and it was the only place in the room that did.** Every other surface spends a body row on `↓ N more below` when content does not fit, on the explicit argument (§9.11, columnCell) that silent clipping is indistinguishable from there being nothing more to say. The help panel is 24 rows on page one and 33 on page two against a hard budget of 17, so at the reference machine's own geometry it was dropping seven lines and sixteen — with nothing on screen to say so, and dropping them mid-word: `…the containment, not a`. A panel whose whole job is to enumerate what the room can do, quietly not enumerating it, is the sharpest available version of §4a.1's rule. **The marker's row is paid for, and the way out is pinned.** `?` is the only documented way back out of this panel, and on both pages it sat at exactly row 17 of a 17-row budget — so a marker taking the last row the ordinary way would have bought honesty with the exit, which is the trade §9.11's footer pass and helpKeys' own budget comment both refuse by name. The exit is now **chrome**, pinned to the last row the way `columnChrome` sits above a transcript, with the marker inside the scroll below it. That makes the guarantee structural instead of a lucky row count. The marker's own row is paid for the way this panel has always paid — by merging two lines that were one category: `ctrl+j` and `esc`, the two compose keys that are not `enter`, one extending the draft and one leaving it alone. Nothing was dropped to make room. **The marker names no key, and neither does the mode line.** `↑↓` do nothing over the help panel — `key()` routes no scroll to it — so the room was advertising an arrow that does literally nothing in the mode a reader is in *when they went looking for what the keys do*. Wiring a help scroll offset was the alternative and it was declined: it buys reachability for a page whose overflow is a paragraph of prose, at the cost of new state, new key routing and a new §7.1 rule-4 surface, when the honest sentence — *there is more, and this terminal is not tall enough* — costs one row and no mechanism. So the panel's mode line names only what works there (`?`, `i`, `q`), which is §9.11's own footer rule applied to a mode it had not been applied to. **The title got the room's grammar.** `council — one brief, several agents, side by side` was the only heading in the product with no rule on it, while the column header, every turn separator and every seat rule on a turn page all draw `labelRule`. A rule *under* the title is what one might expect and it is not what this room does: §9.11 spent a whole item removing exactly that shape on the finding that a heading followed by a horizontal rule says nothing the heading had not, and ruled that a heading carries its own rule. So the title becomes a `labelRule` and costs no row — which is what made it affordable against a budget with none to spare. **The blank above `? close` did not happen, and that is recorded rather than fixed.** It is wedged against the sentence before it and it should not be, but the exit sits at row 17 of 17 and a blank there comes straight out of the legend the page exists for. §9.11's ranking settles it: a rule outranks a blank, the title now carries one, and air is the boundary strength this panel can afford to go without. If the budget ever loosens, that is the first row to spend. **A seat's detail hung ten cells left of its own label.** The per-seat posture section put a seat's name at column 15, under a badge legend at column 15, and then hung the seat's measured detail at column 6 — the child left of its parent, reading as a new statement rather than as the reason for the one above it. Every card in this room has had one grammar since §9.11 (a title at weight, its body hanging under it) and this was the last place still drawing the shape that rule was written to remove. The three hard-coded numbers that had to agree — 13 for the badge column, 15 for the legend's continuation, 6 for the body — are now one `helpIndent`, checked against its own string form at init, because a panel whose continuation rows drift a cell from its key column is invisible in a diff and obvious on screen. **The vendor tag is permanent, and the wide column is now the legend for the narrow one.** §9.18 introduced `CC` / `CX` / `AG` / `CU` as what identity degrades *to* when a strip has no room for a name. Read as a whole product that is backwards: the abbreviation a reader has to know appeared exactly where they had the least context to learn it, and vanished at every width where the room had space to teach it. Drawn always, `CC Claude Code` at 37 cells is the sentence that makes `CC ✓ done` at eighteen readable, and it is the same pairing the HUD's own grid already makes. The tag is **chrome and the name is the anchor**, so the tag is muted while the name keeps the weight that says which column the keys move — asserted, because a tag at the name's weight would put a two-letter abbreviation in competition with the thing a reader is scanning for. It costs three cells of the header row and nothing else, and §9.18's degradation order is unchanged: at widths where the header must truncate, the spelled-out name goes and the two letters stay, which is the strip's one-step collapse performed gradually. **Turn pages and the collapsed-seat notice keep bare names**, and that boundary is the rule rather than an omission: the tag earns its place where columns are *scanned*, and a turn page's seat rule and a notice sentence are prose. `CX Codex (not installed)` inside a sentence is an abbreviation introduced where nothing is being compared. **One stray fact.** `unavailable.txt` drew `final only` under `⚠ Codex is not seated` — a claim about how a vendor behaves *during* a turn, stated about a vendor that cannot take one. Codex was not found on PATH; nothing about its streaming was measured. It was *plausible* — it is what the binary would do if it were installed — which is precisely the class of claim §4a.1 puts at the top of its rejected list. The badge row goes empty for an unavailable seat, and the cost cell with it (a seat that never ran cost nothing). The row stays **reserved**, because §9.11's argument for reserving it is about the grid's rows lining up and is untouched; what changes is that a reserved row now holds nothing rather than something invented. **What was declined.** Wiring `↑↓` to a help scroll offset (above). Giving the help panel its own narrower vocabulary of markers — `↓ 7 more below` is the room's existing sentence and a second one would be the second alphabet §9.11's phase marks were built to avoid. Dropping the badge row entirely for an unavailable seat, which shears the grid for the sake of a row that costs nothing to keep. And putting the tag on turn pages "for consistency": consistency across surfaces that are doing different jobs is how a room ends up with an abbreviation in the middle of a sentence. ### 9.26 one rule glyph was doing four jobs, and the header band re-textured on every dispatch §9.23 made the frame continuous and §9.24 made its margins breathe. What neither looked at is that the room draws horizontal lines at **one weight**, and asks that one weight to be four different things: the frame's own edge, a column header's leader, a turn separator inside a transcript, a seat's heading on a turn page. Every one of them is `─`, so a reader scanning for *where does the room end and the content start* gets the same ink as a reader scanning for *where does turn 3 begin*. A grid with no outline is a grid you have to reconstruct from its contents. **Two weights, one distinction: outline against interior.** `RuleHeavy` is `━` (U+2501) and `=` in the reduced set, and it is spent on **exactly three lines** — the two full-bleed rules that close the frame above and below the reading area, and the turn separator at the top of a turn page. Everything else keeps `─`. Three weights would be a hierarchy nobody can hold in their head; the value of the second one is entirely in its scarcity, which is why the list is closed and `TestOnlyTheFrameAndTheTurnPageDrawTheHeavyRule` asserts it as a *count* on the rendered frame rather than as a property of the three call sites. > **Amended 2026-08-09 (§9.44).** Two of those three lines are now one. The composer is a > bordered box, so the lower full-bleed rule is gone and the frame's closed shape is the header > rule plus the box — closure carried by corners rather than by ink. The scarcity argument here > is unchanged and one line cheaper; what the heavy weight says is now *the chrome stops here and > the seats begin*. The test still asserts a count, and the count is 1. **Why the turn page's rule is the third.** It is the only line inside the frame that bounds a whole document rather than a part of one. §9.23 gave it the *weight* of a root — the label at full intensity while its seat rules recede — on the finding that the page's outline whispered while its entries shouted; this gives it the *form* of one. The grid's copy of that same line is untouched, for §9.23's own reason: there a turn separator sits inside a column already headed by a seat name, so it is the child. The seat rules on a page and the help panel's title stay light for the same test — a heading *inside* the outline that matched the outline would restate §9.23's hierarchy defect one level down. **The weight is a parameter, not a flag.** `labelRuleIn` takes the fill glyph and `labelRule` passes `g.Rule`; a caller that wants the heavy rule has to name it at the call site. That is what makes "exactly three lines" checkable by *reading* the three call sites rather than by grepping for a bool, and it keeps one implementation of the grammar — a label, a rule, optional numbers, two cells of air each side — which is `labelRule`'s own extraction argument. **It is a character before it is a style.** `--ascii` gets `=`, not a fallback to `-`, so the outline survives on exactly the terminals least able to infer it; `NO_COLOR` never touched it, because weight of this kind is a glyph rather than an attribute. `=` is the one unclaimed mark left in the reduced set — `-` is the light rule, the `Range` joiner and the first spinner frame, `|` the separator, `>` the ellipsis, `]` focus, `!` the warning prefix, `^`/`v` the overflow markers, `*` Act, `.` Idle, `:` the prompt, `_` the caret, `+`/`x`/`?` the outcome marks, `/` and `\` the remaining spinner frames, and `#` the HUD's gauge fill. It is also the only unclaimed character that reads as a *doubled* `-` rather than as a different symbol, which is the one property a second rule weight needs. `TestTheHeavyRuleHasAnUnclaimedASCIIPartner` enumerates that whole list so the next glyph cannot be added without meeting it. **The header leader stops depending on phase.** `headerUsesLeader` was false for an idle seat, on an argument that was true at one rule weight: a long `────` between `Claude Code` and `○ idle` was *filling* rather than separating, whitespace does that job for free, and a room with a single rule weight cannot afford ink on nothing. With two weights the leader is no longer "the rule" — it is the interior weight, and its claim on that row is *this name and this state belong to one seat*, which is as true of an idle seat as of a streaming one. The observable defect is the sharper half of the argument. A room where one seat is answering drew the seats' header band as one continuous ruled line across part of the frame and blank across the rest — **one row, two grammars** — and re-textured itself the moment a turn started and again when it ended. §7.1 rule 4 keeps this room still by default, and a band that changes shape on every dispatch is the loudest still-frame change on screen, spent on a fact the state word beside it already states. The air the old comment wanted is not lost: `labelRule` keeps two cells each side of its rule, which is the gap that keeps an ascii spinner (`-`) legible against an ascii leader (`-`). **Golden churn is the whole visible change, and it is two lines per frame plus one.** Every frame's two rules, and every idle seat's header row. Nothing else moved — `PlainStyles` renders both weights as themselves because they are characters, so unlike §9.23's weight half this pass *is* visible in the goldens and had to be read frame by frame. The three test helpers that found the frame by searching for a run of `─` (`fullWidthRule`, `frameBody`) now search for `━`, which makes them stricter rather than merely different: a column header's leader can no longer be mistaken for a frame edge at any width. **What was declined.** A third weight, or a double rule (`═`), for the turn page — the page's rule is already distinguished from its seat rules by its label, its position and its meta, and the frame is the only thing it needs to *match*. Making the frame's rule brighter as well as heavier: §9.23 declined to let the rails' hue mean anything on the argument that chrome competing with content is the wrong trade, and an outline is chrome. And keeping the idle leader off "for quiet": the quiet was bought by making the room's most stable row the one that changed most. ### 9.27 focus was a mark on one row, in a frame the reader had scrolled past §9.12 fixed the focus signal by adding the `▸` and moving the load-bearing half onto the seat name's *weight*. Both of those live on the column header — row one of a body that is twenty rows tall — so a reader forty lines into a transcript, comparing two answers, had nothing on screen at all telling them which column `↑↓` would move. The signal was correct and it was in the wrong place: it described a column and was as tall as a line. **The focused column's LEFT rail thickens.** The gutter cell immediately left of the focused column draws `▌` (U+258C) instead of `│` — same cell, same width, one glyph heavier — for the full height of the band. It is the only mark on this surface that is as tall as the thing it describes, which is the whole reason it is worth a glyph. The `▸` and the name's weight stay: word/glyph-first means two carriers on two rows, not one carrier moved. **The leftmost column has no gutter, so the frame's left pad carries its mark.** `framePad` is two cells since §9.24, and the mark takes cell one — which leaves exactly one cell of air between it and the column, the closest the geometry gets to the gutter's two. Without this, position zero would be the one seat the device could not mark, and a signal with a hole in it is a signal a reader stops trusting. **It rides §9.23's band exactly.** The thick rail spans the rows the thin one would and no others, so focus cannot spear a void either — an idle 120×60 room still has a bare middle, and `TestTheRailRidesTheSameBandTheThinOneDoes` asserts it against the same `bare > 0` test §9.23 wrote. Focus does not get its own answer to a question the frame already settled. **Unfocused columns' prose steps back one contrast level.** `Dim` is `Text` + `Faint`, applied to the *reading area* of a column the keys do not move: the vendor's reply, the §9.14 stand-in for a reply that has not arrived, and the `no turn dispatched yet.` line. That is crush's `Focused`/`Blurred` pair applied to prose rather than to a border, and it is the half of this pass that costs no cell at all. **The faint collapse is accepted, and here is the accounting.** Council has two intensities — `Text` and `Muted` — so a demoted body renders identically to chrome, and inside an unfocused column prose and chrome do arrive at one intensity. What is lost is the *second* signal, on a column the reader is not reading: every distinction between them is carried by shape first (a turn separator is a labelled rule, a trace entry opens `⚙`, a skip line `○`, a note `⚠`), which is §7.1 rule 2 doing exactly the job it was written for. The alternative — a third intensity in `internal/theme` — would spend a shared palette token, on a surface the statusline does not have, for a distinction only the unread column needs. **What the demotion does NOT reach, and each exclusion is a rule rather than a taste.** - **The chrome above the body.** `columnCell` renders the header, the badge row and the gate card with the room's set and only the body with the seat's. A posture badge is a safety claim, and a claim that faded because the reader was looking at the next column is precisely the defect §9.2 wrote the reserved badge row to prevent. - **The prompt echo.** The user's own words stay `Strong` in every column. What a seat was *asked* is the thing a reader scrolls looking for (§9.9), and it is not the vendor's prose to demote. - **Notes and cards.** A failure note, a reattach card, an unavailable card and the thread-cleared sentence under its rule all keep their own styles. This is the one place the ratified shape was **narrowed** during implementation: the thread-cleared sentence is prose in the reading area by position, but it is the body of a card in the room's grammar and it says what the *next* brief will do — an actionable claim about the seat, in the same category as the reattach card whose wording it shares. Leaving one of that pair full-contrast and demoting the other would be two spellings of one fact. **The rail is a columns-tier device, and says so.** The tabs tier has one column on screen with a tab bar above it already carrying `▸` and the selected tab's weight; a rail there would mark the only thing there is. Expanded is the tabs tier by `tierFor`'s own rule, so it inherits that answer rather than needing its own. A turn page is one reading area and has no unfocused seat to demote. **Under `NO_COLOR` and `--ascii` the whole distinction still lands**, and that is the test the demotion had to pass to be allowed at all: `▌`/`[` in the gutter, `▸`/`]` before the name, and the name's own weight all survive both, so a monochrome terminal loses the contrast step and keeps every carrier that was doing the work. `[` is the ascii rail — `#`, the obvious candidate, is refused for the reason `ActOK` refused it (it is the HUD's ascii gauge fill, and one product means one vocabulary), and of what is left `[` is the squarest vertical stroke in the set, faces the column it marks, and mirrors `]`, which is already this room's ascii focus mark. The `[` and `]` in the mode line are key *names* in the footer's prose, never marks in the grid — the same slot argument `Range`'s doc makes for the hyphen. **Golden churn: the rail only.** `▌` is a character, so every columns-tier golden moved by exactly one cell per railed row; `Dim` is an attribute rendered by `PlainStyles` as the identity function, so it moved nothing. The whole diff was verified mechanically — every added line with the rail glyph mapped back to a space is byte-identical to the line it replaced. One golden is **new**: `focus-rail.txt` pins the focused column in the *middle* of the frame, the shape no pre-existing golden reached because all of them render with focus at position zero. **What was declined.** A rail on both sides of the focused column, which is a box and turns a gutter shared between two seats into a property of one of them (§9.23's own last item). Colouring the rail: chrome that competes with content is the trade §9.23 refused. Running the thick rail the full body height so focus always has an unbroken edge: that is the void again. And demoting `Muted` chrome a further step in unfocused columns, which would need the third intensity this section just declined to buy. ### 9.28 the room's one hue exception, and exactly how far it goes `internal/council` has said "adds no hues of its own" since §9.11, and the rule was right: a dispatch room that invented a sixth colour drifts from the visual language the statusline and the HUD share, so council spent WEIGHT (§9.11) and CONTRAST (§9.27) instead, both attributes rather than hues. **This is the one ratified exception (San, 2026-08-07), and it is an exception to the rule rather than a repeal of it.** **The concept the other two surfaces do not have is the SEAT.** Everything council renders that theme already has a token for — severity, identity, chrome — keeps that token. What has no token is *which of four agents is speaking*, because `telltale statusline` and `telltale hud` have no seats to distinguish. `seatHue` returns one ANSI index per vendor: claude `5` (magenta), codex `6` (cyan — theme's identity hue, kept by the seat that already had it), agy `4` (blue), cursor `12` (bright blue), and `theme.ColorIdentity` for anything else. **Why it lives in `internal/council` and not in `internal/theme` — and the stdlib rule is NOT the reason.** These are plain strings; they would compile in theme perfectly well, and citing ADR-002 here would send the next reader to fix the wrong thing. The reason is theme's *own* contract: one hue, one meaning, across every surface that imports it. A per-vendor hue promoted to theme is a token that means nothing on two of the three surfaces, which is how a shared palette stops being shared. **Why 4-bit indices.** theme.go's own argument, unchanged and reused rather than restated: the terminal resolves an index against the scheme the user already chose, so the room looks native in Windows Terminal's default and in a light scheme with no second palette and **no `isDark` fork**. A hex triple would be council asserting a colour over the user's own. **What is off limits, and it is a fence rather than a guideline.** The severity family — `1`/`2`/`3` and their bright twins `9`/`10`/`11` — is the green/yellow/red ramp on every surface, and a seat wearing red would read as a seat that failed, on a row where `✗ failed` is the thing beside it. The chrome family — `0`/`7`/`8`/`15` — is the gauge track and the terminal's own fore/background. That leaves 4, 5, 6, 12, 13, 14; this spends four of them, and `TestNoSeatHueIsASeverity` fails the build if that stops being true. **The honest weakness: 4 and 12 are one hue at two intensities.** agy and cursor are blue and bright blue, which some terminal schemes render close together and a reader can miss. That is acceptable **here and only here**, because §9.25 made the two-letter tags permanent — `AG` and `CU` appear beside every seat name the room scans — so the hue is the second signal it is supposed to be and the tag is carrying the distinction. If a fifth seat arrives wanting blue, the tag is what still works and the hue is what has to be argued for. `TestSeatHuesAreExhaustive` asserts the room seats exactly four vendors, so a fifth cannot be added without somebody reading this paragraph. **Three sites, and the list is closed.** 1. **A turn page's seat rules** (`seatRule`). The highest payoff by a distance: a page stacks every participating seat in one column, one block after another, so position answers *nothing* about who is speaking — which is the exact condition under which a hue earns its place. 2. **The tab bar.** `SeatStrong` selected, `SeatIdentity` unselected, replacing the wholly-muted unselected tab. That is a *promotion*, and the opposite of what §9.27 does to an unfocused column's prose, deliberately: prose in a column you are not reading is content you are not reading, while an unselected tab is a **destination**. It is the one row on that tier whose job is "here are the other seats, pick one", and a menu whose entries are faint makes you read it twice. The selected tab still outranks the rest by weight and by the `▸` in front of it, which is what survives NO_COLOR. 3. **The collapsed-seat notice**, names only. The `⚠` keeps `SevWarn`, the reason in parentheses and the remedy after the bar stay chrome, and nothing there gains weight — it is a sentence, and a sentence with four bold words in it is not one. §9.25's boundary is untouched: the two-letter *tag* stays out of prose, because an abbreviation introduced mid-sentence is one nobody can learn there. A hue is not an abbreviation — it costs no cell and teaches nothing new. **Where it is deliberately NOT spent.** - **Grid column headers.** Position already answers which seat this is, and four coloured names across one row is the circus row this rule exists to prevent — the room's newest signal spent on the one question the layout had already settled. - **Phase marks and status words.** Severity owns those cells (§9.7). - **Rules, leaders, badge rows and every other piece of chrome.** A posture badge is a safety claim (§9.2) and must not compete with a name for the eye. **Constructed to be invisible to the goldens, rather than checked to be.** `SeatIdentity` and `SeatStrong` are `Identity` and `Strong` *retinted*, through one `retint` helper that returns the base style untouched when `Plain` is set. A second pair of literal constructors would have to remember that and would forget it the first time one grew a second attribute. So **golden churn on this pass is zero, and any golden diff on it is a bug** — which is also the whole verification story, since colour is asserted where colour is asserted (§9.5) and never in a golden. **What was declined.** A hue on the grid's column headers (above). Hue on the vendor tag as well as the name, which doubles the ink for a distinction the name already carries. Truecolor, which would override the user's scheme. And a fifth hue held in reserve for "the next vendor": a palette entry with no seat behind it is a decision nobody has made, recorded as if somebody had. ### 9.29 the seats had positions and no way to address one `tab` cycles focus, and at the columns tier that is fine: three seats, at most two presses. At the **tabbed** tier — the narrow terminal, the one a laptop actually runs — one column is on screen and reaching the fourth seat costs three presses, each of which redraws the whole frame and shows you a seat you did not want. The room had four seats sitting in a fixed order, drawn in that order on every surface, and no way to say *that one*. **`1`–`4` focus the Nth VISIBLE seat, in seating order.** Positional, exactly like the columns are. A room with two seats has keys 1 and 2 and nothing else: `3` there is a **no-op**, not a wrap and not a clamp, because a key that quietly lands somewhere else is §7.8's surprise and a wrap would make the number stop meaning the position it is printed at. In **compose** a digit is a digit — the same contract `q`, `f`, `c` and `[` already keep, and it needs no second list: the handler tests whether the key carries text, which is what makes it text in the composer. **The number is drawn where the key acts.** `▸ 1 CC Claude Code ──────── ✓ done 8s` in the seat header, and `▸ 1 CC Claude Code 2 CX Codex` on the tab bar — the two places a seat name heads a reading area, which are the two places §9.25 already put the vendor tag for the same reason. The number is **muted**, on the tag's own argument: it is chrome and the name is the anchor. It sits in FRONT of the tag rather than after the name, because it is what a reader's eye runs down the row of headers looking for, and because a number at the far right would sit beside the state word where every other number on that line is a duration. **It sheds last, and that is a new rung reasoned about rather than an appended default.** §9.18's ladder drops the clock first, then the focus mark, then the tag. The number goes below all three: `1 CC ⚠ unavailable` is exactly eighteen cells, so at `stripColumn` the full form fits every phase word, and below that the tag goes before the number does. The argument is the one §9.18 itself used — it shed the focus mark because the load-bearing half of that signal had moved somewhere free, and it kept the tag because position alone was a weak identity. The number is not a second spelling of anything: it is the key that reaches this seat, at the width where reaching seats is hardest, and **a key nobody can see is a key nobody presses** (§9.10, which is the whole reason this room names keys on its overflow markers at all). **The footer names `1-N`, not `1-4`.** The range is however many seats are on screen. A three-seat room naming a `4` would promise a key that does nothing — §7.8's surprise, which this line already refuses in the other direction for `tab` and `f`. It is the **third rung of the shed ladder**, appended after `[ ]` and `f` so §9.24's order is untouched, and it is the last of the three to go: shedding only bites at the tabbed tier, which is precisely where the number is worth most. `[ ]` sheds first because `g` and `G` still reach the ends of the transcript; nothing else reaches seat 4 in one keystroke. `? help` and `q quit` remain unsheddable, asserted. **A room with ONE seat on screen has no numbers at all**, and that is §9.11's rule applied to a third key rather than a special case: `f` and `tab` are dropped there because they address a choice that does not exist, and a number labelling the only column there is spends a cell on the same nothing. `State.SeatNumber` and the footer's cell run off the same predicate, so the key's label and its advertisement appear and vanish together — a footer naming a key the header did not would be one surprise split across two rows. **Renumbering, and the still-by-default wrinkle it is.** Because the number is a position, a seat folding out **renumbers** every seat after it: with Claude uninstalled, Codex is seat 1. That is a label changing under a reader, which §7.1 rule 4 does not hand out lightly — and it is bounded by *when* it can happen. A seat collapses, or `--vendor`/`/seat` reseats the room, and both already reflow the entire frame: the column widths change, the notice row appears, the grid is visibly a different room. There is no path where the numbers move on a frame that was otherwise going to look the same, and in particular none mid-turn. The help panel says "by position" rather than implying a seat owns its number. **The help panel merged, not grown.** `tab / 1-4` on the row that already named `tab`, because the budget is hard at 17 rows and these are one question asked two ways — step to the next seat, or go straight to one. "move" paid for the characters. The `?` row, the panel's only documented way out, is exactly where it was, and the `↓ 5 more below` marker's count is unchanged. **What was declined.** `alt+1`–`4`, so digits could stay digits in both modes: it buys nothing — compose already routes text keys to the draft — at the price of a chord nobody discovers and that several terminals eat. Numbers on a turn page, which has one reading area and no focus to move. Stable per-vendor numbers that never renumber, which would leave gaps (`1`, `3`, `4` on screen) and make the printed number disagree with the position it is printed at — the number would then be an identity, and identity is what the tag and the hue are for. And a fifth key for a fifth seat: `1-N` already says how many there are. ### 9.30 one question, asked once, instead of four times across the comparison surface Council exists to put several answers side by side, and §9.22 built the page that reads a turn as one document. What neither of them fixed is what the **grid** does with the question itself. §9.9 ruled that the echoed brief is a fact about the COLUMN — a turn can reach two seats and not a third, seats skip turns, and a transcript that filled the gaps would be the room inventing a conversation — so every addressed column echoes it. On a one-seat route, which is the ordinary turn since the default stopped being everyone, that is exactly right. On a committee route it is the same paragraph two, three or four times across the top of the reading area, each copy pushing the answer it belongs to a row further down, with "+ the other seats' last answers were quoted to this one" repeated underneath every one of them on a rebuttal turn. The surface built to compare four answers spent its widest rows agreeing with itself about the question. **So the live turn's brief is drawn once, full width, as a band under the room chrome, and the addressed columns stop echoing it while it is up.** Nothing new reaches `State`, nothing is stored, and no column records anything different: this is a rendering rule over the same echo §9.9 already holds, sanitized and deliberately unredacted for §9.9's own reason — it is the user's own typing shown back to the user, and covering it would hide a secret from the one person who already has it while doing nothing about the copy just sent to three vendors. **History is untouched, and that is the boundary rather than a scope limit.** A finished turn's echo stays inside the column that took it, because §9.9's argument is about a *record*: turn 4 is filed on the two seats it reached and absent from the one it did not, and each column's transcript is that seat's own conversation read top to bottom. The band speaks for one turn — the live one — and it identifies that turn by number (`Column.TurnN == State.Turn`), never by "this column has a prompt". A column's prompt block outlives its turn: a seat that answered turn 4 and sat out 5 and 6 is still displaying turn 4's brief as its current block, and that is its own conversation, not the live turn. **It appears at dispatch and retires at the next one.** Both are keystrokes, and §7.1 rule 4's still-by-default frame is about what moves *without* one. The retirement moment is deliberately the push to history and **not** the instant the last column lands: §9.21 retires the live route there because the header describes the present, but that landing is a vendor finishing, and a band that vanished on it — restoring three per-column echoes and reflowing every column — would be precisely the mid-turn layout jump the rule forbids. Tying the band's life to the same block the per-column echo already had means there is one reflow point, it is the user's own enter, and the frame either side of it is one the reader asked for. **Two seats is the threshold.** One echo on screen is not duplication, and hoisting it would move the user's words away from the answer to them and buy a row of chrome for nothing. The band is therefore a columns-tier device: the tabs tier draws one column at a time, `f` resolves to that tier, and a one-seat room has nothing to compare — in all three the column keeps its own echo. A turn page is excluded for the opposite reason: it already prints the brief once, which is half of what it is for. The help panel is excluded because it replaces the column area outright. **The anatomy is §9.9's echo hoisted, and §9.11's middle boundary under it.** The composer's own `›` at full weight, because the glyph carries "you said this" before the colour does and the user's words are the anchor a reader navigates by; the rebuttal notice once, muted, underneath, only when true, in the same sentence the column prints (one constant, not two spellings); then a **blank row**. Of the three boundary strengths this room ranks — a labelled rule where the turn changes, a blank where the speaker changes, a blank where the kind of content changes — the band's is the second: the user stops and the seats start. A rule was refused. The frame's own full-bleed heavy rule sits two rows above it, and a second horizontal line under that would rebuild §9.11's "one rule per column instead of two three rows apart" at the room's own scale, with §9.26's heavy/light distinction blurred as well. The band also states **no route**: the header already carries `turn 10 → everyone` on the cell that names the turn, and repeating it would be the second copy this whole section exists to delete. What each column keeps is its own turn separator, which is one line saying which turn the lines under it belong to rather than the same paragraph again. **Four rows, and the fourth is the marker.** A brief worth sending to three agents can be a paragraph — the composer grows to six rows for that reason — and a band as tall as the draft would eat the reading area it was written to protect. So the band spends at most four rows on the brief, and when it needs more the fourth row is a **truncation marker** rather than a fourth row of text: how many rows are missing, and that the turn page has the brief whole. Silent clipping is the ambiguity §4a.1 forbids, and it is worse here than anywhere else on screen — a reader cannot tell their own question from a truncated copy of it. The `t` that opens that page is named on the marker in **view mode only**, because `t` is the letter t while composing and a marker advertising it there would promise a keystroke that does something else (§7.8, scrollHint's rule for `f`). The count and the destination survive in both modes; only the keystroke sheds. **The band's rows are room chrome, spent where the notice line is spent.** `resolveLayoutIn` settles the tier first — §9.5's ordering, unchanged, and the band depends on the tier so it could not be spent any earlier — then header, footer chrome, tab bar, collapsed-seat notice, **band**, then the composer, which still yields before the body. The band is budgeted before the composer and tested against the composer's *floor* rather than its current height, deliberately: a band that retired because the draft grew a row would be a layout jump on a keystroke mid-turn, and it would jump back on backspace. **Below a floor the band yields ENTIRELY, and the fallback is a pure function of height.** If spending the band would leave the columns fewer than eight body rows, it is not spent at all and the columns echo the brief themselves — the pre-band frame, byte for byte. Eight is measured from what a column draws before a word of the reply: three rows of `columnChrome` (name, posture claim, one blank) and the live turn's own separator, leaving four rows of answer. All-or-nothing rather than shedding a row or two off the band, because a half-band is the worst of both: a cut question above columns that no longer say what they were asked. One number decides it — `Layout.Band` — and the renderer and the columns both read that one number, so a band without suppression (the brief three times *and* at the top) and suppression without a band (the brief nowhere on screen at all, which is a §4a.1 failure with the user's own words as the missing content) are not states this code can reach. **Scroll detaches from the band, on purpose.** A column scrolled back into history draws its history under the band exactly as it always did, and the band stays — including when *every* addressed column has been scrolled away. It describes the live turn, which is a fact about the room, not about any viewport; a reader who went looking for an older answer is still in the turn they dispatched, and the question that produced it does not stop being true because they scrolled. **What was declined.** Hoisting past turns' briefs the same way, which would flatten §9.9's per-column record into a room-level one and lose which seats a turn actually reached. Keeping a one-line stub in each column to mark where the brief would have been: it costs the row the band was spent to save, once per column, and the turn separator already marks that boundary. Shedding the band down to one row on a short terminal, covered above. And giving the band a rule glyph or a hue of its own — council adds no hues, and the boundary vocabulary this room already has is what a reader has already learned to read. ### 9.31 a word the room did not know was billed to three vendors **The rule: a draft that opens with `/` and names no room command is refused, not dispatched.** Refusing is free — nothing spawns, nothing is billed, the draft stays in the composer — and the alternative is not free at all. **The field report.** Turn 53 of a real room: `/unseat codex` was typed. There was no `/unseat`; §9.17 shipped `/seat ` and nothing else. `roomcmd.go` recognised no command, so the draft fell through **as a brief**, and the committee was billed to discuss the string `/unseat codex` until the user cancelled the turn. Nothing malfunctioned. Every line of that was the documented behaviour: "only a draft that IS a command is intercepted; anything else, including text that merely starts with a slash, dispatches to the vendors as typed." That fall-through was the right call for the *vocabulary* question and the wrong call for the *typo* question, and the two had never been separated. The vocabulary rule exists so the room does not steal words out of the conversation — the argument that kept `/clear` out of `roomcmd.go`, since `/clear` is a real Claude Code command a person means for a vendor. It says what must not be **executed**. It says nothing about what should happen to a draft that executes nothing, and dispatching was only ever the default that was already sitting there. **A leading slash is almost never prose.** It is a command the room does not have, a command a *vendor* has, or a typo for one of ours. Against that, dispatching costs a turn on every seated vendor — on the scarce independence pool as readily as on the cheap lane — for a line the user will retype in five seconds. `addressesRoom` is the whole test: a slash in **column one**. #### The escape hatch is one space, and it had to be A brief that legitimately opens with a slash is a real thing to type — a POSIX path, a regex, `/etc/hosts is wrong`. Prefixing one space sends it, unmodified, to the vendors. It is the cheapest honest escape available, and it is honest because nothing between the composer and the spawn trims it. `sanitizeKeepingSpace` deliberately does not trim ("trimming would make the string on screen disagree with the string about to be dispatched", §7.14's rule applied to the composer), `ParseRoute` returns an unconsumed draft unchanged, and `dispatch` echoes what it sends. The space the user typed is the space the seat receives, and `TestALeadingSpaceSendsASlashBriefToTheVendors` asserts that at the seat rather than at the parser. **The refusal has to say so, in few words.** §9.17's own defect shape is a refusal whose remedy is undiscoverable — the `/flow` write-hop notice that went on naming a flag for two releases after `/write` made it wrong. So the notice carries three clauses **in this order**: what failed, how to send it anyway, then the vocabulary. > `no room command /unseet — a leading space dispatches it · /cd /flow /read /seat /trace /unseat /write` The order is load-bearing. This notice replaces the entire hint stack on the mode line, which truncates from the *right*, so the clause a narrow room loses has to be the one a reader can get elsewhere. `?` lists the room controls; nothing else on screen teaches the space. The quoted word is capped (`unknownVerbEcho`) for the same reason — a pasted 200-character path is one word, and an uncapped echo would push both the remedy and the vocabulary off the end, leaving a refusal that names only the mistake. **The vocabulary in that notice is walked, never written twice.** `roomVerbs` is the one table; the notice reads it and `TestTheRefusalListsTheLiveCommandTable` walks it. A hardcoded list in either place is the copy that goes stale on the next command — and this feature would have been its first victim, shipping `/unseat` with a refusal that did not mention `/unseat`. **What this does to the bare-word rule, which is the one deliberate consequence.** §9.17 made `/read` and `/write` bare-only so that "/read the design doc first" could not silently swallow a turn and run a setting. That rule is untouched: those drafts still do not reach `postureCommand`. What changes is where they go instead — refused with the space named, rather than billed. Both halves are now one test (`TestBareWordOnly`), because either alone is a defect: a "/read the design doc" that ran the setting, and a "/read the design doc" that cost three seats a turn. **`/flow` came along with it.** `dispatch.go` matched `strings.HasPrefix(TrimSpace(draft), "/flow")` — any draft whose first non-space characters were those five letters. So `/flowchart the auth path` was an orchestration, and, worse, `" /flow/gate.log is the file I mean"` would have been swallowed *after* being escaped, making the hatch a lie for exactly one prefix. An escape hatch with an invisible exception is not one. `isFlowCommand` applies the room's single vocabulary rule there too. #### `/unseat `: `/seat` spelled the other way round The typo that started this was reaching for a control that should have existed. `/seat` names who **stays**; `/unseat` names who **leaves**, and it is the argument `-@` makes one control up: the correction a user reaches for mid-session is "not that seat" — one vendor is answering badly or expensively and the other three are fine — and making them retype the complement is arithmetic done at the keyboard, on the one line where getting it wrong quietly reseats the room around seats they did not mean. It is `parseSeatList`, literally: same aliases, same `@` tolerance, same trailing punctuation, same dedupe. A second list parser is how `/seat agy` would work and `/unseat agy` would not. Everything §9.17 ruled for `/seat` holds unchanged, and mostly by sharing the code rather than by sharing the intention: - **It kills nothing.** An unseated seat keeps its thread, its process and every id that would resume it. `/seat all` puts it back mid-conversation, with no resume to fail. - **It refuses mid-turn.** The roster is dispatch state — `frameOwnersFor` decided this turn's grid — so reseating under a live turn would redraw the room around columns that are mid-answer. - **It warns when it removes the default route**, because unaddressed briefs go to claude. The warning lives in `applySeats`, shared with `/seat`, precisely so the subtractive spelling cannot be the one that says nothing. - **Bare `/unseat` reports**, the way bare `/cd`, `/trace` and `/seat` do: a command that half-asks a question answers it rather than doing something. Three refusals are its own. **The last seat**: a room with no seats can answer nothing, so the subtraction that would empty it is refused in `/seat`'s words for `/seat`'s reason. **`/unseat all`**: a sentence someone will type, answered as the empty room it names rather than left to `parseSeatList`, whose honest report would be "no seat called all" — a spelling complaint about a word the room understands perfectly well. **A seat that is not in the room**: distinguished from a typo, and both change nothing, on `/seat`'s argument that a command which quietly did less than it was asked is discovered several turns later as a seat still answering. **Membership is what the room SHOWS, not what it can drive**, and that line took a CI failure to find. `/seat cursor` *forces* an uninstalled seat on screen — "a user who asked for it is owed the card explaining why it is not there" — so that seat is in the room in every sense a subtraction cares about, and the first spelling, which tested membership with `seatsVendor`, could not remove the one card a user is most likely to want gone. On a machine where nothing is installed it could not remove anything at all: every `/unseat` was answered "not in the room" and the roster never moved. Local runs passed because the developer's machine has four vendors on `PATH`; CI, which has none, is the one that reads the rule as written. **So the two questions are separated, and only the second guards the last seat.** *Is it in the room?* is `shows`. *Can it answer?* is `Avail == AvailInstalled`, and the refusal built on it is **conditional on the room having had one**: a room with nothing installed could not answer before this was typed either, and refusing there would blame `/unseat` for a state it did not cause. An empty roster is refused unconditionally, because that is `/seat`'s own "at least one seat" reached by subtraction. #### The focus bug underneath it `/seat` has been able to unseat the **focused** column since §9.17, and nothing moved focus when it did. `State.Focus` indexes `Columns`, so it went on pointing at a seat the grid no longer draws: the focus mark vanished from the room, and `f`, the scroll keys and `y` went on addressing the hidden column. Keys that still work over a transcript nobody can see are worse than keys that stop, because nothing on screen says anything is wrong. `stateWith` already does this once, at launch — "focus lands on a column that is actually drawn" — and the fix is that same rule applied wherever the roster moves (`rehomeFocus`), not a rule of `/unseat`'s own. It is called from `applySeats`, so `/seat` gets it too; a helper that fixed only the new command would have left the older one holding the bug that made the new one worth writing a test for. #### The help row `/unseat` merged onto the row `/seat` already holds, as `/seat /unseat `. The panel's budget is hard — 17 rows to the `?` line on a 24-row terminal — and `helpBody` clips without scrolling, so a control named past the fold is not a demoted control, it is an absent one (§9.20). The merge is also the honest shape rather than only a saving: the two take one argument in one vocabulary and differ only in direction, so a reader who finds either has found both. "times" paid for the width — the row is a list of controls, and `/trace ` is unambiguous without the verb. #### Amendment, 2026-08-17: `/retry` sends the last brief again, to the seats that owe an answer **The gap.** §9.37's 2026-08-17 amendment gave the operator a way to stop ONE seat of an ordinary turn, and it argued from the room's most probable live failure: five seats on an `@all` turn, four answers, one vendor that fails or stalls. The room can now end that turn. It has no way to finish the brief. The only act left is to retype the brief and retype the mentions — arithmetic at the keyboard, on the one line where getting it wrong bills seats that already answered. That is the same complaint `-@` and `/unseat` were built for, one turn later. **`/retry` puts the last dispatched brief back in the composer, addressed to the seats that did not answer it.** The brief comes from the columns' own per-turn record (`Column.Prompt`), unchanged. The mentions are the seats that owe an answer, so the draft reads `@codex @agy ` — a draft the operator could have typed, in the grammar that already exists. **It arms; it does not dispatch, and enter is still what spends the money.** The verb writes a draft, `setDraft` re-derives the route from it, and the footer prices that route through the same `State.SeatsIn` intersection dispatch gates on (§9.21). So the operator reads the bill before paying it, and can edit the draft — drop a seat off the front, fix a word — because it is an ordinary draft. A verb that spawned on the spot would spend up to five quotas on a keystroke that named none of them, and the room would have no surface left on which to say which ones. **What counts as an ANSWER is defined against the four endings**, and it is narrower than "the column looks empty": - **`PhaseDone` answered**, including the seat whose body reads `[Turn completed with 0 text chunks streamed]`. That is a measured zero, not a missing reply (§4a.1), and re-sending on it would be the room overruling a vendor's honest empty answer — and billing a seat that did the work. - **`PhaseFailed` and `PhaseCancelled` did not.** Cancelled covers `ctrl+c` and the per-seat give-up together, deliberately: a seat the operator cut is the case this verb exists for, and after the turn the room holds one cancelled phase for both. - **A seat that SAT THE TURN OUT is not a candidate at all.** `dispatch` never calls `startTurn` on it, so its `Column.TurnN` still names the last turn it took and the scan skips it. That is load-bearing rather than incidental: without it, a `/retry` after an `@codex` turn would widen the bill to four seats the operator deliberately did not address. **Bare-only, on `/read` and `/write`'s rule.** The verb takes no argument, and "/retry the failing test" is a sentence someone types. A verb that swallowed that argument would run a re-send and discard the brief — §9.17's vanishing-brief failure. The bare draft is the command; anything longer is refused with the space escape named, which costs nothing. **Three refusals, three sentences, and only the first keeps the draft.** *A turn in flight*: the phases this verb reads are not settled, so any list it produced would be a claim about a turn that has not ended — and the operator still wants the verb one turn later, which is `postureCommand`'s own reason for holding the draft. *No brief on record*: turn 0 and the degenerate turn whose brief sanitized away to nothing are one sentence, because they are one fact. *Every seat answered*: there is nothing to re-send, and saying so beats a composer the operator has to clear by hand. The verb sits in `roomVerbs` like every other, so §9.31's walked refusal teaches it for free — no second copy of the vocabulary was added, and the notice fits the reference width with one cell to spare. **`--brief` stays unfiled, and this verb is careful not to file it.** It re-sends the brief UNCHANGED: no re-briefing, no edit, no automatic second attempt, and it never reads `Model.brief`. §9.17's sweep left `--brief` out on the ruling that first-turn context is a different feature from re-briefing; that question is still open and still to be decided on its own. **After a race it re-sends the race's brief as an ORDINARY turn.** `Column.Prompt` holds the brief the racers were given and never the `/arena` draft that wrapped it, and this verb invents no grammar to put the wrapper back. The composer shows exactly what will be sent before enter, which is where the operator reads that the worktrees are not part of it; `/arena ` races again. **The help panel does not name it**, on `/adopt` and `/arena drop`'s precedent: the room-controls row is at its budget (§9.20), and the verb is taught by the slash refusal and by this block. Verified offline only. `roomcmd_test.go` pins the four endings producing the right seat list, the measured zero counting as an answer, the three refusals, the bare-word rule, the verb's appearance in the walked refusal table, and — at the dispatch level, with `countSpawns` — that the re-send spawns one process per seat owing an answer and none for a seat that already replied. The `slash-refusal` goldens and their `--ascii` twins carry the new word. No test here spawns a vendor. A live `/retry` on the Windows reference box is not owed as a separate payment: nothing here is a claim about vendor behaviour, and every process the verb can cause is an ordinary turn's spawn. ### 9.32 the room remembered where it was and forgot who was in it **The ruling, San's, 2026-08-08, and it is the line every field in `room.json` is now cut along:** > `room.json` records the room's **SHAPE** — workspace and roster — and restores it. > **AUTHORITY** — write posture, gate cadence — is never restored; it must be typed. The saved > posture field exists only for the reattach-mismatch notice, so it records the room **as it > stood** (live write + live asking, both sides at once so the notice can't fire spuriously). Two defects, one on each side of that line, and they are opposite failures of the same file. **Shape was half-saved.** `/cd` moves the room and the file follows; `/seat` moves the room and the file never heard. So `/seat claude,agy,cursor` — evicting a Codex that was dark on quota — died with a restart, and Codex walked back into the room and started billing the next unaddressed turn. That is the expensive-default defect returning through a reboot, on a control built specifically to kill it, and the room said nothing while it happened: the header drew four seats and the user had typed three. **Authority was half-recorded.** `savedPosture(m.st.Write, m.opts.Auto)` reads one live field and one launch flag. Press `a` in a gated write room and the file goes on saying `write-gated` about a room with nothing left asking. §9.17's own closing rule is the one that was broken — *a flag with an in-room twin stops being the answer to "what is the room doing" and becomes only the seed* — and §9.17 named this exact call site as a legitimate launch-time read. **This section amends that.** `savedPosture` is not a launch-time decision; it is a description of a running room, and it is the third miss of the same shape after `/write`'s confirmation card and the request path. #### The roster is keys, not content — which is why it may be saved at all ADR-008's ninth amendment ratified council writing exactly one file and ruled what may be in it: **keys, not content.** A roster passes that test rather than being excused from it. It is at most four vendor ids out of the closed set `addressableVendors()` — the same words `--vendor` takes on the command line and the footer prints on every frame. It says *who was in the room* and not one syllable of what was said in it. If the file leaked, the roster discloses which of four public CLI tools the user had on screen, which is strictly less than the workspace path sitting beside it already discloses. `TestTheSavedRoomHoldsKeysAndNeverContent` is the guard, and it fails closed by pinning the exact key set — so adding `seats` had to be a deliberate act that broke a test and got read. It now also asserts the roster's *content* is names, so a field added to `Seats` later that carried a note or a reason reaches this file through the same tag and gets caught there. #### Saved when it moves, not at the next dispatch `c`'s rule, in `clearSeat`'s own words: the room file is what a reattach reads, so a change held only in memory is undone by quitting — the user ends a thread and finds it waiting for them. A roster is the other thing a user deliberately takes out of the room, and it earns the same treatment for the same reason. **The save is an observation on `roomCommand`, not a call inside `seatCommand`.** `c` could put its `saveRoom` inside itself because there is exactly one way to clear a seat. The roster has `/seat`, has `/unseat`, and will have whatever narrows it next — and a save per command is a save the third one forgets. So `roomCommand` snapshots the roster, runs the command, and saves if it moved. Any command reachable from there inherits persistence without knowing the wrapper exists, which is what let `/unseat` be written in a parallel lane and compose with this without either side being told about the other. Two consequences worth stating rather than discovering: - **Only a change writes.** Bare `/seat`, a typo, and a `/seat` refused mid-turn all report without reseating. Rewriting the file on those would refresh `saved_at` — the age a reattach shows — for a room that answered a question and did nothing. - **A room that has never dispatched still writes nothing.** `saveRoom` returns at turn 0 and `readRoom` refuses a turn-0 file, both unchanged. A `/seat` typed before the first brief rides out on that brief's own save, which is the only save there was ever going to be. #### `--vendor` overrides the saved roster, and then rewrites it `--cd`'s rule and `--cd`'s reasoning: **an explicit launch control someone typed today outranks a file from yesterday.** `seatsFor` mirrors `Run`'s workspace switch line for line, down to sharing the same `re.Active() && !re.Offered` — a room `--fresh` declined restores neither half of the shape. The rewrite needs no code, and that is worth saying because it reads like a missing branch: `stateWith` copies the answer into `State.Seats` and `saveRoom` writes `m.st.Seats`, so the first completed turn records the room the user actually got. The same one line is what makes `/seat` persist. Leaving it out would be worse than not overriding at all — the file would go on describing a room that is not on screen, and the *next* launch would restore it. **Restoring is unconditional on the roster's own content**, including the zero value: the default room saved as the default room. A saved roster that could only ever *widen* would be a `/seat` you could not undo by quitting. **Back-compat is the absence of a field, and it is exact.** A `room.json` written before this section has no `seats` key; that decodes to the zero `Seats`; the zero `Seats` is the full detected table. So an old file opens the room it has always opened, and no version bump is owed — `roomVersion` is bumped when a field *changes meaning*, and additive fields are handled by the zero value, which is the rule `roomVersion`'s own comment already states. Pinned by a hand-written v2 fixture rather than a round-trip, since a file this build saved would carry the field and prove nothing. **An unknown seat name is dropped, not obeyed.** The roster is the one restored field whose value is a *name*, so it is the one a hand-edit or a downgrade can fill with a word this build has no seat for. Obeying it would seat nobody, fall through the everything-collapsed fallback in `VisibleColumns`, and hand the user the default room while the file claimed a narrowed one — §4a.1's collapse in the surface this section exists to make trustworthy. Dropped rather than refused, because a roster is shape: the sessions are still perfectly reattachable and refusing the whole file over the seating plan would cost four conversations to fix a screen. #### The posture field records the room, and still never restores it Both arguments are live now — `m.st.Write` and `m.st.Asking()` — and **the writer and every reader moved in one change**, because the field has exactly one consumer. A writer reading the state while a reader read the flags would compare a description of the live room against a description of the launch argv and report a change to a user who made none. That is the spurious fire the ruling names, and it would have been *introduced by the fix* had either side moved alone. Recording the room accurately is the opposite of restoring it, and nothing about the restore changed. `TestReattachRestoresNoPostureAndNoGate` extends the old `TestReattachDoesNotRestoreWritePosture` to the gate as well: a room saved `write` reopens read, a room saved with the gate off reopens asking, and the WRITE marker is asserted absent on the rendered frame rather than on a field. Both halves are witnessed, so it cannot pass by the room being read-only for some reason of its own — the same fixture reopened with `--write --auto` gets exactly what was typed. *A posture that can arrive from a file is not one anyone typed* survives this section unchanged; what it never said is that the room may not write down what it did. #### Declined **Restoring the gate.** It is authority, it is on the far side of the ruling's line, and `a`'s own section already argues that a safety property whose default is off is the wrong way round however carefully the constructor sets it. A gate that can arrive from a file is that mistake with a longer fuse. **Recording *why* the roster is what it is** — flag, command, or file. It is the room's own history rather than its shape, `saved_at` already dates it, and a `reason` string is the first thing in this file that would be prose. **Bumping `roomVersion`.** A bump costs every user their reattach, and it is reserved for a field that changes meaning. Nothing here changes what an existing key means. #### Amendment, 2026-08-16: the restore was correct, the fallback was silent, and neither had a seam A live 5/5 room reopened with the workspace cell reading `~` instead of the repo it was saved in (`STATE.md`, the 2026-08-15/16 drive). The roster came back and the workspace did not. `STATE.md` recorded the cause as undetermined: either an earlier session saved `~` honestly, or the restore dropped the field. This amendment answers that question with a measurement. **The restore does not drop the field.** `TestASavedWorkspaceIsRestored` plants a `room.json` whose saved workspace exists, opens the room, and asserts the workspace comes back. It passes against the *untouched* decision logic. The ordinary reopen was always correct, so the file itself held `~`. **The mechanism is the fallback that put it there.** `Run` verified the saved directory with `os.Stat` and replaced it with the current directory when that failed. It said nothing specific about the replacement. The reattach notice printed one sentence for two different events: *the room was in A; it is now in B* is what a `--cd` override prints, and a vanished workspace printed the same words. Nothing distinguished a directory the user moved to from a directory that no longer exists. **The fallback then persisted itself, and that is the real cost.** The room opens in the current directory. The next completed turn calls `saveRoom`, which writes `m.st.Workspace`. So one launch against a missing path overwrites the only record of where the room was. A renamed repo, an unmounted drive or a removed git worktree costs the saved workspace permanently, and the room never names the path it lost. That is `Reattachment.Offered`'s argument applied to the workspace: the destruction is silent and total, so the room must state it once. **Nobody could measure any of this, which is the other half of the finding.** The decision was a `switch` inside `Run`, and `Run` enters the alternate screen. No test could reach it. An untestable decision is one whose failures are all reported by the operator, which is exactly how this one arrived. **The fix.** `openWorkspace(opts, re)` is that decision as a function. It returns the directory AND the saved path it refused. `Run` carries the refused path on `Reattachment.WorkspaceGone`, and `reattach` gives it its own sentence: *the room was in ~/code/x, which no longer exists — it opened in ~/code/y instead*. The `--cd` sentence stays for the case it actually describes. - **One stat, one answer.** `openWorkspace` is the only place that stats the path. `reattach` reads the carried fact instead of statting again. Two reads a moment apart can disagree, and the room would then choose its workspace on one answer and describe it with the other. - **`WorkspaceGone` is never written to disk.** It is a fact about this launch, not about the room. `room.json` stays the keys and nothing else, per ADR-008's ninth amendment. - **A file sitting where the directory was is gone too.** `isDir` asks whether the path is a directory, not whether it exists. `os.Stat` succeeds on a file, and a room pointed at a file would dispatch four agents against it. - **A saved workspace is resolved rather than trusted as written.** `resolveWorkspace` makes it absolute. Every other consumer of a workspace in this package is handed an absolute path. - **The `--cd` refusal is unchanged.** A typed path that is not a directory stays a plain error before the alternate screen. The user named that path, so a silent substitution would act somewhere they did not ask for. `seatsFor` still mirrors this decision on `re.Active() && !re.Offered`. The mirror moved from a `switch` in `Run` into `openWorkspace`, and the shared condition is unchanged. **Declined, and named rather than left implicit.** The room still writes the fallback over the saved workspace at the next completed turn. Preserving the old path would make `room.json` describe a room nobody is in, which this section already refuses for the roster and refuses here for the same reason. The notice is the answer chosen instead. #### Amendment, 2026-08-16: the save choke point observed half the shape This section opens by saying `/cd` moves the room and the file follows. That was true of the FIELD and not of the WRITE. `saveRoom` reads `m.st.Workspace`, so whatever save came next recorded the move — but `/cd` made no save of its own, and `roomCommand`'s choke point compared only the roster (`sameSeats`). So the move reached disk at the next completed turn, or at teardown, or never. **Never is the case that matters, and the per-turn save already named it.** `endTurn` writes rather than leaving it to the way out, in its own words, because the failure it exists to survive is the room not getting a clean exit: a crash, a closed terminal, a machine that went down. A `/cd` had exactly that hole. It survived a quit, and a room that crashed after a `/cd` and before its next completed turn reopened in the directory the user had moved out of. The workspace is the field beside the session ids in the same file, on the same half of the ruling's line, with none of the protection. **A `/cd` is a deliberate operator statement about where the room is.** That is the roster's argument — `c`'s argument in `clearSeat`'s words, a change held only in memory is undone by quitting — reaching the other half of SHAPE. It should write when it happens, not when something else happens to write. **The fix is one line at the choke point, and deliberately not a `saveRoom` inside `cdCommand`.** `roomCommand` now snapshots the workspace beside the roster and saves if EITHER moved. The wrapper exists precisely so a command inherits persistence without knowing it does, and putting the call inside `cdCommand` would have been the per-command save this section already rejected — correct for `/cd` and absent from whatever re-points the room next. - **Each half compares in its own terms.** `sameSeats` because a slice does not compare with `==`; `sameDir` because two spellings of one directory are one directory, case-folded on Windows. Comparing the workspace with `!=` would write on a `/cd` that changed nothing but the capitalisation. - **The refusal semantics are unchanged, and the observation is what keeps them.** `resolveCD` rejects an unknown path BEFORE the workspace is assigned, so a bad path is never a value the file could briefly hold — the refusal is not layered on top of a write, it is upstream of one. A `/cd` mid-turn, a `/cd` to the directory the room is already in, and bare `/cd` all return with the workspace untouched, so none of them writes. That is the roster's rule verbatim: rewriting the file on a command that answered a question and did nothing would refresh `saved_at`, the age a reattach shows. - **Turn 0 still writes nothing.** `saveRoom` returns before the first dispatch, so a `/cd` typed before the first brief rides out on that brief's own save, and a room opened in the wrong directory and quit still drops no file into `~/.telltale/council`. Stated as a test rather than left to be found. - **The write path is `saveRoom`'s, unchanged** — the same atomic temp-file-and-rename, the same best-effort failure stated in the footer. No new writer, no second serialization of the same file. **Measured against the crash rather than against the field.** `TestCdIsPersistedWhenItHappens` drives `/cd` through `roomCommand` and then reads `room.json` off disk with no teardown and no completed turn — the simulated crash. It fails on the pre-amendment code with *nothing was saved*, which is the defect in one line. ### 9.33 the cursor seat's per-turn cost, split at last — and the seam that was hidden from `--help` §9.8 gave the Claude seat one live process and measured what it bought. The obvious next question was whether the Cursor seat could have the same thing, and the standing instruction in `STATE.md` was to read a trace before optimising anything. This is that reading. **Version pinned first, because the last capture's lesson was that a rule is only as general as the capture it came from (§9.6c).** Everything below is `cursor-agent` **2026.08.04-aaa8809**, the bundle's own `--version`, on Windows 11. That is **not** the version the rest of this seat was measured against — `vendors/cursor.go` cites 2026.07.23-e383d2b throughout — and one of the findings is a direct consequence of the gap. **Instrument:** the vendor's own `node.exe` against `index.js`, argv identical to the seat's read posture, with every stdout line stamped against the moment of launch. Two trials per arm. The `result` event carries the vendor's own `duration_ms`, which is the cross-check: it agrees with `system/init` → `result` on every trial, so the split below is the vendor's arithmetic as much as this instrument's. #### What the 44 seconds actually decomposes into `STATE.md` already established that spawn is 13 ms and that `wait` is where the time goes, and said outright what it could not do: `wait` bundles the vendor's startup with the model's time-to-first-token, and nothing then in the room could separate them. Stamping raw lines separates them, because `system/init` lands *before* the model is called. Print mode, no `--resume` — trivial prompt, `reply with exactly: OK`: | trial | launch → `system/init` | `init` → `result` (vendor `duration_ms`) | `result` → exit | total | |---|---|---|---|---| | 1 | 5.666s | 5.779s | 2.299s | 13.742s | | 2 | 5.617s | 5.361s | 1.818s | 12.792s | Print mode, `--resume` against a real prior session (created by the trials above, so nothing of anyone's real work is in this record): | trial | launch → `system/init` | `init` → `result` | `result` → exit | total | |---|---|---|---|---| | 1 | 5.196s | 5.042s | 3.104s | 13.337s | | 2 | 5.551s | 4.298s | 3.080s | 12.928s | And the startup itself, taken apart with progressively less work asked of the same bundle: | what ran | to first output | |---|---| | `node.exe -e "console.log('x')"` — interpreter only | 0.078s | | `node.exe index.js --version` — interpreter + bundle load + arg parse | 1.204s | | a turn in an **untrusted** directory (aborts at the trust check, before any model call) | 2.139s | | a real turn, to `system/init` | ~5.6s | **Three things follow, and only the first was already known.** **The standing diagnosis was right, and `--resume` is not the expensive half.** "`--resume` restores context, not process warmth" is confirmed and now has a number against it: resumed startup (5.196s, 5.551s) is *no larger* than cold startup (5.666s, 5.617s). Restoring a conversation is free. What costs is the fixed startup underneath it, paid identically either way. **Process cost is ~8.1s per turn and none of it is the model.** ~5.6s before `system/init` plus ~2.5s after `result` — the process lingers after answering — against a model turn the vendor itself clocks at 4.3–5.8s. Of the ~5.6s startup, node is 0.08s and loading the bundle is ~1.13s; the remaining ~4.4s is the vendor resolving auth, config, trust and workspace, and it is the largest single item in the seat's budget. **The honest proportion, stated so the number is not oversold.** On these trivial prompts the 8.1s is ~60% of the turn, but a trivial prompt is the arm that flatters the finding most. Against the real room traced in `STATE.md`, where `cursor` totalled 25.014s, the same fixed 8.1s is ~32%. The *absolute* figure is what is load-bearing: it does not shrink as the question gets harder, and it is paid again on every single turn. #### The seam: what print mode cannot do, and what the hidden subcommand can Persistence needs two halves. The output half the seat already has — `--output-format stream-json` is what §9.6c parses. The input half is the one that decides it: a way to hand turn N+1 to a process that is already running. **Print mode cannot be that channel, and the measurement is unambiguous.** Turn one was written to an open stdin and then the pipe was *held*. Nothing happened for sixty seconds. Only when stdin was closed did `system/init` appear, 3.6s later, and the turn ran — one turn, on the joined contents of stdin, then exit. **Print mode drains stdin to EOF and treats the whole of it as one prompt, so the EOF that starts the turn is the same EOF that destroys the channel for the next one.** There is no `--input-format` in `--help`, and none in the bundle either: enumerating every flag the bundle defines turns up hidden development flags (`--ian-dev`, `--sb-debug`, `--tool-gallery`), which is what makes that absence evidence rather than an unsearched corner. **One correction to this repo's own record falls out of the same probe.** `vendors/cursor.go` said no code path in the bundle reads the prompt from stdin, and that there is no `-` sentinel and no `--prompt-file`. That was true when it was measured; at 2026.08.04-aaa8809 the first clause is **false** — a prompt piped in with no positional argument produced a normal turn. Nothing in the seat depends on it (council always passes the prompt in argv), so this changed no code; the comment is corrected because a stale measurement left standing is how the next reader inherits a wrong premise. **The channel exists, and `--help` does not mention it.** The bundle registers a subcommand marked hidden: ``` Ce.command("acp",{hidden:!0}).description("Start the Cursor Agent as an ACP (Agent Client Protocol) server") ``` This is the `--permission-prompt-tool stdio` situation from §9.8 exactly — absent from the help text and real — so it was driven live rather than believed. **Two turns, one process, one session:** | trial | `initialize` | `session/new` | turn 1 | turn 2 | |---|---|---|---|---| | 1 | 1.944s | +0.994s | 5.285s | 5.335s (this turn ran a tool call) | | 2 | 1.701s | +1.040s | 5.365s | **1.177s** | The shape, recorded rather than the content: JSON-RPC 2.0, newline-delimited, on stdin/stdout. `initialize` returns `agentCapabilities` — including `loadSession: true`, the resume equivalent. `session/new` takes a `cwd` and returns a `sessionId` plus `configOptions`, among them a `mode` select whose values are `agent`, `plan` and `ask`. Turns are `session/prompt` requests carrying that `sessionId`; output arrives as `session/update` notifications (`agent_message_chunk`, `agent_thought_chunk`, `tool_call`, `tool_call_update`) and the request resolves with a `stopReason`. The second turn correctly answered a question about the first, from the same pid, so this is one conversation in one process and not two conversations that happened to share a parent. **The prize, stated as measured:** a follow-up turn costs **1.18s** where a print-mode turn costs ~13s, because the ~8.1s of process cost is paid once at `initialize` and never again. #### What this section does NOT authorise, and why it stops here The gate this work was run against was "build persistence only if the cost is process warmth *and* a live-verified seam exists." Both are now true, so the finding is **build**, and it is worth building. What the measurement also established is that the build is **not** the change it was expected to be — mirroring §9.8's shape onto this seat — because ACP is a *different protocol*, not the same protocol with an open stdin. Three forks come out of that, each a design decision rather than a detail, and each one is recorded here instead of guessed at: - **`Persistent` as written cannot express ACP.** `Turn(prompt) ([]byte, error)` is stateless: it returns the line for a turn. ACP needs `initialize`, then `session/new`, then a `sessionId` captured out of a *response* before any turn can be encoded at all — and `runner.Session` pipes lines and correlates nothing. Server→client requests (ACP's `session/request_permission`, the natural home for §9.8's gate) have no channel back at all today. That is a change to shared runner plumbing, not to one adapter. - **Posture and cwd stop being argv-bound, which un-founds the respawn rules.** `persistent.go` respawns a seat when the room moves or a `/flow` hop needs a different posture, and the comments there rest on both being fixed at spawn. In ACP, `cwd` is a `session/new` parameter and `mode` is a session `configOption` — so a `/cd` could open a new *session* in the same live process, and a posture change might not need a respawn either. Whether it *should* is a product question about what the badge is allowed to promise, not a mechanical one. - **Every measured claim on this seat was measured against print mode.** The §9.6c dedup rule, the `tool_call` oneof wire shape, the `--mode plan` badge and the Windows sandbox finding are all facts about a surface ACP does not use. Worth noting precisely because it is *not* yet a finding: across these two ACP turns, `agent_message_chunk` arrived once per turn with **no whole-message repeat** — which would mean the dedup rule is unnecessary here. That is a two-turn capture with one tool call in it, and §9.6c is the standing warning against generalising exactly that. It is a hypothesis for whoever builds this, not a rule. So the seat keeps its print-mode invocation for now, and the next lane starts with a number, a verified seam, and three named decisions instead of a guess. #### 2026-08-15: the same rig, pointed at the codex seat, and what a warm thread saves The rig above measured one vendor. This block runs it against a second one, and answers a question `STATE.md`'s 2026-08-08 trace could not. That trace shows the codex seat paying `wait=3.688s` before its first byte, while a cold binary start measures 190ms. Nothing said where the other ~3.5s went. **This is measurement only. It authorises no seat change.** **Version pinned first, and the subcommands were driven before they were believed.** Everything below is `codex-cli 0.147.0` on Windows 11 (`codex --version`). That is a NEWER build than the one `vendors/codex.go` cites. `codex app-server` and `codex app-server generate-json-schema` both exist on this build and both ran: the schema command wrote 46 files to a directory, and the server answered a live `initialize`. A subcommand named in `--help` is not evidence of a subcommand that runs, which is this repo's twice-earned lesson, so both were executed rather than read. **Instrument:** the installed `codex` binary, argv identical to the seat's first turn in `vendors/codex.go` (`-s danger-full-access --skip-git-repo-check --cd -`), with every stdout line stamped against the moment of launch. The prompt is **brief-shaped**: it opens with `brief.go`'s own `--- operating context ...---` fence and carries the request under it, because a greeting-shaped probe measures a transport the product never uses. One trial per arm, which is half of what §9.33 spent. Treat every figure below as one observation. #### The three-way capture, one identical turn `codex exec` (human), one trial. The seat does not use this renderer; it is here because it is the only arm that shows what the `--json` arm drops. | stamped line | at | |---|---| | spawn returned | 0.030s | | banner (`OpenAI Codex v0.147.0`) | 0.928s | | `hook: SessionStart` | 3.709s | | `hook: SessionStart Completed` | 4.201s | | the model's answer (`OK`) | 6.800s | | `tokens used` / `13,543` | 9.036s | | process exit | 15.519s | `codex exec --json`, one trial. This is the seat's own invocation. | stamped line | at | |---|---| | spawn returned | 0.014s | | `{"type":"thread.started",...}` | 1.153s | | `{"type":"turn.started"}` | 1.554s | | `{"type":"item.completed",...,"text":"OK"}` | 5.250s | | `{"type":"turn.completed","usage":{...}}` | 5.327s | | process exit | 13.266s | `codex app-server`, one process, one thread, two turns. The fixed half is paid once: | stamped line | at | |---|---| | spawn returned | 0.016s | | `initialize` response | 0.246s | | `thread/start` response, and the `thread/started` notification | 0.572s / 0.573s | Then the two turns, both on that one open thread: | turn | `turn/start` sent | `turn/started` | first `item/agentMessage/delta` | `turn/completed` | wait | stream | total | |---|---|---|---|---|---|---|---| | 1 | 0.576s | 0.836s | 5.085s | 5.259s | **4.509s** | 0.174s | 4.683s | | 2 | 5.263s | 5.303s | 6.467s | 6.705s | **1.204s** | 0.238s | **1.442s** | **A warm turn costs 1.44s, and 1.20s of that is the model.** Turn 2 asked what turn 1 had answered and got it right from the same pid, so this is one conversation in one process. Against the same prompt through `codex exec --json` the comparison is 1.442s against 5.327s to the last line, or against 13.266s to exit. **Four separate items make up the difference, and only one of them is process start.** 1. **Process and thread start is 0.573s, not 3.5s.** `initialize` answers in 246ms and `thread/start` in a further 326ms. Spawn itself is 16ms, which agrees with the 190ms class of figure and confirms again that spawning was never the cost. 2. **A `sessionStart` hook runs before the model does, and it is the operator's own.** `hook/started` at 3.151s and `hook/completed` at 3.840s, and the notification names its source: `"sourcePath":"C:\\Users\\sanle\\.codex\\hooks.json"`, `"source":"user"`, `"durationMs":838`. The human arm shows the same hook as `hook: SessionStart` at 3.709s. **This item is machine-specific.** A box with no `hooks.json` would not pay it, so it must never be quoted as a property of the vendor. 3. **Five MCP servers start on the same path.** `mcpServer/startupStatus/updated` fires for `node_repl`, `context7`, `github`, `kb-agent` and `codex_apps`, and two of them go `starting` to `cancelled` to `ready` across the turn. This is also operator config, and the same caution applies. 4. **The process lingers after it answers.** `exec --json` printed its last line at 5.327s and exited at 13.266s, which is **7.94s** of linger. The human arm shows 6.48s of the same. §9.36's "kill, never wait" rule was written for a different vendor and a ~2.5s linger. **This vendor's linger is larger, and nothing here checked whether council waits on it.** **One comparison this block does NOT make.** The `exec --json` arm reached its first line at 1.153s, well below `STATE.md`'s `wait=3.688s`. That trace ran a real brief through a real room with four seats starting at once, and this trial ran a trivial prompt alone. The 3.688s is not reproduced here and must not be treated as refuted. #### The hook question, answered by the captures **`codex exec --json` is the only one of the three surfaces that hides hook activity.** The human renderer prints `hook: SessionStart` and `hook: SessionStart Completed`. The protocol emits `hook/started` and `hook/completed`, each carrying the hook's id, event name, source path, source and `durationMs`. The `--json` stream emitted **four lines in total** for the whole turn (`thread.started`, `turn.started`, `item.completed`, `turn.completed`) and not one of them mentions a hook, an MCP server, or a rate limit. The protocol also carries two things the seat currently reads off disk instead: ``` {"method":"thread/tokenUsage/updated","params":{...,"tokenUsage":{"total":{"totalTokens":21130,...},"modelContextWindow":258400}}} {"method":"account/rateLimits/updated","params":{"rateLimits":{"limitId":"codex","primary":{"usedPercent":2,"windowDurationMins":10080,"resetsAt":1787369304},"secondary":null,"planType":"plus",...}}} ``` Those are the same fields §3.4 verified in the rollout files, arriving live on a socket. **Nothing is built on that here.** #### Seat-move viability, recorded and not acted on 1. **Thread continuity exists.** `thread/started` arrived live at 0.573s. `thread/resume` is a real request and its schema documents three routes (`thread_id`, `history`, `path`), plus `thread/fork`. The `thread/start` result carries `thread.id`, an identical `sessionId`, and the rollout `path` under `~/.codex/sessions/...`, which is the same id the adapter already reads. 2. **The sandbox channel exists on this path, and it is wider than `-s`.** `thread/start` takes a `sandbox` parameter, and the live response echoed `"sandbox":{"type":"dangerFullAccess"}`. `turn/start` takes a per-turn `sandboxPolicy`. That is strictly more than `codex exec` offers, where `-s` is first-turn-only and `codex exec resume` rejects it outright. Only `danger-full-access` was driven here. 3. **The Windows `danger-full-access` finding does NOT carry over, and needs its own re-check.** The evidence is direct rather than inferential: this protocol has a Windows sandbox surface that `codex exec` has no equivalent for. `windowsSandbox/setupStart` and `windowsSandbox/readiness` are client requests, and `windowsSandbox/setupCompleted` and `windows/worldWritableWarning` are server notifications. `vendors/codex.go`'s finding rests on `-s read-only` failing every process spawn, and `read-only` was never sent on this path. Until somebody sends it, the seat's badge rule stands unchanged. **Spend:** four billed turns. Two arms of one turn each, plus two turns on the app-server thread, because a single app-server turn would have reported turn 1's 4.509s as if it were the warm number and oversold nothing or undersold everything depending on which row was quoted. #### 2026-08-16: the linger gets an owner — what rides the tail, and what the column says now The block above ended with item 4 and a named gap: *"this vendor's linger is larger, and nothing here checked whether council waits on it."* It did wait on it. This is the check, the answer, and the fix. **The room's side of it, read from source first.** `vendors/codex.go` parsed `turn.completed` into a bare `KindMeta`, which `applyEvents` uses to adopt a session id and record a cost — codex sends neither on that line. Nothing else consumed it. A spawn-per-turn seat retired only on `KindDone`, the process exit. So the answer-complete marker was arriving, being parsed, and changing nothing: the column held `streaming` from the last token until the process died, and the elapsed it kept was the process's lifetime rather than the answer's. **Re-measured, because §9.33 spent one trial on this and the number is the whole argument.** Same build as §9.33 (`codex-cli 0.147.0`, Windows 11), so nothing here is confounded by a version change. Instrument: the installed binary, `codex exec --json` with the seat's own argv shape, a brief-shaped prompt behind `brief.go`'s fence, in a throwaway directory outside any repo. **`-s read-only`, not the seat's Windows `danger-full-access`** — a probe does not need write access, and the prompt tells the model to use no tools, so §9.33's Windows spawn refusal is never reached. It ran clean twice, which is a small measured aside worth keeping: `-s read-only` on Windows breaks codex's *child* spawns, not codex itself. | | trial 1 | trial 2 | |---|---|---| | `thread.started` | 1.914s | 0.537s | | `item.completed` (the answer) | 6.417s | 4.405s | | **`turn.completed`** | **6.619s** | **4.499s** | | last rollout write | 6.670s | 4.723s | | stdout/stderr close | 10.869s | 8.554s | | process exit | 10.870s | 8.555s | | **linger** | **4.251s** | **4.056s** | **Nothing rides the tail, and that is the finding this change rests on.** The concern was that receipts or vendor-side bookkeeping might be paid after the answer, which would make an early settle a lie about a turn still in progress. It is not: on both trials `turn.completed` is the LAST line on stdout, stderr stays empty throughout, and the vendor's own rollout file under `~/.codex/sessions/` takes its final write 51ms and 224ms after that line — roughly four seconds *before* the exit. The rollout was polled at 250ms rather than diffed before-and-after, so this is an observation of when writes stopped, not an inference from a file that changed. The linger itself is 4.06s and 4.25s here against §9.33's 7.94s on the same build. **The size is not stable and must not be quoted as a constant**; the shape is, and the shape is what the render has to survive. **What changed, and the line it does not cross.** `turn.completed` now carries `EndsTurn`, and `applyEvents` grew a third branch for a spawn-per-turn seat that names its own end of turn — a case that did not exist before, because the batch CLIs all ended a turn by dying. That branch **settles the column without retiring it**: phase to `done`, elapsed stamped at the answer, body completed, and the seat stays in `m.turn.live`. The split is the whole design, and the reason is mechanical rather than tasteful. `turnColumnFinished` cancels the turn's context; `runner.Start` kills the child on that context. Retiring the column at the marker would therefore **kill the process four seconds early**, and that is refused even though the tail measured empty — both probe turns used no tools, so a turn that ran commands is unmeasured, and probing one needs `danger-full-access`, which a probe does not get. Shortening a vendor's life on evidence that does not cover the case is the inference-dressed-as-measurement move §4a.1 exists to refuse. The exit still arrives, still runs `KindDone`, and still retires the column. It just no longer decides what the column *says*. Two consequences fell out, and both are corrections rather than costs: - **The turn clock stops at the answer.** `clock.observe` already ended a turn on `EndsTurn`, and its comment said a spawn-per-turn child never sets it. That comment is now false and is fixed. The old reading billed the linger to the vendor's turn time, so the seat was timed on how long its process took to get out of the way. - **`Busy()` and the footer came apart, and had to.** `Busy()` is derived from column phases, so it goes false the moment the seat settles — correctly: nothing is working. But the turn is still live, and `q` is refused while a turn is live, so the footer's `Busy()` test would have advertised a key that answers with a notice (§7.8). The footer now asks `InFlight()` (`Busy() || Settling()`) instead. `Busy()`'s own doc claimed it drove the meaning of `ctrl+c`; it never did — that key reads `Model.turn` — and the stale claim is corrected rather than preserved. **The linger is rendered, not hidden.** Between the two moments the room would otherwise go completely quiet — no spinner, every column reading `done` — with the composer still locked. That is a room that looks wedged for a different reason than before, which is not a fix. So a settled seat renders `done 4s exiting`: a WORD, after the clock so it cannot be read as part of the figure, and a word rather than a glyph or a colour because what it prevents is a reader concluding the room is stuck, and that reader may be on `--ascii` or `NO_COLOR` (§7.1 rule 2). **Two defects the review caught in this change, both now guarded.** A killed process drains its buffered stdout, so a `turn.completed` can arrive after the column it belongs to is already terminal — `giveUpSeat` kills an arena racer and retires its column as cancelled, and the queued line lands behind it. The phase write was guarded from the start; the BODY write was not, so a cancelled seat's note-bearing body was replaced with `[Turn completed with 0 text chunks streamed]` — a cancelled column asserting that its turn completed. Every write on the branch now sits behind one guard, and the guard is wider than a phase test: it also checks `m.turn.live`, which covers the same line arriving after the turn boundary entirely, where it could otherwise settle a *fresh* turn's column on the strength of the previous turn's answer. Second, the vendor-reported failure path restamped `Elapsed` unconditionally, so a seat that answered and then exited badly recorded the process's whole lifetime — the exact figure this block exists to stop billing. It now fills only a zero, which is the rule `finishColumn` already followed. **What is NOT claimed.** The linger's cause is still unknown — this block measures when it starts and ends and what does not happen during it, and nothing about why the vendor holds the process open. `codex app-server` (§9.33's third arm) has no linger at all because the process is meant to stay up, so the seat-move case gains an argument here; it is recorded, not acted on, exactly as §9.33 left it. **Spend:** two billed turns, both trivial prompts on `read-only`. #### 2026-08-16: the gap that block named, on the seat that fails inside its stream The block above closed the answer case and left one gap open, written into `InFlight`'s own doc comment: **a seat that FAILS in its stream was terminal inside a live turn.** This closes it. No vendor was run for this change. The whole finding is a source read, and the section says so rather than borrowing the authority of the measurement above it. **What the source says.** `vendors/agy.go` parses a `result` line carrying `status: "ERROR"` into a `KindError`. It fills `Note` from the vendor's own `error` field. It sets no exit code and no error, because nothing has failed at the process level and nothing has exited yet. It does not set `EndsTurn`; no agy event does. `applyEvents` therefore reached its `KindError` tail, wrote `PhaseFailed`, and retired the column only on `ev.ExitCode != 0 || ev.Err != nil`. This event carries neither. **One evidence boundary, because the phrase is easy to over-read.** "Exit code 0" here is the EVENT's field, read off the adapter. It is not a measurement of what the `agy` process exits with after a failed turn, and no such capture exists — `vendors/testdata/wire/README.md` records that the one probe that could have produced an agy error frame does not produce one, because that CLI answers a lost thread with success. Nothing here depends on the process's exit status. `KindDone` and a failing `KindError` both retire the column, so either exit ends the turn. **The room that produced.** The column read `failed`. The turn stayed live, because nothing retired it. So the seat was neither `Busy()` nor `Settling()`, the spinner stopped, every column on screen read terminal, and `q` was still refused with *"a turn is in flight"* — which was true. The footer offers `q` on `InFlight()`, so the room named a key that answers with a notice. That is §7.8's surprise, and it is the same defect the block above fixed for the seat that answers early, reached through the other terminal phase. **The fix is that block's split, applied to `failed`.** The column settles instead of retiring: phase to `failed`, the vendor's sentence kept, `Settling` set, and the vendor left in `m.turn.live`. The exit still arrives, still retires the column, and still clears the word. The seat renders `failed 5s exiting`, which is the same honest sentence `done 5s exiting` makes about a different outcome. **Retiring the column was the alternative, and it is refused for §9.33's reason rather than for a new one.** `turnColumnFinished` cancels the turn's context, and `runner.Start` kills the child on that context. Retiring here would kill a process that is still winding down. Codex's linger was measured before that argument was accepted; this vendor's linger was not measured at all when this block was written, which made the case stronger rather than weaker. A room may not shorten a process's life on a number nobody has. > **Amended 2026-08-16 (§9.43).** That number now exists, for the SUCCEEDING turn only: agy's tail > is 0.049s, 0.135s and 0.314s on `agy 1.1.13`, against codex's 4.06s and 4.25s. The settle above > stands unchanged. The measurement makes the ruling cheaper rather than wrong — there is almost > nothing left to cut short — and it does **not** cover this branch, because all three trials > succeeded. The failing turn's tail is still a number nobody has. **Moving the fact onto `State` was the second alternative, and it is refused too.** `State` cannot see `Model.turn`, so "a turn is live" would have to be copied onto it and reset on every path that ends a turn. `Settling`'s own doc comment already rules on that shape: a second home for a fact is a second thing to forget to reset. `Settling` IS the state-side fact this needed, and the failure path was simply not setting it. **The guard is the same one review found for the answer case.** A killed process drains its buffered stdout, so this line can land on a column that is already terminal, or after the turn boundary entirely. The settle is therefore behind a phase test AND `m.turn.live`, read before the phase is written. Without it a late failure line would hold `InFlight` true with nothing running — the room wedged the other way, where the footer never offers `q` again. `failedturn_test.go` pins all three states: the settle, the retirement on the exit, and the late line that must change nothing. **Scope, stated because the branch is vendor-neutral and the case is not.** Any spawn-per-turn seat whose adapter reports a turn failure in-stream reaches this branch. agy is the only seat that does so today: codex and grok have no structured error frame, so their failure IS the exit and the `ExitCode != 0` leg already retires them; Claude runs persistent; the Cursor seat is ACP. The fix is written on the shape rather than on the vendor id, and a seat that adopts the same shape later is covered without an amendment. *(Amended 2026-09-01: codex adopted the shape. codex-cli 0.151.0 puts a `turn.failed` frame on stdout, and the adapter now parses it into exactly this event. The amendment is not free: a spawn-per-turn failure now produces TWO events, the vendor's sentence and then the exit, and the exit must not overwrite the sentence. [§9.58](#s9-58) records that rule.)* ### 9.34 the rebuttal stopped naming its authors A `ctrl+r` turn used to quote each seat's answer under its vendor's name: *"quoted reply from Codex"*. The fence's security framing was right and is untouched (quote.go); what was wrong is subtler — **the receiving model was told who wrote what**, and models weigh an argument differently when it arrives under a name they recognise. That is the self-preference / identity-bias class that peer-review setups blind for, and the reason llm-council anonymises its review stage before models rank each other. A rebuttal exists to test the argument, so the argument is now what crosses: *"quoted reply from participant A"*. The mechanics, because each carries a decision: - **Labels are positional per receiver** — seat order with self skipped — and a seat with nothing to quote this turn keeps its letter reserved, so a quiet seat does not shuffle every neighbour's identity. The letters exist so a multi-turn argument stays attached to a consistent speaker; letters that agreed BETWEEN receivers would need a shared assignment written somewhere the models could correlate, and nothing downstream may join on them anyway. - **The blinding is label-deep, and says so.** A reply whose content self-identifies ("as Claude Code, I…") has identified itself, and editing another participant's words to hide it would be the censorship the fence refuses — the room shows what was said. Best-effort blinding, stated as such, over silent redaction. - **The user is not blinded.** Columns stay labelled by vendor; the blind applies to what the models read, never to what the person sees. Nothing in the room's rendering changed at all. - **One test inverted, on purpose.** `TestQuotedMaterialIsFencedAsUntrusted` used to fail with "quoted material is not attributed to its author"; attribution to the model is now the defect. The name-absence checks in the nothing-quotable test were re-grounded on the fence itself at the same time, because a vendor-name check passes vacuously against a prompt that never contains vendor names — a guard that cannot fail is not a guard (the same one-level-up rule the architecture repo's test policy states). What this deliberately does not add: a ranking stage, a chairman, or any synthesis hop. llm-council's stage 3 collapses the answers into one; §9.2's position is that independent answers ARE the product, and a synthesis is available today as an explicit `/flow` hop the user types. Blinding sharpens the comparison; it does not delegate the verdict. ### 9.35 a running chain can be told to stop after this hop `/flow` shipped with one control over a chain already moving: ctrl+c, the turn-cancel key. That is a hard abort — it interrupts the hop mid-sentence — and it was also a lie in three parts, which is where this section starts, because the honest baseline had to be measured before a gentler control could be designed against it (the flow-autoadvance plan named this gap and nothing tracked it). **What cancelling mid-chain actually did, measured 2026-08-08.** Ctrl+c during hop 1 of 3 killed the hop and the chain never advanced — that much was right, and stays. Everything around it was wrong: the current step stayed `running` forever, the header went on claiming `hop 1/3` over a room doing nothing, and the next brief the user typed was **eaten** — dispatch saw a live chain, tried to continue it, failed with `flow start error: cannot start step in state running`, and threw the user's words away. The second enter worked. And the corpse was not the cancel path's alone: **a chain that COMPLETED left the same corpse**, and the first brief after a successful flow died as `cannot start step in state returned`. The happy path was charging the same tax. **The fix under everything else: a chain ends whole, or it has not ended.** `endFlowChain` retires the chain, its draft, its carry, its pending flags and its header marker in one move, and every ending — finished, failed, cancelled, stopped, refused at the write gate, start error — goes through it. Half of these paths used to clear only the marker and half only the chain; each half-state was a room asserting something false about the other half. The teardown in `turnColumnFinished` is the backstop for the endings `finishFlowHop` never sees (a cancelled hop, a vendor failure, a hop that returned nothing), because teardown is the one place every turn's death already passes through — and it says which death it was: `flow cancelled at hop 1/3 — 2 later hops not run` against `flow stopped at hop 1/2 — 1 later hop not run`. Cancelled and failed are different facts, and a stopped chain must never render as a finished one (§4a.1). **The control: `s`, stop after the hop that is running now.** The middle ground ctrl+c cannot offer: the current hop finishes on its own terms — artifact saved, receipt verified, Returned or Published exactly as if nothing had been pressed — and the chain ends there instead of handing off. Pressed while the hop streams, because that is when the decision forms: you are reading hop 2's output when you learn hops 3 and 4 are no longer worth their quota. Reading the columns is part of deciding — §9.16's own argument for its gate — so the key lives in view mode where the columns stay scrollable, and interrupts nothing. The grammar decisions, each against a precedent: - **A key, not a room command.** `c`'s reason (§9.17): a key takes no vocabulary from the composer, and `/stop` is a word people address vendors with. Not `c`'s confirmation though — `c` spends a `y` because a dropped thread is irreversible, and this destroys nothing: the hop completes, every artifact lands, and the cost of a stray press is one keystroke to undo. So `s` is a toggle, `a`'s shape, and pressing it again re-arms the handoff. - **The armed state lives on the chrome, not in the notice.** The WRITE badge's argument: a notice scrolls away and the promise persists. The hop cell reads `hop 2/4 @codex (stops here)` — words, so it survives `--ascii`, and on the hop cell because it is a fact about the hop: this one runs, its successors do not. The busy mode line offers `s stop after hop` while a chain is live and flips to `s continue chain` when armed, the label-renames-itself rule `a` set. `TestFlowStopIsAToggleAndTheArmedStateRenders` pins all four states. - **The last hop refuses to arm.** The chain ends there whether or not `s` is pressed; a key that "worked" would claim credit for an outcome it did not cause. The refusal says so — `hop 2/2 is the last — the chain ends here anyway` — and a press with no chain running is answered rather than swallowed, §9.12's attribution rule for a key that did nothing. - **Not in helpKeys, deliberately.** The key exists only while a chain runs, and the mode line names it on every frame of exactly those moments — the contextual-control surface `y`/`n` and the card's `a` already use. The help panel's 17-row budget has no row for a key that is dead in every room the panel is usually read in; the §9.17 sweep's rule was about controls reachable from inside the room, and this one is announced there. **What was declined.** Pausing — a stopped chain that could resume where it left off — is a real feature and not this one: the carry artifact is consumed at dispatch, the draft would need re-parsing against a chain whose position moved, and stop-then-retype is the honest v1. A stop that also killed the current hop was declined because ctrl+c already is that, and two keys with one meaning apart is how a keymap grows synonyms. And ctrl+c stays exactly as hard as it was: the interrupt semantics did not move, only the lying state it left behind. The general lesson, in this file's own terms: §9.16 built the chain's authority grammar and §9.17 built the room's mid-session controls, and the seam between them — a chain that is neither obeying nor gone — belonged to nobody, so nothing asserted on it. The corpse survived every ending, including the successful one, because the tests all stopped at "did not advance" and none typed the next brief. The regression tests here end by dispatching one. ### 9.36 the cursor seat re-founded on ACP: what the wider capture said, and what it cost §9.33 ended with a build verdict, a verified seam and three named decisions. This is the build. The ruling that shaped it was **wholesale**: ACP replaces the spawn-per-turn path rather than sitting beside it, there is no fallback, and git history is the record of what went. A fallback would be a second protocol to keep honest, and the numbers below are the reason nobody would want to fall back to it. **Version pinned, and it is the same one.** `cursor-agent` **2026.08.04-aaa8809**, the version §9.33 measured, on Windows 11 — so nothing here is confounded by a bundle change. Instrument: a throwaway JSON-RPC client driving `node.exe index.js acp`, every line stamped against launch, across **thirteen arms**. Then the finished seat re-verified through the room's own code. **One environment note, because it cost an arm and will cost the next reader's.** The first capture had every tool call blocked by this machine's own `PreToolUse` credential guard, whose wrapper fails closed when cursor-agent is launched from a Git Bash parent on Windows (a known upstream wrapper bug, agent-ops ADR-012). Nothing was wrong with ACP. **Drive cursor-agent from a PowerShell or cmd parent on Windows**, or every tool in the capture will read as failing. #### Phase 1: what two trivial turns could not have told us §9.6c's lesson is that a rule is only as general as the capture it came from, and §9.33 flagged its own two-turn no-repeat finding as a hypothesis for exactly that reason. So the capture was widened first, and the widening changed three conclusions. | what was asked | trials | what came back | |---|---|---| | a turn that runs a TOOL | 3 | `tool_call` then `tool_call_update`(in_progress) then `tool_call_update`(completed). `title` always populated; `rawInput` **empty** for Read/Find/grep and populated for shell | | several model calls in one turn | 2 | four tool calls and three message segments in one turn, interleaved, with no envelope around a "call" | | a long streamed reply | 1 (300 words) | 24 `agent_message_chunk` in 2.6 s — ~95 chars each, ~9 a second | | an interrupted turn | 1 | `session/cancel` (a notification) and the open `session/prompt` resolves `{"stopReason":"cancelled"}` **23 ms** later; the process took a further turn 1.1 s after that | | resume in a NEW process | 2 | `session/load` works, and **replays the whole prior conversation** onto the update stream before it answers | | a dead thread | 2 | `-32602 … Session "…" not found` in **0.45 s**, and the process survives — a fresh `session/new` answered 0.45 s later | | cwd binding | 1 | **per SESSION, not per process.** One server ran a session in `ws1` reading ws1's file and another in `ws3` reading ws3's | | workspace trust | 2 dirs | **does not apply.** Print mode refused the same directory with "⚠ Workspace Trust Required"; the ACP server wrote a file into it | | a permission prompt | 2 | `session/request_permission` **blocks**; `allow-once` ran the command, `reject-once` did not and nothing was created | | an edit | 2 | ran ungated — **no permission request at all**, in a never-trusted directory, under the user's own `approvalMode: allowlist` | | plan mode | 1 | `session/set_mode {"modeId":"plan"}` accepted; asked to create a file the seat declined and **no file landed** | | the dedup hypothesis | every arm | **no whole-message repeat anywhere, and no `model_call_id` field in ACP traffic at all** | Timings across the twelve arms that ran a handshake: `initialize` **1.43–4.30 s**, `session/new` a further **0.85–2.55 s**, `session/load` a further **0.89–1.37 s** — so resume is once again no more expensive than a fresh conversation, which is the same shape §9.33 measured for `--resume`. Warm turns: **1.12 s, 1.79 s, 1.82 s**, against §9.33's print-mode ~13 s. **Three of these overturn something.** **The dedup rule is not carried over, and it is now a measurement rather than a hypothesis.** §9.6c's rule exists because print mode sent a model call's deltas and then that call's complete message, so appending both rendered the passage twice. Across a turn with four tool calls and several model segments, ACP repeated nothing and carries no `model_call_id` at all. §9.33 was right to refuse to generalise from two turns; the wider capture is what earns the conclusion. **The safety net that rule leaned on is gone with it.** §9.6c named the fallback explicitly — "the failure mode is a column that fills at the end, never one that is wrong" — because print mode's `result` carried the whole reply. An ACP turn resolves with `{"stopReason":…}` and nothing else: no reply, and no token usage either. So a broken chunk parser here gives an **empty** column, not a late one. There is no mitigation that would not be invented, so it is stated instead — in `cursor.go`, in `dispatch.go` where the old special case was, and in `STATE.md`. **Workspace trust does not apply on this path.** This is the one finding that makes a claim *worse*, and it is on the badge for that reason. The tightest form of it: the directory print mode had just refused was written to over ACP, with no prompt, minutes later. #### Phase 2: what was built **runner grew a second protocol shape, and the stream-json path did not move.** `ParseFunc` sees a line and has nowhere to reply to, which is enough for a monologue and cannot express ACP: a turn cannot be *encoded* until `session/new` answers, the child asks questions that block it, and ids come in two independent namespaces. So `runner.Protocol` is a stateful per-process driver that owns both directions and returns lines rather than writing them — which keeps it replay-testable exactly as a `ParseFunc` is. `StartSession` and `StartRPCSession` are two wrappers over one body; the Claude adapter was not touched and its tests did not change. `Session` grew `SendTurn` and `SendAside` in place of a bare `Send`, and the split is the turn clock: an ACP protocol may **take** a turn it cannot yet encode, and the person who pressed enter is waiting from that moment whether or not a byte has moved. An answer to a question the vendor asked mid-turn goes the other way — it belongs to the turn it is holding up, so it must not start a new one. **The seat.** `vendors.Conversational` is a sibling of `Persistent`, not a subtype: `Open` returns a spec plus the protocol. The invocation is now the single word `acp`. Posture arrives as `session/set_mode`, the workspace as `session/new`'s `cwd`, and the brief as a JSON string — so no prompt text can reach argv by any path, which retires the shell-shim question this seat used to have to reason about. **Re-measured, not inherited.** | claim | verdict | |---|---| | granularity `tokens` | **re-earned, with a caveat.** ACP chunks are ~95 chars at ~9/s — coarser than print mode's real tokens ("P", "ONG") and *finer in time* than the Claude seat that already carries this word (§9.7: ~80 chars, ~3/s, flagged there as an overstatement). The word stays with its existing looseness and no new looseness; fixing it is one change to both seats at once, which is why §9.7 left it as a separate change to a separate surface | | `ro:requested` | **level unchanged, reasons replaced.** Plan mode did better than print mode's ever did, and it is one trial of a mode the model obeys | | `--sandbox enabled` | **gone.** ACP takes no sandbox parameter on any OS, so the badge is no longer split by platform. On Windows nothing was lost — the flag was measured killing the turn. On macOS and Linux what was lost is a *request* whose enforcement was never observed | | `gated` | **withheld, deliberately.** `canGate` used to read "is this a live process", off the registry. That was right only while those two questions had one answer. This seat can ask *and does not ask about edits*, and `gated` promises that nothing which changes anything runs without a keystroke — so it keeps `WRITES`, and its detail says what the cards cover and what they do not | | the §9.6c dedup rule | **retired**, above | | the `result` fallback | **retired**, above | **The gate, such as it is.** Council answers every `session/request_permission` — an unanswered one blocks the vendor forever, which is a column that never finishes. In a write posture the request becomes the room's ordinary approval card; in a read posture the adapter refuses it itself and records the attempt in the trace, because a read-only seat asking to change something is not a question for the user, it is already answered. `allow-always` is never selected in any posture: it writes a permanent rule into the user's own `~/.cursor/cli-config.json`, which is the line this adapter already declines to cross by never passing `--trust`. Two shapes fall out of the capture and both are in the code beside the lines that produced them. A **rejected** call arrives as `completed` with no output at all — indistinguishable on the wire from a completion that said nothing — which is §9.8's `ActDenied` argument in a sharper form: the room records the refusal from its own keystroke and `recordAct` refuses to let the echo overwrite it. And ACP's rejection carries **no message field**, so unlike the Claude seat this one cannot ask the model not to retry; it was measured saying "DONE" afterwards as though nothing had happened. **`session/load` replays history, and dropping it is load-bearing.** A loaded session streams the entire prior conversation back — old prompts, old tool calls with their real output, old replies — *before* it answers. A parser that appended it would refill a reattached column with the whole previous room. The gate is the pending response rather than the `replay-` prefix those ids happen to carry: a prefix is a spelling, the pending request is the protocol. **Two hazards this protocol has that a one-way stream does not, both found in review and both ending in a room nobody can quit.** They are recorded because neither is visible from the wire format alone. - **A turn's end is a RESPONSE, so a turn that was never sent can never end.** Anything that holds a brief — the handshake, the `session/set_mode` round trip — is a window in which there is no outstanding `session/prompt` for the vendor to answer or for a cancel to abandon. So the protocol refuses an interrupt in that window rather than reporting a quiet success, which is what makes the room fall through to killing the seat; the alternative was a turn that never ends, a room that then refuses every further brief, and a `q` that will not quit. - **A failed handshake is TERMINAL, because the server does not exit on one.** An ACP server that refuses `initialize` answers and stays up — and a live process is exactly what §9.8's stale-exit guard correctly reads as a healthy seat. Without a terminal state the room would keep handing that process briefs forever. The protocol refuses instead, the seat is killed and forgotten, and one retry inside the same dispatch gives the user a working column rather than an error naming a handshake they cannot see. The likeliest trigger is an unauthenticated CLI: somebody's first run. The same class of care applies to the `set_mode` window in the other direction: a turn *taken* there must wait too, or it would go out under the server's default `agent` mode while the badge said `ro:requested` — invisibly, because a reply from the wrong mode looks exactly like a reply from the right one. **A refused thread now costs two round trips instead of a process.** The one-attempt rule the ninth amendment established is unchanged and is simply cheaper here: the id is spent, the same process opens a new conversation 0.45 s later, and the brief still runs. Reattachment therefore never fakes a restored thread — if the load is refused the seat honestly starts fresh and `settleRestoredThread` says so, exactly as it does for the other three seats. #### The forks, and the one that was decided rather than measured §9.33 named three. The first (Persistent cannot express ACP) and the third (every claim was measured against the wrong surface) are settled above. The second is a **choice**, and it is called out here because a reader would otherwise find a measurement in the code and wonder why it was not acted on: **`cwd` and posture are no longer argv-bound, and the seat is respawned anyway.** Measured: one process really did run two sessions in two directories. So a `/cd` *could* cost a new session (~1 s) instead of a new process (~3 s). It costs a process — because what a move actually costs the user is a new conversation either way; because one rule across four seats is worth more than three seconds; and because re-opening a session inside a live process has failure modes (a half-moved session, a queued turn addressed to the old one) that nothing has measured. The argument is on `seatProc`, the behaviour is pinned by `TestAMovedRoomReplacesTheCursorSeatToo`, and it is revisitable with a measurement rather than with a preference. The **stale-exit guard** (eleventh amendment) applies to this seat unchanged and is re-asserted for it: a terminal event names a vendor, not a process, and acting on a predecessor's exit would fail the live turn and leave a real process running and invisible. #### Verification Fixture replay in the #62 style over synthesized shapes (`vendors/testdata/cursor-acp-turn.jsonl`), lifecycle pinned by **process counts** rather than by anything the adapter says about itself, and one live multi-turn conversation through the merged seat (`-tags=live`): ``` turn 1 phase=done elapsed=9.744s act "Read File" → ok body: github.com/sanlee-ys/telltale turn 2 phase=done elapsed=1.120s same process body: github.com/sanlee-ys/telltale ``` Turn one read a file it could only have read by running a tool in the workspace; turn two answered a question only turn one's history could answer, from the same process, in 1.12 s. That is the whole of §9.33's prize, measured through the room rather than through an instrument standing beside it. **Not verified here: macOS.** Every arm ran on Windows 11, and the Mac's ACP seat is unmeasured. That belongs in `PARITY.md` rather than in this section. ### 9.37 /arena: the seats race in worktrees, and the human picks the winner `/arena ` is one brief raced across every seated vendor, each attempt in its own git worktree, compared by diff instead of by prose. It is §9.2's thesis — independent answers ARE the product — applied to code, where "independent" stops being free: four writers in one shared tree are not four answers, they are one trampled tree. The isolation the manager lane built its whole category on (Crystal's same-prompt sessions, claudexor's best-of-N envelopes, parallel-code's AI Arena) is what makes four *write* attempts comparable at all. Ruled 2026-08-08, four decisions and their reasons: - **Per-turn, typed at the room** — not a launch posture. §9.17's rule; a race is something you want *about a brief*, not about a session. - **Every attempt is a FRESH session.** All three comparable products race fresh, a continued thread would anchor each seat on its own prior answers, and whether resume even survives a cwd change is measured for none of the spawn-per-turn seats — so fresh is also the only option that costs no new vendor measurements. Mechanically: every seat goes through the `FirstTurn` one-shot it already implements, the persistent seat included. The room's live process, saved ids and conversations are untouched — dispatch guards the session-id capture so a race's throwaway ids can never replace the room's saved threads (the reattach-swap bug, killed in a test before it could exist). - **Worktrees are KEPT until the user deletes them**, named `-arena-t-` as SIBLINGS of the workspace with branches `arena/t/`. Siblings, not a state directory: kept-until-deleted means the user must SEE what is kept, it matches the README's own worktree convention, and /cd's sibling resolution makes `/cd repo-arena-t7-codex` work with zero new code. - **Comparison lands in-column** — `git diff --stat` against a base SHA recorded once before any seat spawned, rendered in the transcript's boundary grammar; `y` yanks the full diff (capped at 1 MB, truncation stated). Three outcomes, three renders: a diff, a measured "no changes against ", and "diff unavailable: " — zero, absent and degraded stay three different facts (§4a.1). Two mechanics carried in from the deep-read of claude-squad's `session/git/diff.go`, because they are the difference between a diff surface and a lying one: the diff anchors on the **recorded base SHA**, never HEAD, so an attempt that commits mid-turn cannot show an empty diff; and `git add -N .` runs before diffing so an attempt whose whole answer is a NEW file cannot read as "no changes" — the false zero, again. Posture is `PostureWrite` for every racing seat, stated rather than hidden: a one-shot process has no channel to be asked on, so the gate structurally cannot exist here, and the containment is the worktree — which is the whole reason the worktree exists. A read room refuses `/arena` with the in-room remedy named (`/write lets it`), per §9.17's tell. **What council deliberately does not do, having read the competition:** claudexor AUTO-ADOPTS the winning patch into the live tree. This room offers the diffs and the human picks — adoption is a git command the user runs against a kept branch, never an action taken for them. And a race is not routable in v1 (`@codex`-only arenas): the value is the comparison, and a one-seat race is an ordinary turn in a worktree, which `/cd` already provides. Deliberately deferred, each its own change judged against this section — and every one has now landed, each as its own change (the 2026-08-09 amendments below): ~~commit-per-turn inside arena worktrees~~ (with its undo), ~~`.worktreeinclude` seeding (the first real arena run on a repo needing `.env` will surface it)~~ (landed on exactly that argument, ahead of that repo showing up), and ~~a deletion guard stronger than git's own refusal to remove a dirty worktree~~ (landed as `/arena drop`'s counted refusals, beside `/adopt`). This paragraph briefly existed as two half-struck copies of itself — two same-day changes each struck their own item and a text merge kept both variants — collapsed back to one on the same day. Verification note: the git mechanics (worktree creation from one base, add -N, the three collection outcomes, the session-id guard, the renders, the yank) are all pinned by offline tests against a real temp repository. ~~No live vendor has raced yet.~~ **The first live race ran 2026-08-09** — turn 4 of a real room on the Windows box, four seats dispatched, `/arena` against this repo at 422b1c3 — and it paid the debt this note carried while measuring exactly where the predicted risk lived: - **The core is verified live.** Worktrees created as named siblings, three seats raced fresh, ranks rendered in host-observed order (agy 1st · 7s, codex 2nd · 15s, claude 3rd · 19s), and the zero-render said "no changes against 422b1c3." on every finisher — honest zeros, since the brief was a harness check that asked for no changes. The room's threads survived intact. - **The cursor seat cannot race, by its own design.** The ACP refounding (§9.36) gives `Cursor.FirstTurn` a deliberate refusal — "driven as a live ACP process, not as one child per turn" — which arena's uniform one-shot path duly surfaced on the column. The fix is a follow-up with its own shape: an EPHEMERAL ACP session per race (spawn in the worktree, one turn, kill), which is §9.36's machinery pointed at a throwaway session. ~~On the deferred list below until someone builds it; until then a race is honestly 3-of-4 on Windows.~~ **Built 2026-08-09, in exactly that shape — second amendment below.** - **The write seat hit the allowlist-prefix trap.** claude's one probe — `git -C status --short --branch` — met `autoAllowedTools`' `Bash(git status:*)` rule and failed the prefix match, so an ungated print-mode seat had an approval request and nobody to ask (act rendered ✗, correctly). The fix is NOT a blind `Bash(git -C:*)` — that constant also serves `--auto` in real workspaces, where pre-approving every `-C` form is a wider grant than the verbs it lists — and the vendor file already warns its rule grammar has not been driven. ~~What this needs first is one measured probe of whether the rule syntax can scope a verb behind `-C` at all; the finding is filed, the measurement is the next step.~~ **The probe ran, 2026-08-09, and closed this.** A four-arm probe on the reference box measured the matcher as prefix-only, with no rule spelling that scopes a verb behind `-C`, so `Bash(git -C:*)` stays rejected and a seat runs plain `git` — its cwd is already the workspace. The record is `STATE.md`'s "Closed without code" entry with its 2026-08-10 amendment, plus `autoAllowedTools` in `internal/council/vendors/claude.go`. Do not re-open it without a new measurement. **Amendment, 2026-08-08: the finish line and the `d` key.** Two deferrals came off the list: - **Every racer's arena block now carries a finish line** — *"2nd of 4 · done · 25.0s"* — and each part keeps its own epistemics. The rank is the order the ROOM saw seats land (finishColumn call order, host-stamped; event batching bounds the resolution, which is the honest limit of what was measured — a vendor's own claim about when it finished is an inferred value wearing measured clothes and is not consulted). The phase word is welded to the rank on purpose: "2nd · failed" and "2nd · done" are different facts, and a bare number would let a fast crash read as a podium. A DNF ranks — it landed, just not well. The elapsed is the column's own measured clock. parallel-code's results screen is the pattern source, minus its star rating, which is a judgment no gauge here is allowed to render. - **`d` flips the focused seat's arena block from stat to the full patch** and back. Per column, because reading A's stat against B's whole diff is a legitimate way to compare. The frame renders at most 400 patch lines (`arenaDiffScreenLines`) — a render cost bound, not a data bound — and the cutoff names how many lines it dropped and both routes to the rest (`y`, and the worktree itself). Three refusals with three sentences: no race this turn, a measured nothing-to-show, and an unreadable diff carrying its reason. ~~Plain text, no diff colouring yet: +/- prefixes are the first signal and survive `--ascii`; colour through the existing palette is a later, separate change under style.go's no-new-hues rule.~~ **Coloured 2026-08-09, through the existing palette and nothing else** (`Styles.ForDiffLine`): added lines wear `SevOK`, removed lines `SevCrit`, headers (`diff --git`, `index `, `---`/`+++`, `@@`) the muted chrome style — no new hue, per style.go's rule. Classification reads the raw prefix with headers matched first, so `+++` never wears the addition's green. The `+`/`-` prefixes stay the first signal: `PlainStyles` renders the same bytes as before, which is why no golden moved, and `--ascii`/`NO_COLOR` see exactly the frame they always did. The stat blocks (interim and final) stay unstyled — a stat is a summary, not a patch line. **Amendment, 2026-08-09: the live stat — the race shows the diff growing.** Until now a racer's stat appeared only when its column finished; the audience of a 20-second race watched three spinners and then a scoreboard. The pattern is the manager lane's event-triggered diff refresh (codeg's), rebuilt under this room's honesty rules (`internal/council/arenalive.go`): - **Event-triggered, throttled, off the loop.** Stream activity on a racing column (text or a tool call — a session id arriving is not evidence the tree moved) ARMS a re-read of that seat's worktree — `git add -N . && git diff --stat`, the same two claude-squad mechanics the finish-time read carries, stat only (the full patch stays a finish-line deliverable). An armed seat is read at most once per `arenaRefreshInterval` (2 s: the read is a subprocess pair, the audience is human, and the first live podium ran 7 s/15 s/19 s — faster buys frames nobody can distinguish), timed off the tick-stamped `State.Now` so the throttle is testable and Render stays pure. The read runs as a Bubble Tea command (goroutine → `arenaStatMsg`), never inline in Update, never in Render; one read in flight per seat, a due refresh that finds one running SKIPS rather than queues. An idle seat never arms, so an idle seat is never read; a seat whose worktree failed setup has no refresh slot at all, by construction. - **The interim marker ruling.** A mid-race read is a measured value at a moment that is already past, so the block's label is `arena · so far` — the "so far" is the whole marker, the same honesty spend as an estimate's `~` — and it withholds the finish line's receipt (branch, tree path, rank), which would dress an interim block in the final's clothes. Three states stay three renders (§4a.1, mid-race edition): no read yet is the nil pointer and renders NOTHING (absence, not a zero); a read that returned empty says "no changes yet against " (the "yet" is what separates a running seat's measured zero from the final's settled one); a failed read carries git's first stderr line, never dressed as no-changes. A failed read degrades only the live stat — the race runs on — and `arenaRefreshMaxFails` (3) consecutive failures end the seat's live stat WITH THE STOP NAMED on the column ("stopped re-reading … the finish-time diff still runs"), because a gauge that quietly freezes goes on reading as live. A success resets the count: the likeliest failure is the refresh contending with the vendor for the worktree's own index, and one contended read is not evidence the tree is unreadable. - **The finish-time `collectArena` read stays the authoritative final and REPLACES the last interim — cleared, never merged.** The refresh state lives on the turn, so teardown ends all refreshing with no cleanup path to forget; a read that outlives its turn or its seat arrives as a stale message and is dropped by comparison (turn number, final-already-landed), not by hoping the timing worked out. The one collision the feature introduces is named in the code: an interim `add -N` holds `index.lock` for milliseconds, so a final read that fails while a refresh is in flight is retried once — reporting "diff unavailable" for a lock this feature itself held would be the refresh degrading the read it exists to complement. Verification note: the mechanics — arming, the throttle, single-flight, the three interim renders, replacement by the final, stop-on-turn-end and stop-on-repeated-failure, the `add -N` false-zero property of the interim read — are pinned by offline tests (`arenalive_test.go`), the git ones against a real temp repository. ~~No live race has watched the stat move yet~~ **Half paid, 2026-08-09, and the halves are worth keeping apart.** The block APPEARING mid-race and reading honestly is live-verified by the give-up amendment's own race below: the stuck cursor racer's `arena · so far` read "no changes yet against \" for 26m40s, which is the interim empty state observed live for longer than anyone wanted. What is still owed is the other half — a "so far" block that GROWS and then swaps for the settled block at the finish — because no live race is recorded as having watched a NON-empty interim stat change. One `/arena` against a brief that changes files pays it, and it is stated here rather than implied paid. **Paid, 2026-08-15/16, race t9 — the first 5-of-5, and the growing half both.** The owner raced a file-changing brief across all five seats. The Antigravity and Cursor columns both drew a non-empty `arena · so far` that CHANGED on a later refresh (one file, then two) and was replaced by the settled block at landing — the growing half, watched twice over. The race's full record: three clean finishes with ranks and receipts (claude `1st of 5 · 50s · committed 2770c0c`, grok `2nd · 1m8s · 4aba168`, codex `3rd · 1m10s · 6874ff2`), and two seats given up with `x` after ~11 minutes (agy `4th · cancelled · committed 91fc53e`, cursor `5th · cancelled · committed 6b94b55`) — the give-up's second and third live exercises, and both cut seats kept their commit receipts exactly as the finish-line design promised. The cause of the two stalls was measured from OUTSIDE the room before the cuts: both racer processes were alive with ~zero CPU over an 8-second sample and no go toolchain process existed anywhere, so the work was done and the vendors' own turns had stalled — agy inside a `manage_task`/`schedule` poll loop that stopped polling, cursor after its final tool step. `/adopt claude` then exercised the dirty-room refusal live (the probe turn's two throwaway edits held the tree; the card named them; the operator restored and re-ran) before cutting `adopt/t9-claude` and landing the `--no-ff` merge cleanly — the refusal path's first live run. **One new gap, found by the same race:** `/trace` was armed before the turn and recorded NOTHING for it — the trace file holds only the preceding ordinary turn's line — so an arena turn produces no per-seat spawn/wait/stream split, and grok's timing on that axis stays unmeasured. Recorded as an unowned gap in STATE.md. **Amendment, 2026-08-09: the cursor seat races too, on a throwaway session.** The deferred follow-up the verification note filed is built, in the shape it predicted. For an arena turn — and only there — dispatch recognises the Conversational seat and, instead of the `FirstTurn` one-shot it deliberately refuses, launches a throwaway `cursor-agent acp` server rooted in that racer's worktree, runs exactly one `session/new` and one `session/prompt` through §9.36's own protocol driver, and kills the process when the column lands. It is `startEphemeralRacer` (persistent.go), a sibling of `spawnSeat` reusing `Cursor.Open`, `acpProtocol` and the counted `startRPCSession` spawn — no second ACP implementation exists to drift, and where the client was welded to the room seat's lifecycle the seam extracted was placement, not protocol. - **The room's conversation is untouchable by construction, not by discipline.** The race session opens with an empty resume id (never persisted, never resumed), registers on the TURN (`turnState.arenaEphemeral`) rather than in the seat-process registry — so a live persistent cursor seat and its racer coexist without either being mistaken for the other — and the throwaway session id it reports is refused by the existing arena guard before it can reach the saved threads or room.json. - **Kill, never wait.** §9.33 measured this vendor's process lingering ~2.5 s after answering, so the racer is killed at its own finish line — before the diff is read, making the receipt a snapshot of a stopped attempt — and on a protocol-reported failure (an ACP server survives its own refusals, so no exit event would ever have come), on ctrl+c, and at room teardown. Its context is the turn's rather than the room's, which is the backstop on every one of those paths; a seat cannot be cleared mid-race at all, because `askClearSeat` refuses while a turn is in flight. - **Two processes now wear one vendor id during a race**, so the eleventh amendment's stale-exit guard grew an attribution rule on the same liveness test it already trusts: an exit that arrives while the racer is alive can only be the room's idle seat dying in the background (forgotten, race untouched); one that arrives after the racer is dead is the racer's own, and must not be eaten by the guard reading a live room process as "this seat is fine" — that would hang the race column forever. - **The exits keep their epistemics.** A racer that dies without its end-of-turn response FAILED — on this seat the turn's end is a response, so a bare exit means no answer arrived, and rendering it done would be the empty-success this seat's missing result line makes possible. A turn that ends cleanly having streamed nothing lands done with a note naming the ambiguity, because a silently-working racer and a broken chunk parser are identical on this wire (§9.36's stated loss). Token usage stays what ACP makes it: absent, never zero. And the containment phrase every racer carries — write posture, contained by the worktree — is stated at its weakest here: §9.36 measured workspace trust not applying over ACP, and an arena worktree is a freshly created, never-trusted directory, so nothing but the session's cwd scopes the attempt. The posture detail already says so; the worktree gives it more force, not less. Verified offline only: fixture-driven tests (arena_cursor_test.go) pin the spawn choice, the untouched room thread, the kill on finish / protocol failure / cancel / teardown, both degraded exits, and the exit-echo not re-ranking the race. cursor-agent was not installed where this was built, so ~~this amendment owes a live race on the Windows box~~ **the amendment owed a live race; two have now run it, and what they paid is narrower than the word "verified" would suggest.** The throwaway racer has been spawned live and rooted in its own worktree — it streamed for 26m40s on the race the give-up amendment below records, and was cut loose mid-race with `x` on race t9 — so the spawn, the live interim read against its tree, and the kill path are measured. **A clean completion is still owed**: no live race is recorded in which this seat's own `session/prompt` resolved and the racer was killed at its own finish line with its diff read. Until then the *finishing* half of this amendment stands on `arena_cursor_test.go` alone, and that is the honest split. *(The 2026-08-15/16 5-of-5 race cut this racer a second time — alive at ~zero CPU with its edits complete and committed on the cut, ~11 minutes in — so the debt stands and gained a second data point: two live races, two stalls, zero self-finishes.)* **Amendment, 2026-08-09: every attempt survives as a commit, and `u` takes one back.** The commit-per-turn deferral came off the list, and it brought the rollback it makes possible (mechanics stolen with attribution: Crystal's commit-per-turn checkpoint, cc-haha's turn-level undo). Two halves that stack: - **Commit-per-turn.** The moment a racer lands and `collectArena` has read its diff, the worktree's whole state is staged and committed onto `arena/t/` — subject `arena t: ` (64-byte cap) — so every attempt is durable on its branch: diffable, adoptable, rollbackable, and the worktree itself becomes deletable without losing anything. Staging everything is correct *there and only there*: the tree contains nothing but the racer's own output, so the reason blanket staging is wrong in a real workspace does not exist in that one. The sha the column renders is exactly what `git rev-parse HEAD` returned, shortened for display only. On a machine with no git identity anywhere (CI runners, fresh boxes) the commit carries a fallback identity via per-command `-c` flags — never a config write, which on a worktree would land in the shared repo config, i.e. in the room's repo. A commit that cannot land (a stale ref lock, a signer that cannot run) degrades **that seat's receipt only**, as `not committed: ` — the diff was still read, the race and the other racers are untouched. A racer that committed for itself mid-turn already parked its attempt; its own tip is reported rather than papered over with an empty commit — and the diff still answers against the recorded base, so the mid-turn commit cannot hide the work. - **The empty-commit ruling.** A zero-diff attempt commits **nothing**, and that is a ruling, not an omission: an empty commit would be a receipt claiming work that did not happen — §4a.1's false zero, mirrored into the write path. The seat renders no commit line and no failure either (nothing was owed); "no changes against \." stays the whole story, and the branch tip staying at base is the machine-readable form of the same fact. - **Undo-the-whole-turn.** `u` on a focused arena seat, y/n-gated exactly like `c` (a stray keystroke must cost a `y` before it costs an attempt), runs `git reset --hard ` inside the racer worktree **only**. Branch and tree agree by construction rather than by a second command: the worktree has its arena branch checked out, so `--hard` moves that ref itself. The safety argument is an explicit path guard, not trust in recorded state: the reset runs only on a path equal to the arena-tree name recomputed from the room's *current* workspace, turn and vendor — a name that structurally cannot be the workspace itself — and anything else refuses before git runs. Refusals are four sentences for four facts: no race this turn; the attempt changed nothing (nothing to take back); already undone (pressing again is not more undone); and the reset itself failed, carrying git's own first stderr line. After an undo the stat stays on the column — the measured record of what the attempt changed — under an "undone" line saying the tree and branch no longer hold it. - `u` landed on the help panel's room-controls row at its exact 114-cell budget by trading the word "worktrees" for it: the arena block prints the worktree path on every race, so that clause restated something the screen already teaches, while an undo key documented nowhere is a control nobody finds. Verification note, on the same terms as the section's own: the git mechanics — the commit landing on the branch with the racer's files, the per-seat degrade, the zero-diff skip, the self-committed tip, the undo round-trip (files restored, created files gone, branch back at base), the path guard, and every refusal sentence — are pinned by offline tests against real temp repositories. ~~No live race has exercised either half yet.~~ **Commit-per-turn is paid; the undo is not, and it is now the last unpaid item in this section.** Race t9 (Windows box, 2026-08-09) landed the claude racer's attempt as `cf69634` — subject `arena t9: write table-driven tests for the small render helpers …`, parent `e1bf983`, which is the base the race recorded — onto `arena/t9/claude`; that commit outlived `/adopt`, `/arena drop` of its worktree, and a merged PR (#164), with the branch `adopt/t9-claude-helpers` left as the receipt. **`u` has still never run against a real racer commit.** The debt is one `/arena` against a brief that changes files, then `u` on the finisher between turns. **Amendment, 2026-08-09: `.worktreeinclude` — a race carries the files git ignores, when the repo names them.** The seeding deferral came off the list, on the schedule the original note predicted (a real race on a repo needing `.env` fails falsely on every seat at once, so the fix is worth landing before that repo shows up). What was built, and the rulings inside it: - **The file and the copy.** A `.worktreeinclude` at the room repo's root — gitignore-style patterns, one per line, `#` comments and blank lines ignored — and during `arenaSetup`, after each racer's worktree is added, every matching file is copied from the room repo into that tree, relative paths preserved, parent directories created. The grammar is a documented subset of gitignore's (bare names match at any depth, anchored patterns from the root, `*` within a segment, `**` across segments, a directory pattern takes its subtree; no negation). - **Copy only, never execute — the half of agent-deck deliberately not taken.** agent-deck (the pattern source) pairs seeding with repo-carried setup scripts that run after the copy. That half crosses a trust boundary this project has explicitly parked: byte-level trust gating is on the parked list pending an audit, and a repo that can run code on the machine by merely containing a file is a different product with a different threat model. Copying bytes into a tree the room already owns is containable; execution is not. - **Candidates are untracked files only** (`git ls-files --others`, ignored files included — exactly the set a fresh worktree lacks). A tracked file already arrives with the checkout, and seeding the room's possibly-dirty copy of one would plant the room's own edits in every seat's diff — a lying diff, §4a.1's class. Known limit, stated: a seeded file that is untracked but *not* git-ignored still surfaces through collection's `git add -N .`; name git-ignored files and it cannot. - **Containment.** Patterns resolve from the repo root and matches come from git's own enumeration, so they structurally cannot leave it; absolute and `..`-carrying patterns are refused by name anyway, per pattern, so one bad line disables only itself. Symlinks are never followed — Windows is primary and symlink semantics differ per platform — a symlink match copies nothing and says so. - **The budget: 64 MiB per seat** (`seedBudgetBytes`). Exists for the node_modules pattern — an over-broad line must fail loud and named, not hang the room copying a dependency tree into four worktrees. Over-budget is refused wholesale (copying *some* of the file's list would hand every seat a tree that half-works), with the measured total in the sentence, and the budget is enforced again on actual bytes during the copy, because files grow between stat and copy. - **Honesty in the column.** "seeded 3 files" is the count actually copied into that seat's tree; no `.worktreeinclude` means no line at all (zero and absent stay two facts). A pattern that matches nothing is a named notice, not silence — an allowlist-shaped file fails both ways — and not a failure either. A copy error degrades that one seat with the path and the error's first line, through the same per-seat lane a failed worktree add uses; the race runs on, and the half-seeded worktree stays on disk (kept-until-deleted receipts include the broken ones). Verification note, same shape as this section's original one: the mechanics — the copy into each racer tree, nested parents, every named refusal, the budget, the per-seat degrade channel, zero-vs-absent on the seed line — are pinned by offline tests against real temp repositories. **No live race on a repo that actually needs a `.env` has run yet; that run is the debt this amendment carries**, and it is the same debt the original note carried for the core, paid the same way. **Amendment, 2026-08-09: the end of a worktree's life — `/adopt` and `/arena drop`.** The deferred deletion guard came off the list, and adoption moved from a printed suggestion to a typed room command; both are §9.17 verbs, reachable mid-session with no flag and no relaunch (`lifecycle.go`). The shapes borrow from the two products that had already worked this seam — claude-squad's adopt, Pane's deletion guard — with one deliberate fork each: - **`/adopt ` merges; it does not check ~~out~~ the racer's branch out.** claude-squad's adopt is a checkout of the attempt's branch over the user's — which moves HEAD, rewrites the tree wholesale, and leaves the user's own branch behind. That is more state than the act requires, and §9.37's whole posture is offer-never-take, so council does the least-magic git operation that lands the work: `git merge --no-ff arena/t/` in the room's repo, ~~on the branch the user is already standing on~~ **on a fresh branch council cuts for it — ruled 2026-08-11, the last block in this section; the merge itself is unchanged**. `--no-ff` keeps the adoption a visible event in history — the merge commit is the receipt saying where the work came from. Because arena seats leave their work uncommitted (commit-per-turn is still deferred), a dirty attempt is first committed in its OWN worktree, on its OWN arena branch, under the user's own git config — and the y/n card says so, naming the exact command(s) y will run, the flow write gate's contract. Hard precondition, refused by name with the path count: the room tree must be CLEAN (`git status --porcelain` empty) — a merge writes into that tree, and adopt must never eat the user's uncommitted work. A racer that changed nothing refuses (an empty merge commit would claim work that does not exist); a merge that conflicts is `git merge --abort`ed — tree restored, attempt intact on its branch — and the notice hands the merge to a human. A merge that failed before starting reports the tree as untouched instead: the two endings are different facts (§4a.1). Posture is not consulted: read/write governs the seats, and an adopt runs on the user's own y, the same footing as /cd. - **`/arena drop ` (or `all`) deletes tree + branch, guarded; the force is a spelling, not a keystroke.** Two guards, each refusing with exactly what would be lost and the way forward: a worktree holding uncommitted changes (counted), and an arena branch holding commits the room's HEAD cannot reach (counted, with `/adopt ` offered beside the force). The force form is a trailing bang — `/arena drop codex!` — re-run by the user, chosen over a second y/n on purpose: y is one keystroke answered against a notice half-read, while the bang travels in the command, records that destruction was asked for, and cannot be produced by a stray key. (`/adopt` keeps y/n because its act is additive and revertible; drop orphans work.) Mechanics are `git worktree remove` (git's own `--force` only when the user spelled it) then `branch -D`, argv via gitOut — and the path check is mechanical: a tree is only ever removed if it re-derives, from the recorded race's own workspace/turn/seat through the same `arenaTree` that minted it, to exactly the recorded path. A receipt entry that fails that check is refused even under force. `drop all` degrades per seat rather than refusing wholesale: clean trees go, survivors are named with their reasons. - **The target is the RACE'S receipt, not the column.** `Column.Arena` is a per-turn fact the next dispatch clears; the worktrees are kept until deleted. So dispatch records the race — workspace, turn, base, each racer's tree — on the model (`arenaRace`), in memory only: room.json stays keys-and-numbers, and a room reopened after a quit finishes the lifecycle by hand with the same git commands, against worktrees that are visible siblings precisely so no session state is needed to find them. Grammar note: only the exact two-word form `/arena drop [!]` is the verb; anything longer after `/arena` is a brief and races as prose, the roomcmd vocabulary rule applied inside the one command that takes free text. The help panel's room-commands row is at its width budget and does not name the two verbs; they are taught by the slash refusal (which lists `/adopt` in the live table), by bare `/adopt` and bare `/arena drop` answering with usage, and by every guard refusal naming its remedy. Verification, ~~owed~~ **paid 2026-08-09**: the git mechanics — merge, commit-then-merge, conflict abort, both guards, the force, the path check, `drop all`'s partial degrade — are pinned by offline tests against real temp repositories (`lifecycle_test.go`), and ~~no live adopt has run on the Windows box~~ **race t9's winner went through `/adopt` and then `/arena drop` on the Windows box**, which is why no `arena/t9/*` branch and no `telltale-arena-t9-*` sibling exist on it while every earlier race's leftovers sit exactly where they were left. What that adoption then cost in hand-run git is the open question at the end of this section, filed rather than ruled. Two guard paths a SUCCESSFUL adopt cannot reach still rest on `lifecycle_test.go` alone: a merge that conflicts, and a drop refused for unmerged commits. **Amendment, 2026-08-09: the race numbers itself off the refs, and a failed race says why.** A live `/arena` at turn 3 (Windows box, real room) failed on all four seats, and each column's whole explanation was `arena: Preparing worktree (new branch 'arena/t3/')`. Two measured defects, one incident — the second is what made the first expensive: - **The collision.** Kept-until-deleted cuts both ways: arena branches and worktrees outlive the room, but the turn counter — and the in-memory race receipt `/arena drop` needs (`Model.lastRace`) — reset with every launch. So a fresh room's turn 3 minted the exact names an older room's turn 3 had already parked, `git worktree add -b` refused every seat, and drop could not reach the old trees because their receipt had died with the old room. The only remedy was hand-run git. The fix reads instead of guessing: at setup the race number is `arenaRaceNumber` — one past the highest N among the repo's existing `arena/t/...` branches (`git for-each-ref` over `refs/heads/arena/`, argv via gitOut), floored at the turn number, so a repo with no leftovers keeps racing as `t`. The refs are the one record that shares the leftovers' lifetime, which is what qualifies them to number the race; a scan that cannot run degrades to the turn-number floor with the race running, because a broken for-each-ref must not brick `/arena`. The number is recorded once (`turnState.arenaRaceN`, `arenaRace.raceN`, `ArenaResult.RaceN`) and EVERYTHING that mints or re-derives a name reads it — the branch on the receipt, the `arena t:` commit subject, `/adopt` and `/arena drop`'s re-derivations, undo's path guard — because the turn and the race now legitimately disagree, and one call site still reading `Column.TurnN` would aim a verb at names the race never created. The arena block's render is untouched: it already shows the branch name, which carries the (now honest) `t`. - **The lie about the collision.** gitOut surfaced the FIRST stderr line of a failed command, and `worktree add` prints progress chatter ("Preparing worktree ...") before its `fatal: a branch named '...' already exists` — so the column showed the narration and swallowed the diagnosis. gitOut now prefers the first line git itself marks as the problem (`fatal:` / `error:`), falling back to the first non-empty line only when no marked line exists (some refusals print bare prose). git's own prefixes are the measured marker of which line is the error; the old rule displayed the nearest string to the failure instead of the failure. If a residual collision still happens despite the scan — a sibling directory an old room left with no branch to be scanned, a ref minted between scan and add — the seat's error now carries the fatal line plus the named remedy (`git worktree remove` / `git branch -D`), since those leftovers are exactly the state no receipt can reach. Both fixes are pinned by offline tests against real temp repositories: the two-line-stderr collision fixture (the live transcript, replayed), renumbering past stale `t3` branches with every seat racing clean, the turn-number floor on a failed scan, the residual-collision sentence, and adopt/drop/undo driven end to end against a race whose number outran its turn — leftovers untouched throughout. ~~The live re-race is owed~~ **The live re-race ran, and has kept running ever since**: every race after the fix has been numbered over a growing pile of leftovers — 27 `arena/t/` branches and 28 sibling worktrees, t2 through t8, are still on the reference box — and race t9 raced all four seats clean over them, which is the claim. One consequence worth knowing before the next race: t9's branches were dropped, so the highest surviving `arena/t` is t8 and the scan will mint `t9` again. **Amendment, 2026-08-09: `x` gives up on one racing seat, and the race runs on.** The second live `/arena` (same day, Windows box) measured the gap: four seats raced, three landed (7m51s / 26m28s / 5m07s), and the fourth — the cursor throwaway ACP racer — streamed for **26m40s** with the live stat honestly reading "no changes yet against \" the whole time. The operator sat ~20 minutes after the race was effectively decided, because one stuck racer holds the WHOLE turn hostage: ctrl+c is the only exit and it cancels everything. The room displayed the truth and offered no per-seat act on it. The act built: - **`x` on a focused, still-racing arena seat, mid-turn** — the one per-seat key that runs while a turn is in flight, because mid-flight is the only time it means anything. y/n-gated exactly like `c` and `u` (a stray keystroke must cost a y before it costs a process), and the question names the vendor and what y does. On y, the room kills THAT racer only — the ephemeral ACP session when one is racing, else that vendor's one-shot process through `turnState.arenaHandles`, the per-vendor record dispatch's arena branch now keeps beside the flat `handles` list (the flat list stays: cancel and teardown are all-or-nothing acts and never address a single process; the give-up is the first act that does). The kill lands on the racer's side of the two-processes-one-vendor-id split — the room's idle seat behind the same id survives, per applyEvents' existing attribution rule. - **A given-up seat lands like any other finisher, wearing the honest phase.** The stream tail is flushed, the elapsed stamped, the note says "given up after \ — anything it wrote is in the diff", and the column retires through `finishColumn` with the CANCELLED phase — the same phase and render ctrl+c's cancel produces ("cancelled — the output above is partial" is that path's wording; this one names the give-up instead). Everything the finish line already does happens unchanged: the racer dies before the diff is read (the receipt is a snapshot of a stopped attempt), a dirty tree commits its receipt onto the arena branch, a clean tree stays a measured zero with no commit, the rank is stamped in host-observed landing order — a DNF finished too, and the render welds the rank to the phase word so "4th · cancelled" cannot read as a result — and the interim stat clears. The seat leaves the turn's live set through the same drain every landing uses, so **the turn ends when the remaining seats land** — which is the whole point. - **Three refusals, three sentences** (the undo key's rule): no turn in flight; ~~an ordinary turn — its seats share one fate by design, this key is arena-only and **ctrl+c remains the whole-turn act**, said in the refusal~~ **(REVERSED 2026-08-17 — the block below; the key now runs on an ordinary turn, and the refusal that took this one's place is a turn ctrl+c is already cancelling. ctrl+c is unchanged and is still the whole-turn act)**; and a seat that already landed (its result is settled — a y arriving after the seat lands under the question refuses the same way, killing nothing and re-ranking nothing). The help panel's room-controls row is at its exact 114-cell budget and does not name the key; it is taught by these refusals and by this amendment, the way `/adopt` and `/arena drop` are taught by theirs. Verified offline only, the section's standing debt shape: real-temp-repo plus fake-session tests (giveup_test.go) pin the ephemeral kill and the cancelled landing with rank and committed receipt, the keyed one-shot kill with every other racer's handle surviving, the room process surviving its racer's give-up and the exit echo landing inert, the turn ending when the remaining seats land, all three refusals, the y/n/stray gate, and compose leaving `x` a letter. ~~A live give-up on the Windows box is owed~~ **Paid on race t9** (Windows box, 2026-08-09): the cursor throwaway racer stalled, was cut loose mid-race with `x`, and the turn ended when the remaining seats landed — which is the whole claim, measured. The three refusals and the y/n/stray gate stay offline-pinned, and always will be: a live race has no way to exercise a refusal it never trips. **Amendment, 2026-08-17: `x` gives up on one seat of an ORDINARY turn too — the one-fate line is reversed.** The amendment above refused the key outside a race on a stated design position: *an ordinary turn's seats share one fate by design*. The owner reversed that position on 2026-08-17, and the reason is that the position was **written for the four-seat room**. It said, in effect, that a brief and its answers are one act, so a seat that is still working is the turn still working. The five-seat room (§9.39) supersedes it: with five vendors on an `@all` turn, **one stalled vendor while the other four have answered is the most probable live failure this room has** — it is the failure two live races already produced inside `/arena`, and nothing about it is a property of worktrees. The hostage argument that built the key does not change when the brief is prose instead of a race. Only the cost of the cut changes, and that is what the room now says per seat kind. - **How a seat is stopped is per seat kind, and the three arms are not interchangeable.** A batch seat (codex, agy, grok) is KILLED, through `turnState.seatHandles` — new plumbing that mirrors `arenaHandles`, not a widening of it, because `arenaHandles` also answers `arenaRacing`'s question about whose exit a `KindDone` is while two processes wear one vendor id, and an ordinary handle in that map would send every ordinary exit down the racer's branch. The persistent claude seat is **INTERRUPTED** (`interruptSeat`), never killed: killing it would work and would also throw away the conversation and the session-init cost that bought it, so cutting one turn would silently make the next one expensive — `cancelTurn`'s own argument, applied per seat. The next brief resumes it. The ACP cursor seat needs no third arm: on an ordinary turn it IS a persistent seat and takes the interrupt, and the throwaway racer the arena kills only exists during a race. The flat `handles` list is untouched, for its own reason: cancel and teardown are all-or-nothing acts that never address a single process. - **Four endings, four sentences.** The cut column lands CANCELLED through `finishColumn`, keeping everything it streamed, and its note has to stay distinguishable from the three other ways a column ends with no answer: *not addressed* ("not addressed in turn N", with `Column.Skipped` set), *ctrl+c* ("cancelled — the output above is partial", still the whole-turn act), and *a measured empty answer* (a body reading "[Turn completed with 0 text chunks streamed]" under `done`). The give-up's own note names the elapsed, whether anything had arrived **when it was cut**, and what became of the seat — past tense on purpose, so a killed child's last buffered chunk landing after the column retires cannot make the sentence false. **A seat that streamed nothing must never acquire the placeholder**: that would be §4a.1's false zero in its sharpest form, a seat the operator stopped claiming to have measured nothing. `testdata/golden/given-up-vs-zero.txt` pins the two side by side. Holding that took one fix outside the give-up itself, and it is the same defect the eleventh amendment's end-of-turn branch was already fixed for: the placeholder is a claim that THIS turn completed, so only a column still in a live phase may acquire it. The phase write on the `KindDone` exit path had always been guarded and the BODY write had not, so an exit landing on a column that had already ended overwrote it. The give-up makes that reachable — a cut seat's child can exit after the turn boundary, past `turnState.givenUp`'s lifetime — and the guard is now on the body write too. - **The cut seat's own late events land inert, by guard rather than by luck.** `turnState.givenUp` is recorded BEFORE anything is stopped, because both stops provoke one more event: a killed child drains its buffered stdout, and an interrupted persistent seat answers with its own failed `result` (measured is_error true, terminal_reason "aborted_tools"). Unguarded, that error would overwrite "given up after 4m12s …" with the vendor's abort text and record a vendor failure against a seat the user stopped. The guard still does the PROCESS bookkeeping — an interrupted seat whose process later dies for real is forgotten, so the next brief does not write into a closed pipe. - **The refusal set changed by one.** "This turn is not a race" is gone. Its place is taken by a turn ctrl+c is already cancelling: every seat is going anyway, so a per-seat act would only re-label one of them. The other two are unchanged — no turn in flight, and a seat that already landed — and the re-check on `y` is unchanged too, because events drain between the card arming and the answer and the seat can land while the question is up. Verification, stated honestly and in two halves. **Offline**: `giveup_test.go` pins both seat kinds on a real `@all` turn — the batch kill reaching exactly one process with every other seat still working, the persistent seat interrupted rather than killed and still registered, the interrupted seat's own abort error not overwriting the give-up, the cut seat that streamed nothing never acquiring the placeholder, the turn ending when the remaining seats land, the four endings reading as four different sentences, the new refusal, and `--ascii`/`NO_COLOR` parity on the cut column. No test here spawns a vendor (`countSpawns`, per the council-test rule). **Live: a LIVE ordinary-turn give-up on the Windows reference box is OWED**, as a dated payment before 2026-09-30. Offline tests cannot exercise it: what is unmeasured is whether a real vendor's interrupt lands on a real persistent seat mid-turn and whether that seat's NEXT brief actually resumes the conversation, which is the whole claim the persistent arm makes and the one thing a fake session cannot witness. Until that date and that run, the interrupt arm stands on `giveup_test.go` and on `cancelTurn`'s already-measured interrupt, and this paragraph is the record that it does. **Amendment, 2026-08-09: the brief carries the conduct line — the one place the room adds words.** A write-posture racer's confinement is its worktree, but the machine's git and gh credentials are ambient, so a racer can reach GitHub — and one did: the codex seat took the t5 gofmt brief, pushed its arena branch, opened a PR, waited out CI, and merged it into main, then announced the same plan on the very next race. This section's founding ruling — the room offers the diffs and the human adopts, never an auto-adoption — binds this codebase and cannot bind a vendor that runs `gh pr merge` on its own initiative. The operator ruled the same day: every `/arena` dispatch now prepends `arenaConduct` (arena.go) to the brief — *"This tree is a race attempt in its own git worktree. Do not push, open pull requests, or merge — the operator compares the attempts and adopts the winner. Do the work, verify it locally, and stop."* The bend to the brief-verbatim promise is bounded three ways, each pinned by test: races only (an ordinary turn's prompt is byte-verbatim), a constant — the same published line for every racer on every race, prepended so a long brief cannot bury it, leaving the cross-seat comparison undisturbed — and recorded here rather than discoverable only in a wire capture. Stated honestly for what it is: an instruction, not a control. A vendor can ignore it, and the mechanical version of this boundary — credentials a racer cannot reach — is a different, harder change that this amendment deliberately does not claim. ~~The live measurement owed: the next race showing the codex seat stopping at its commit.~~ **Measured on race t9**: the codex seat raced under the preamble and ended with "nothing pushed." One race is one race — an instruction a vendor obeyed once is still an instruction, so the sentence above stands exactly as written and this measurement does not promote it to a control. **Open question, 2026-08-09 — RULED 2026-08-11, option (b): where should an adoption land?** Filed from the first live `/adopt` rather than decided, because the answer depends on a convention the room cannot see. `/adopt` merges into the room repo's current branch — for most workspaces, local `main` — and that is the smallest honest act the verb can perform. But an operator whose repos are run branch→PR (this project's own convention, and this operator's standing rule across every machine) then holds a merge commit on a local `main` that must never be pushed as-is; the first live adoption ended with four hand-run git commands turning the merge into a PR branch and resetting `main` back to origin. Options, ~~none ruled on~~ **(b) ruled, 2026-08-11 — the block below**: (a) keep the current shape and document the hand-off (an adoption is a local act; publishing is the operator's, as it already is for every other commit); (b) `/adopt` onto a NEW branch cut from the room's HEAD (`adopt/t-`?), never touching the current branch — closer to branch→PR, but the room minting branch names in the operator's repo is a bigger footprint than one merge; (c) a flag or second verb for each. The founding posture — offer, never take — leans (b) no further than it leans (a): both are one revertible act on the operator's own y. Whoever picks this up starts from the measured friction: four commands, once per adoption, only on branch→PR repos. **Ruling, 2026-08-11: an adoption lands on a fresh branch, never on the branch the workspace has checked out.** The owner ruled option (b), and the reason is the convention the room could not see: the owner's workflow is branch-then-PR on every repo and every machine, and a commit straight to `main` is never made. The room's own posture does not decide this — both options are one revertible act on the operator's y — so the measured friction decides it, and the friction is four hand-run git commands per adoption on every branch→PR repo. A fresh branch turns that hand-off into one `gh pr create`. What was built: - **The name is `adopt/t-`, cut from the room's current HEAD and checked out.** The race number, never the turn, like every other arena name (the renumbering amendment above). The seat is joined by a dash rather than a slash so `arena/t9/claude` and `adopt/t9-claude` differ in the last segment, which is where a reader of `git branch` looks. `--no-ff` and the commit-then-merge order are untouched: what moved is WHERE the merge lands, not what it does. - **A collision takes the next free suffix** (`-2`, `-3`, … to 50, then a named refusal). The collision is ordinary rather than exotic: race numbers repeat once a race's branches are dropped, since the scan that numbers a race reads `refs/heads/arena/` alone, and an operator can adopt, revert and adopt again. Reusing the name would either fail the checkout or land the merge on an older adoption's branch, where the PR would carry work nobody asked about. The free name is resolved WHEN THE CARD ARMS and carried to the `y`, because a card that named `adopt/t9-claude` and then cut `adopt/t9-claude-2` would be the card describing something other than what it ran — the one contract this gate exists to keep. One `for-each-ref` over the adopt namespace answers every candidate at once and separates "no such branch" from "git could not answer"; a scan that cannot run degrades to the plain name with the adoption still running, and `git checkout -b` reports the collision with git's own fatal line. - **A failed adoption leaves nothing behind.** The branch is cut for a merge, so a merge that conflicts or refuses ends with the branch deleted and the room back on the branch it came from — an empty branch handed over as the receipt of a failure would be the verb charging for its own failure. The two failure endings stay two facts (§4a.1): a conflict aborts and says a human merge is needed, a merge that never started says the tree is untouched, and both now also say where the room stands. A restore that cannot finish says THAT instead, naming the branch the room is left on and the command that puts it back. - **The alternative reading, recorded rather than taken**: leave the room standing on the fresh branch after a conflict, so the human merge happens there. It is defensible — that is where the resolution belongs — and it was not chosen, because the ruling covers where a SUCCESSFUL adoption lands and a failure quietly moving the operator to a new branch is a state change nobody asked for. If a live conflict makes the restore feel wrong, this is the line to amend. - **Existing refusals are untouched**, and so are their tests: the dirty-room gate with its untracked-bystander rule, the zero-change refusal, the mid-turn refusal, and `/arena drop`'s unmerged-commit guard, which reads the room's HEAD and therefore counts zero as soon as the adopt branch holds the merge. Verification, on this section's own terms: the mechanics — the branch cut and checked out, the room's own branch not moving, the suffix on a taken name, the conflict restore, and the notice naming the branch and `gh pr create` — are pinned by offline tests against real temp repositories (`lifecycle_test.go`). **Paid 2026-08-14** by race t14 on the reference box: the card named the branch and the merge before the `y`, `adopt/t14-claude` was cut and checked out, `git merge --no-ff arena/t14/claude` landed as `91c5f3e`, `main` did not move, and the notice named `gh pr create`. The push half was deliberately not run — the racer's payload was a throwaway date comment, and opening a PR to prove a verb works would put noise in the repository to record that the repository is fine. #### A warm seat's racer could not finish, 2026-08-13 — measured, then fixed **A race against a seat the room had already used never ended.** Race t10 on the reference box: one Claude racer, a one-line edit, and the room rendered `streaming` for **21 minutes** after the racer had exited. The vendor was not slow and did not fail. Its transcript ends at 52 seconds with a complete reply, and by then no process wearing that vendor id was left alive except the room's own persistent seat. Nothing downstream ran — no diff, no commit, no rank, no seed receipt — because all of them live inside `finishColumn`, and `finishColumn` was never called. **Both of the column's exits were closed at once, which is why nothing caught it.** A one-shot racer ends its turn by exiting, so `KindDone` is its only retirement signal; the two earlier paths do not apply to it, and each declines for its own correct reason. `arenaEphemeral` is populated only for a `Conversational` seat, so `ephemeralRacer` is nil for this vendor. The arena spawn never sets `turnState.persistent`, correctly — the racer *is* a spawn — so `isPersistent` is false. That leaves `KindDone`, and `KindDone` reaches the stale-exit guard first: a terminal event names a **vendor**, not a process, so a live entry in `m.procs` reads as "this seat is fine" and the exit is discarded as a predecessor's. The room's persistent Claude seat is exactly such an entry. **The failure mode was already written down, one path over.** `KindDone`'s own attribution comment names it for the ACP racer — a guard "reading a live ROOM process as *this seat is fine* would leave the race column streaming forever and the turn unable to end" — and fixes it with the `ephemeralRacer` check. `giveUpSeat`'s comment then observes that the guard eats a racer's exit when a room process wears the id, and judges it harmless. For `giveUpSeat` that judgement is right: a given-up column is already terminal when the exit lands. On the ordinary path the column is not, and the same swallow is the difference between a race that finishes and one that cannot. Two processes wear one vendor id for every racing seat, not only for the one whose racer happens to be a live session. **The trigger is a warm seat, and it is why this survived several clean races.** `m.procs` must already hold a live process for that vendor when the race dispatches. Race before the room's first ordinary brief and the guard never fires, which is what t9-on-2026-08-09 did when all four seats raced and a winner was adopted. The drive that found this sent two ordinary briefs first, so the seat was warm. It also means codex, agy and grok racers were never affected: none of those vendors holds a persistent process, so their exits pass the guard untouched. The bug reached exactly one seat, and it is the seat the room dispatches to by default. **The fix is attribution, not a hole in the guard.** `arenaRacing` is the one-shot sibling of `ephemeralRacer` — keyed presence in `arenaHandles`, because a handle is not a session and cannot be asked whether it is alive, which is precisely the case at hand: the process has already exited and the map is the only record of whose exit it was. A racing vendor's `KindDone` retires its column; a vendor that is not racing keeps the stale-exit guard exactly as it was. `dropProcess` is deliberately not called on that path — the exit belongs to the racer, the room's own seat is still running, and forgetting a live process would leave it running and invisible, which is the state this product refuses. `TestAWarmSeatsRacerRetires- OnItsOwnExit` was verified to fail without the change, reproducing the hang as `phase = streaming`; its sibling pins that a non-racing vendor's predecessor exit is still discarded. **Verified live the same day.** Race t13, one Claude seat, first dispatch of a cold room: the column retired in **18 seconds** and drew the whole settled block — `arena arena/t13/claude`, `seeded 1 file`, `no untracked file matches ".env"`, `1st of 1 · done · 18s`, `committed 1ef0f00.` and the diff stat. Eighteen seconds against twenty-one minutes is the measurement. Race t14 then landed the same way with a warm seat behind it, which is the arm that matters: it is the state the bug needed. **And it paid the three debts that had been stuck behind it.** `u` on t13 reset the branch to `ba2d00b` with the stat kept above an `undone` line, and a second press refused as already undone; the reset was confirmed against git rather than against the room's own claim. `/adopt claude` on t14 cut `adopt/t14-claude`, ran the `--no-ff` merge, left `main` at `ba2d00b` and named `gh pr create` — the first live adoption under the shape ruled 2026-08-11, and the debt this section states three paragraphs above its own dated block is now paid. **Amendment, 2026-08-16: a racer's turn says which race it was.** `STATE.md` recorded the gap on 2026-08-15/16: an operator armed `/trace` before a race, and the file held only the preceding ordinary turn's line. The report named the consequence correctly — grok's spawn/wait/stream split stayed unmeasured, because the arena is the one turn shape that dispatches grok. **The mechanism is a seam, not a dropped record.** A probe reproduced the arena's exact spawn shape against the real runner — `runner.Start`, a one-shot child, its own `Dir` — and the record came back: `grok spawn=5ms wait=43ms stream=2ms total=51ms`. So the clock does run for a racer, on both arena spawn paths. What the record could not say is that it was a racer. `runner/clock.go` states its own rule at the top: a `TurnClock` is keyed by the seat and the moment, because those are the only facts that package holds. The race number is not one of them. It is read off the repo's own arena refs at setup (`arenaRaceNumber`) and it lives on `turnState`, which is gone by the time the record is written — the runner emits at process exit, on its own goroutine. So four racers wrote four lines that named neither the race nor the worktree, and each line was byte-identical in SHAPE to an ordinary turn's line for the same seat. The trace held the race and could not point at it. **The fix carries the label the room already has.** `runner.Spec` gains `Race`, the one field in that struct the runner does not use and only carries. Council stamps it on both arena spawn paths: the batch seats through `FirstTurn` in dispatch's arena branch, and the merged cursor seat inside `startEphemeralRacer`, which builds its own spec through `cv.Open`. Both arms are stamped because a label applied at the obvious call site alone would leave that seat as the one racer nobody could find. `newClock` takes the race and fixes it for the life of the process, which is correct by §9.37's founding ruling: every attempt is a FRESH one-shot session, so no second turn on that process could belong to a different race. **Nothing is derived, and the ordinary line does not move.** The tag is `arena/t` (`arenaRaceTag`), minted from the same race number as the branch, so a trace line and the worktree its attempt is parked on cannot disagree — the vendor is already a column on the line, and the two together spell `arenaBranch` exactly. The field is APPENDED after `total=`, so every reader that already parses a trace line keeps its field order; a test pins that position rather than trusting whoever edits `String` next. An ordinary turn appends nothing at all. That is §4a.1 rather than terseness: an unmeasured `Span` prints `-` because the stretch existed and was not measured, but an ordinary turn is not a race whose id went missing — it is not a race, so there is no field to mark absent, and `race=-` on every ordinary line would invent a category for the room's normal case. **What is verified, and what is not.** The council half is pinned at the spec, and that split is forced rather than chosen: a council test never spawns a vendor (`CLAUDE.md`), so the clock cannot run there and the spec is the whole of what that package contributes to the record. The runner half spawns for real and asserts the race survives onto the emitted line, with the spawn/wait split still measured beside it. `TestArenaSpecsCarryTheRaceIntoTheTrace` was verified failing before the change, on all four racers and both spawn paths. **The live half is owed.** No race has been run against this build, so the claim that a real `/trace` now holds an attributable racer line rests on the probe and the tests, not on a race. One thing this amendment deliberately does not claim: it does not explain the operator's empty file. The records are emitted, so the reported absence has some other cause, and the same drive recorded two candidates beside it — a room that opened on workspace `~`, and `/trace` resolving a relative path against it. That stays open and belongs to whoever runs the next live race. **Amendment, 2026-08-17: the worktrees are cut off the render loop, and the room stays a room while they are.** Until now `arenaSetup` ran inline in `dispatch`, which runs inside `Update` — so for as long as git took, council drew no frame, read no key and answered nothing. The operator measured the failure the way these things are always measured, by living in it: parallel sessions against one repository, a `index.lock` held by another of them, and a room frozen with **ctrl+c unread** — the one act that could have ended the wait, sitting in a queue whose only drainer was blocked inside `git worktree add`. A room that cannot be stopped is worse than a slow one, and this is the same class of defect as §9.37's own give-up amendment: the room displayed a true thing and offered no act on it. **The setup is a command now, and the seat order is unchanged.** `arenaSetup` runs on a goroutine and reports back through a channel (`internal/council/arenasetup.go`); `dispatch` stops at the point of preparing and returns, and the turn is born later in `applyArenaSetup`, which calls the extracted `sendTurn` with what the setup measured. **The per-seat `git worktree add` calls stay SERIAL, and that is a ruling rather than an unfinished optimisation** — those adds write the repository's own refs and administrative files, so N at once contend for exactly the lock this change exists to survive, and the parallel version would turn one slow setup into N racing ones each able to fail the others. Everything stamped when the turn starts — its clock, its context, the snapshot of the previous replies — is stamped at the SPAWN rather than at the keypress, which is the honest reading of every duration the turn then renders. **The frame names the step and refuses to name the progress.** Each stage reports the words for what it is about to do — `reading the base commit`, `numbering the race`, `reading .worktreeinclude`, then `preparing worktree for codex` and `seeding worktree for codex` per seat — and the footer draws that sentence with the spinner beside it. There is no percentage, no "2 of 4" and no elapsed figure, because council cannot measure how long a checkout takes and a number it did not measure is a number it may not draw (§4a.1). The spinner is the second signal and is liveness, not progress: a step that takes a minute prints the same sentence throughout, so without a moving cell a working room and a dead one render identically — which was precisely the old lie. `TestSetupStepsNameTheWorkAndNeverTheProgress` fails on any step carrying a digit or a `%`, which is the rule stated as a test rather than as an intention. **The deadline, and the measurement behind it.** The setup carries one context with a **90 second** deadline over the WHOLE of it, enforced through `gitOutCtx` — a context-carrying sibling of `gitOut` used by the setup path and nowhere else. Every other git call council makes stays un-deadlined by construction, because `gitOut` takes no context to hand one: a diff read, a config probe or a commit killed by somebody's guess at a timeout is a worse outcome than a slow one everywhere the room is not blocking on it. One deadline over the whole setup rather than one per call, because the number an operator experiences is how long the room was unusable, and a per-call bound times five seats is a total nobody chose. The 90 is measured against, not guessed. On the reference Intel Mac (macOS 26.5.2, 2026-08-17), a five-seat setup against a synthetic repository built to this repo's own shape — 540 files in 60 directories, ~8 MB of content, against telltale's 526 tracked files and 8 MB — ran end to end in **2.3s cold and 1.3s / 1.4s warm**, worktree adds included. The deadline is therefore ~40x the measured case, and the margin is the decision rather than the number: the failure it exists for is a lock another session holds, which is unbounded by nature and says nothing about how large the repository is. What the deadline must never be is tight enough to kill a setup that would have finished, since a `git worktree add` killed mid-checkout leaves a half-created tree the operator clears by hand. **Every ending hands the room back.** A deadline hit or a git refusal ends the setup WHOLESALE — a clock is a fact about the room's patience, not about a seat, so recording it as four per-seat skips would blame four vendors for one timer and then race whatever survived as if the operator had asked for a 1-of-4 race. It lands on the room's existing arena notice, opened the way a refused race always was, and the sentence leads with the STEP before quoting git verbatim (`arena: preparing worktree for codex: fatal: …`): the git line names a lock and not which of eight calls met it, and the step is the half the operator cannot reconstruct. A process the context killed is never quoted as if git had refused — `cmd.Run` reports the signal there, and dressing "signal: killed" up as git's own sentence would be §4a.1's bug pointed at a failure. The brief returns to the composer and the room composes again, so the same enter retries it. ctrl+c stops the setup in every mode and does NOT quit, which is the keystroke the freeze ate; the trees already added are **kept and named** rather than swept up, per this section's founding ruling that worktrees live until the user deletes them, and the next race numbers itself past them anyway (`arenaRaceNumber`). Verification note: the mechanics are pinned offline against real temp repositories — the room drawing and reading keys mid-setup, the step vocabulary, the serial adds (measured, not assumed: when seat N's add is announced, seat N-1's tree already exists on disk), the deadline's wholesale stop, the failure handing the room back, ctrl+c, a stopped setup's messages being dropped by comparison, and the rendered frame. **The live half is owed**: no race has been run against this build, so the claim that a real held `index.lock` now ends in a notice instead of a freeze rests on the tests and on an expired-deadline fixture, not on the lock that started this. **Amendment, 2026-08-29: the brief also arrives as `AGENTS.md`, for the seats that were measured reading one.** The candidate (competitor sweep 2026-08-18) proposed AGENTS.md as the one cross-vendor context channel needing no per-vendor prompt plumbing, on the strength of agents.md's own "read natively by 20+ tools". That is a docs claim, and ADR-001 does not accept docs claims about vendor behavior. **The measurement came first, and the build was conditional on it**: one headless probe per vendor CLI on this box, from a scratch directory whose only content was an `AGENTS.md` naming a codename nothing else on the machine knew, asked for the codename and nothing else. | seat | version | result | | --- | --- | --- | | codex | codex-cli 0.149.1 | **answered `ZEPHYR-9`, no tool call** — the file reached the model as context | | grok | grok 1.0.5 | **answered `ZEPHYR-9`, no tool call**, and named its source on the wire: *"From the always_applied_workspace_rules, the Agents.md file says"* | | claude | Claude Code 2.1.251 | **answered `ZEPHYR-9` by going to look** — both trials ran `ls -la` then `cat`, recorded in the probe sessions' own transcripts | | agy | 1.1.25 | **reads it by going to look**, once `--add-dir` names the tree ([§9.59](#s9-59), 2026-09-03): asked only to create a file, it ran `list_dir` then `view_file AGENTS.md` before writing, and then wrote the file the brief in AGENTS.md asked for. The claude row's fact, not the codex row's. The codename probe itself was not run | | cursor | — | **unmeasured** — this seat races over ACP on a throwaway session, and no probe of that path ran | Two seats demonstrably ingest the file unprompted, which is the bar the sweep set, so the feature is built. The claude row is deliberately NOT counted as the same fact: it is a real read of a real file in the cwd, in a directory holding exactly one file — the easiest possible discovery — and nothing here claims that seat auto-loads AGENTS.md. Whether the answer was shaped by the operator's global `CLAUDE.md` load order cannot be separated out by this probe either; what the transcripts DO show is that the words came from the file on disk, because the model went and read it before answering. The ruling that follows from that table is what shapes the feature: **council writes the file for every racer and claims it for none.** No column, no notice and no snapshot field says a seat was briefed via `AGENTS.md`, because the room cannot tell per race which seats ingested it — and two of five are unmeasured. Writing it costs a seat nothing; claiming it would be the room narrating a fact nobody measured (§4a.1). The file is offered exactly the way the worktree is. - **Identical for every seat, by construction.** Marker, `arenaConduct`, then the brief — the same bytes in all five trees. `arenaBriefText` takes no seat parameter at all, so the per-seat constraint text the candidate pitched is unrepresentable rather than merely discouraged: it collides with `arenaConduct`'s standing position that the room's added words are a CONSTANT so the cross-seat comparison stays undisturbed, and a later change that wants divergence has to argue for it here. - **The attempt's receipt stays the racer's.** Council's file would otherwise land in the stat through `git add -N .` — §9.37's own "lying diff" known limit, arriving from the other side. The three reads that could pick it up (the finish-time `collectArena`, the live `collectArenaStat`, and `commitArena`'s stage plus its dirty check) append one pathspec, `:(exclude)AGENTS.md`, measured on git 2.55.0.windows.3: the file stays untracked, stays out of the commit, and a tree holding nothing but it still reports clean — which is what keeps the empty-commit ruling working. **`/adopt` therefore merges a branch that never held the file**, and the operator's repo cannot acquire a stray `AGENTS.md` from a race. - **The lifecycle verbs read the racer's tree too, and all three reads were wrong until they carried the same pathspec.** This was found by building the feature, not by reasoning about it, and each one is a different bug: `/adopt`'s arming read (`lifecycle.go`) would have offered to adopt a seat that changed nothing, because council's file made a clean tree look dirty; `/adopt`'s OWN commit — the one it makes for a racer whose work never reached `commitArena`, which is the give-up path race t9 exercised twice — would have staged the file with `add -A` and merged it into the operator's repo; and `/arena drop`'s refusal would have named council's write as the operator's uncommitted work and demanded the `!` spelling on every clean attempt. - **`/arena drop` takes council's file back before git sees the tree.** `git worktree remove` counts an untracked file as a dirty worktree and refuses, so the pathspec alone was not enough: an ordinary drop failed at git. `removeArenaBrief` deletes the file only while the marker still stands, so a racer's own `AGENTS.md` keeps the refusal it has earned. That is the ONLY deletion — the worktree is kept until the user drops it, and until then the file is the visible record of what that seat was told. - **The exclusion is conditional on the marker, re-read per call.** `arenaBriefArgs` opens the file and checks it still starts with council's marker. A racer that REPLACED it authored a file, and it appears in the stat like any other; a file council never wrote is never excluded. A stale flag recorded at setup would have hidden that authorship. - **Council never overwrites an `AGENTS.md` the checkout or `.worktreeinclude` seeding already put in the tree.** In a repo that ships one, no racer gets council's copy, every seat reads the repository's own instructions identically, and the comparison is as uniform as it was before this existed. The pathspec is off there too, so a racer's edit to the repo's own file is in the diff. - **A write that fails skips that seat**, named on its column through the existing `seatErr` channel — the `.worktreeinclude` rule, applied for the `.worktreeinclude` reason: a tree the room KNOWS holds a different brief from its siblings races a different question, and that is not the comparison the operator opened. Nothing new is rendered for it. Known limit, stated rather than hidden: while the marker stands, a racer that APPENDS to council's `AGENTS.md` is excluded from its own diff on that path. A racer editing the room's brief file is not an answer to the brief, and the alternative — an exclusion that lapses on the first stray edit — would drop council's own file into every stat instead. Not built, and not by omission: `telltale doctor` reporting whether a repo carries an `AGENTS.md` was part of the same candidate. It is a different surface with a different reader and it is left for its own change. Verification note, on this section's own terms: the mechanics — the identical file in every tree, the file never reaching the stat, the patch or the commit, the zero-diff attempt staying a measured zero with council's file in its tree, the racer-authored file NOT being hidden, the repository's own file being left alone, the skip on a failed write, the ended-context stop, and all three lifecycle reads (`/adopt` still refusing a brief-only racer, `/adopt` not merging the file, `/arena drop` needing no force) — are pinned by offline tests against real temp repositories (`arenabrief_test.go`), and no test spawns a vendor. **The live half is owed**: the per-vendor probes above were run headlessly in a scratch directory, not inside a racer worktree during a real `/arena`, so no live race has yet watched a seat act on this file. **Amendment, 2026-08-29: `/adopt` says what it is about to merge INTO, before you say y.** The card named the act — the branch it cuts and the exact `git merge --no-ff` it runs — and named nothing about the room the merge lands in. Everything the operator needed in order to weigh the `y` was in a second terminal: how far the racer's branch had drifted from the room, what had landed in the room since the race was cut, and whether the two had written the same files. So the answer was "yes because I trust it" or "no because I don't", which is §9.41's finding about the room's *other* gate, arriving a second time at the one gate that merges. **The card now leads with measured git state and then names the act.** ``` adopt codex? vs main: 1 ahead, 1 behind · 1 overlapping path (a.txt) · y cuts adopt/t4-codex and runs git merge --no-ff arena/t4/codex · n cancels ``` - **Every count carries its baseline, and the baseline is the room's own HEAD** — because that is the commit `/adopt` cuts the adopt branch from, so it is genuinely what the merge lands in. It is named as the branch when one is checked out (`vs main`), as the short commit on a detached HEAD, and as `vs the room's HEAD` when git could not answer at all. The clause is never dropped: a bare `1 ahead` is a number with no question attached. `behind` is the half the operator had no other way to see, and it is the whole point of the line — it is everything that landed in the room while the race sat there, including an earlier adoption from the same race. - **One `rev-list --left-right --count HEAD...` answers both counts**, and `ahead` is the same figure `unadoptedCount` was already reading for the zero-change refusal, so the preview costs the card one git call rather than two. Measured at git 2.55.0.windows.3: the left count is what only HEAD holds and the right is what only the branch holds. - **"Overlap" is a read; "conflict" would be a claim.** The overlapping paths are the intersection of two `diff --name-only` reads over the same merge base — `HEAD...` is the incoming half git actually applies, `...HEAD` is the room's own half — so the card states that both sides wrote a path and stops there. A repository can overlap on a path and merge cleanly. The word "conflict" belongs to a merge that ran, and the reactive path below still owns it. - **Three overlap states, kept apart (§4a.1).** A read that returned nothing renders `no overlapping path`; a read that returned paths renders the count and names the first; a read that failed renders `the overlap check could not run:` with git's own line. An unreadable ref never renders as a clean one. The counts and the overlap fail differently on purpose: the counts are load-bearing, so a failed read refuses the whole command by name, exactly as the older `unadoptedCount` call did; the overlap is advisory, so a failed read degrades to its own sentence and the card still arms. A broken preview must not brick a verb (`arenaRaceNumber`'s rule). - **The preview states its own limit rather than leaving it to be discovered.** Every figure is read off COMMITTED state, so a racer whose worktree is still dirty has work none of the figures cover — and that card adds `these counts exclude 1 uncommitted path` beside the clause that already says `y commits its worktree`. `TestAdoptConflictAbortsCleanly` is exactly that case: an uncommitted racer edit conflicts against a room commit while the overlap read correctly reports nothing shared. Without the clause, `no overlapping path` would be read as a promise about the merge. **Two shapes recorded rather than taken.** - **`git merge-tree --write-tree`, which computes a REAL merge result.** It would let the card say "conflict" honestly. It is not here because the claim would need a live measurement at a pinned version on this box before it could ship (this section's own rule), it needs git ≥2.38, and it writes objects into the repository — which puts a preview on the write side of a room whose posture is offer, never take. The reactive abort already owns the real merge result, and it owns it after the operator asked for one. - **Folding the racer's uncommitted paths into the overlap set**, by parsing `git status --porcelain`. It would close the limit named above, and it was declined because those paths are a prediction of a commit nobody has made yet — council reading a tree to guess what a future commit will contain, where §4a.1 asks it to read what exists. The exclusion clause states the gap instead. **The preview leads the line, and the cost of that is stated.** The notice truncates from the right at a narrow width, so leading with the measured state can cost the action clause its tail — and the action clause is the older contract. It leads anyway: an operator who can read only the first clause can still press `n`, and the preview is what makes that `n` a decision rather than a mood. The alternative, recorded and not taken, is a second sheddable cell on the status line (the mechanism `st.ArenaSetup` already uses), which would drop the preview whole instead of slicing it — a new render surface for one notice, in a file this change otherwise does not touch. Verification, on this section's own terms: the mechanics are pinned by offline tests against real temp repositories (`lifecycle_test.go`) — the counts against an unmoved room and against one that moved, the named overlapping path, the uncommitted exclusion, the two overlap failure states held apart, and the baseline on a named branch, on a detached HEAD and on no answer at all. No golden moved, because the card is a notice string and no golden renders one. **The live half is owed**: no real `/adopt` has been armed against this build, so every sentence above rests on the fixtures rather than on a race. **Amendment, 2026-08-29: `/adopt` can take the winner plus the parts of the runner-up you point at.** A race ends with four attempts and one decision, and the decision the room offered was all-or-nothing: adopt one seat whole, and retype by hand whatever the runner-up got right. The ask is not ours — it is the one users put to Cursor's own multi-agent judging thread, in those words: synthesize a best-of-both instead of picking one wholesale. No surveyed tool ships it. Council already owns the substrate — per-attempt worktrees, one base SHA, commit receipts, and a y/n card that names exact git commands — so this is a grammar and four refusals, not new machinery. **The grammar is one more argument, and the fork is PER-PATH.** ``` /adopt claude +codex internal/council/helper.go docs/council.md ``` - **Per-path, not per-hunk.** The sweep's own evidence is users asking to mix at path OR hunk level, so the choice was open. Per-hunk needs an interactive picker inside the room — a full-frame body with its own scroll, its own keys and its own mode word — which is a new render surface for a v1 whose value is that the operator can take one file from the runner-up. A path is also the unit already in front of them: the column's `git diff --stat`, and this card's own overlap clause, both speak in paths. **Per-hunk is deferred, not rejected**, and this shape does not block it: a hunk picker would narrow what `+` contributes and leave the grammar alone. - **`+` glued to the donor seat.** A bare `+` as its own word would make `/adopt claude + codex` legal, and that reads as a request for two whole attempts — which this verb cannot do and must not appear to offer. `/adopt` already takes its whole argument (roomcmd's `parseCommand`), so the longer form needs none of the vocabulary handling `/arena drop` needed. - **User-typed, never computed.** §9.34 rejected a synthesis hop, and that ruling binds here: council applies the paths the operator named and chooses nothing. The refusals below are how that promise is kept mechanically rather than by intention. **Four refusals, and not one of them resolves anything.** Each names the path and a way forward (§9.17's tell), and each fires before the card arms, so a `y` is always one that can be honored: - **A path BOTH racers wrote.** This is the founding refusal. `git checkout -- ` would discard the base attempt's answer with no merge and no conflict marker, so council refuses by name and the operator decides — drop the path, or merge it by hand afterwards. - **A path the ROOM wrote since the race was cut.** The same silent clobber one level out: the merge machinery never sees a path taken by checkout, so the room's own work there would vanish. - **A path the donor did not write.** Taking it would land the base attempt's own content under a receipt saying it came from the donor. - **A path the donor deleted.** A hybrid takes files a racer wrote, never a deletion — a stated v1 limit rather than a `git checkout` pathspec error discovered after the branch was already cut. **The card says exactly what will be merged from where, composed with the divergence preview above.** The preview still leads, for that amendment's reason; the leading question gains the hybrid's own scope, and the action clause gains its second half: ``` adopt claude + 1 path from codex? vs main: 1 ahead, 0 behind · no overlapping path · y commits both worktrees, cuts adopt/t4-claude+codex and runs git merge --no-ff arena/t4/claude, then takes helper.go from arena/t4/codex · n cancels ``` Every path is named rather than counted-with-an-example. The count-plus-first grammar the overlap clause uses is right for a measurement the room took; these paths are the SCOPE the `y` authorizes, and a card that authorized "2 paths (a.txt)" would leave the second one unread. The operator typed them, so the list is short by construction. **The branch name carries both seats: `adopt/t-+`.** This is a naming decision and the arena record (§9.47) forced it, because that page derives everything it knows from these refs. The alternative — keep `adopt/t-` and let the commit message carry the donor — would leave one seat's name alone on a branch holding another seat's work, in the one place `git branch` shows a reader, and the record would then count the base seat as having won the race outright. `+` is the joiner because it is legal in a ref name, because `-` is already the collision suffix and because `/` is already the arena namespace. `freeAdoptBranch` suffixes a collider identically, from the same single scan, so the two spellings cannot disagree about what "taken" means. **The receipt names both sources.** The base arrives as the unchanged `git merge --no-ff`, and the paths arrive in a second commit whose message names both arena branches, lists every path, and says what council refused to do: ``` adopt race t4: arena/t4/claude whole, plus 1 path from arena/t4/codex the merge below this commit carries arena/t4/claude whole. this commit adds the paths that came from arena/t4/codex, and it adds nothing else: helper.go telltale council took no path that both seats wrote. a shared path is refused by name, and the operator merges it. ``` The notice says it a third time, because that is the last moment the operator is still looking: `adopted claude onto adopt/t4-claude+codex, with 1 path from codex (helper.go)`. **The arena record renders a hybrid as its OWN state, and credits nobody.** The refs can say a race was decided and which two seats the adoption was cut from; they cannot say which paths came from where, because that lives in a commit message the page does not read. So a hybrid raises a fourth per-seat count and moves no rate at all: the race counts as decided, both seats count as having entered it, and neither seat's `adopted of decided` moves. Crediting the base seat would score it for work the donor wrote; counting it against both would score two seats down for a race the operator resolved in both their favour. A seat whose only decided races were hybrids reads `no attempt adopted whole part of 2 hybrid adopts`, which is the true statement — and the window sentence carries the difference a reader adding the seats up would otherwise not find: `3 decided by you (2 by a hybrid adopt, counted for no seat)`. A whole adoption of the same race still outranks a hybrid of it, because an operator who adopted whole, reverted and then took a hybrid did adopt it whole once, and both receipts survive. **One fork from the divergence-preview ruling above, recorded because it is a fork.** That ruling declined to fold a racer's uncommitted paths into the OVERLAP set, on the grounds that they are a prediction of a commit nobody has made. The hybrid's path checks DO read them, and the difference is what the read is for. There it was a preview of a merge RESULT; here it decides a refusal about paths the operator named, and `y` commits both worktrees in the same act with `git add -A` — so the set is `tracked changes ∪ untracked-and-not-ignored`, which is what that commit will contain by definition rather than by forecast. Refusing to read them would refuse every hybrid on an ordinary race, because arena seats leave their work uncommitted (commit-per-turn is deferred). The card still states the older ruling's limit, in the same clause it already used. **A conflicted hybrid restores exactly like the conservative whole adopt.** The base merge is unchanged, so a conflict aborts, the branch is deleted, the room goes back to the branch it came from, and the donor's paths are never written. A failure in the second half restores the same way, with one extra step named rather than hidden: a `git reset --hard` before the checkout back. It is bounded to a branch council cut, at a commit council made, over a room tree measured clean before any of it — the only content it can discard is a half-checked-out copy of files that exist whole on the donor's own branch. **Two shapes recorded rather than taken.** A `+` with no paths, meaning "take everything of the donor's that does not collide" — declined because it makes the scope a thing council computed rather than a thing the operator read, which is the whole contract of the card. And renaming the donor's paths on the way in — declined as a second grammar to learn, for a case `git mv` already handles after the adoption. Verification, on this section's own terms: the mechanics are pinned by offline tests against real temp repositories (`hybrid_test.go`) — the merge plus the named path landing while the donor's other file does not, the receipt naming both branches, all four refusals, the grammar's own refusals, the collision suffix, a committed donor read off its branch instead of its tree, and a conflicted hybrid restoring the room. The record half is pinned in `record_test.go` against ref lists, with its own golden in both glyph sets. **The live half is owed**: no real hybrid has been armed against a race on this box, and it is owed on the same keystroke as the divergence preview's live debt above. ### 9.38 paste lands whole, and never sends (2026-08-09) The ask, in the operator's words: *"how i can paste things into the area i can type in."* The answer required measuring what a paste even was in this room, because the two obvious guesses — it works, or it fires a send per pasted line — were both wrong. **What was measured about today's behaviour.** All of it read off the pinned module source, not vendor docs. bubbletea v2.0.8 enables bracketed paste unless a view opts out (`cursed_renderer.go` writes `SetModeBracketedPaste`; council's `View()` never sets `DisableBracketedPasteMode`), and ultraviolet's terminal reader buffers everything between the paste markers into ONE `PasteEvent` — a newline inside the paste lands as `\n` in its content, never as an Enter keypress, and the win32-input-mode encoding Windows Terminal uses is decoded into the same buffer (`terminal_reader.go`, ultraviolet pinned at v0.0.0-20260703014108). So in a bracketed-paste terminal a paste could never have fired a send. What it did instead was NOTHING: council's `Update` had no `PasteMsg` case, the message fell through the type switch, and the clipboard's offer was silently discarded. The composer never learned a paste happened. The fires-sends failure is real on exactly one path: a terminal with NO bracketed paste replays a paste as keystrokes, each pasted newline arrives as an Enter keypress, and compose mode's enter dispatches — a five-line paste is up to five turns, each to live vendor CLIs. Council cannot distinguish that replay from typing without a timing heuristic, which would be inferred behaviour, and this product does not ship inferred behaviour (§4a.1). So that path is left as it is and named here instead: the text chunks are flattened safely by `sanitizeKeepingSpace`, the enters are enters, and the fix on such a terminal is the terminal. Windows Terminal — the reference environment — brackets its pastes. **What was built** (`paste.go`): the room's half of the contract the runtime already offers. - **One `PasteMsg` case in `Update`.** The content goes into the draft and nowhere else — a paste never dispatches, never answers a gate, never quits. Enter, a keystroke from a person, remains the only way a brief leaves the room. Pasted control characters cannot act: a pasted `\x03` is not ctrl+c, a pasted `q` is the letter q; controls without width are dropped. - **The multiline ruling: newlines are PRESERVED, raw.** The composer has been a block since ctrl+j existed (`State.Draft` may hold newlines; `wrap()` honours them; the compose area grows to `maxComposerRows` and elides with "N more above"), so there is no single-line prompt to protect and no need for a `⏎` display glyph — a pasted paragraph renders as the rows it is, in both glyph sets, and dispatch hands the vendors the draft with its real newlines. The string on screen is the string sent (§7.14). CRLF collapses to `\n` (the Windows clipboard's line ending; splitting it would gift every line a trailing space). The one lossy rewrite is stated rather than hidden: a tab becomes one space, because a cell grid cannot budget a tab and a guessed tab stop would be fidelity theatre. - **A cap with a named refusal.** `maxPasteRunes` (8,192, over draft-plus-paste) refuses atomically — nothing lands, not a truncated prefix — and the notice carries both numbers and the remedy: *"paste refused: 20481 chars against the composer's 8192 — put long text in a file and name the path in the brief."* The number is anchored to the narrowest pipe a brief must fit through (the Antigravity seat's prompt rides argv; Windows caps a command line at 32,767 UTF-16 units; 8,192 runes is at most half that even all-surrogate-pair) and to the point where a footer composer that deletes rune-by-rune stops being an editor. - **A paste from view mode inserts and opens compose.** A paste is not a keystroke — view mode's letters are commands because they are keys; pasted text can only be material, and the only place material goes is the draft. The mode line states the switch on the next frame. The exception is a pending y/n (tool gate, `c`, `/write`, a flow write hop): the paste is refused by name and the question stays exactly where it was — nothing about a pending request happens implicitly. No golden changed and none was added: a pasted draft produces the same `State` shape ctrl+j already produces, and a golden that did not change is the claim that the room's appearance did not either. The tests (`paste_test.go`) drive `Update` with the real message shape and assert the observables the flow security tests trust — spawn count, draft, pending flags — end to end through enter, which must deliver the pasted newlines to the seat intact. **The live verification owed.** No test in this container can observe Windows Terminal bracketing a paste — that is the terminal's half of the contract. The check, one minute at the real machine: open `telltale council` in Windows Terminal, copy a three-line snippet, paste into the room. Expected: one insertion, three rows in the composer, zero dispatches; then enter sends it as one brief. If the paste instead lands as separate turns, the terminal did not bracket it — record the terminal build in PARITY.md, because that is a measured vendor fact, not a council bug. **Paid, 2026-08-13, on Windows Terminal 1.24.11911.0.** A three-line snippet pasted into a live room landed as ONE insertion, three composer rows, and **zero dispatches**; `ctrl+u` then cleared it and reported the count, which is the second half of the same gesture and is why the clear is evidence too — a paste that had dispatched would have left nothing to clear. The terminal build is named because the bracketing is the terminal's half of the contract and a version is the only thing that claim can be pinned to; PARITY.md stays out of it, since that file records a machine BEHAVING DIFFERENTLY and this machine behaved as specified. **One half of the check is still unexercised**, stated rather than rounded up: the draft was cleared instead of sent, so "enter sends it as one brief" has not been observed on a pasted multi-line draft. The property that was owed — a paste never sends — is the one that was measured. **The other half paid, 2026-08-15/16, and it found a render bug on the way.** A three-line brief was pasted and SENT with enter in a live 5/5 room: ONE dispatch, the turn counter moved by one, and the newlines reached the seat as bytes — verified in the vendor's own session transcript (`now.\nThen add a second one…`), not read off the wrapped column echo, because the echo's row breaks are ambiguous between newlines and word wrap. **What did not hold was the composer render: the three-line draft drew as ONE row before enter.** The 2026-08-13 check drew three rows in a one-seat room on the same terminal build (1.24.11911.0), so the suspect is the composer's height behavior under a full seat strip, not the paste path — the wire is proven right and the drawing is proven wrong, which is the exact split this section exists to keep. The render defect is recorded as an unowned gap in STATE.md; this section's claim is amended to say the paste property holds ON THE WIRE, with the row rendering owned separately. **Amendment, 2026-08-16 — the render was measured, and the seat-count suspect is REFUTED.** The bullet above named a suspect: the composer's height behavior under a full seat strip. That suspect is wrong. `TestAMultilineDraftNeverCollapsesSilently` sweeps the drawn geometry with a three-line draft and asserts that every line is on screen, or that the frame says how many are not. The sweep covers one, three and five seats, widths from `MinWidth` to 240, heights from `MinHeight` to 40, both glyph sets, and the expanded projection. All 588 combinations draw all three rows. `composer-multirow-five-seats.txt` is the 5/5 room at 200x24 as bytes. **The seat count changes nothing about the compose area**, and the reason is structural: the composer is full-width chrome, so `composerRows` wraps against `promptWidth(st.Width)` and never against a column width. Five columns narrow the COLUMNS. They do not narrow the composer. Two further things were measured, because a refutation is only useful when it also closes the paths it rules out. - **The one-row compose area is the only silent collapse in the render, and it is unreachable.** `composerLines` flattens newlines to spaces when `lay.Prompt == 1`, with no marker. Its comment calls that unreachable for a real draft. The claim holds. `resolveLayoutIn` clamps `Prompt` to 1 only below a height of 9, and `Render` refuses to draw a room under 60x10 at all — it prints `council needs 60x10 (have WxH)` and no compose area. So no drawn frame reaches the flatten. The branch stays as it is, and this paragraph is the measurement that says why. **This bullet is WRONG and the 2026-08-17 amendment below replaces it.** It is kept as written because the amendment is about how a sweep can prove the wrong thing, and the sentence that did it is the evidence. - **The paste transport carries the newlines, on this platform's own path.** ultraviolet, pinned at v0.0.0-20260703014108, buffers a bracketed paste into ONE `PasteEvent`. The win32-input-mode path Windows Terminal uses converts a pasted Enter record to a literal `\n` inside that same buffer (`terminal_reader.go`, the `isWin32 && event.Code == KeyEnter` case) rather than emitting a key event. bubbletea v2 maps that event to `tea.PasteMsg`, which `Update` routes to `paste()`. `sanitizePaste` keeps the newlines and `setDraft` stores them. `TestAWindowsPasteKeepsItsLinesAndLosesItsCRs` already drives `Update` with LF, CRLF and bare-CR content and pins the draft that results, so this half needed no new test. **What is left, stated rather than rounded up.** The live observation is not explained. The draft held newlines on the wire, the render draws newlines at every geometry it will draw, and the two cannot both be true of one frame. The leading remaining hypothesis is an observation error at the composer, and §9.38 already names the trap that produces one: the echo's row breaks are ambiguous between a newline and a word wrap. That warning was applied to the column echo and the wire was checked against the transcript because of it. The composer's own row count was still read by eye. **The decisive next measurement is cheap: reproduce the paste in a live 5/5 room and record the terminal's rows and columns with it.** A geometry at or above 60x10 makes the render correct by the sweep above, which moves the defect off this section entirely. A geometry below it means the room drew the floor refusal, and the report is about a different frame than the one assumed. **Amendment, 2026-08-17 — "correct at 60x10 by the sweep" was not true, and the sweep is why.** The sentence above says a geometry at or above 60x10 makes the render correct. It does not, and neither does the bullet two paragraphs up that calls the one-row flatten unreachable. Both rest on the same sweep, and the sweep could not see the cell they were claiming. **What the sweep pinned false.** `TestAMultilineDraftNeverCollapsesSilently` built every one of its 588 cells from `fiveSeats()`, and that fixture carries no pending gate and no collapsed seat. Those are the room's two chrome rows, and `resolveLayoutIn` spends both out of the SAME budget the compose area is measured from — the needs-you strip (§9.40) first, because it does not yield, then the collapsed-seat notice. Pinning both absent held the compose area two rows taller than a real room's, in all 588 cells at once. The sweep was not measuring a geometry the frame reaches. It was measuring a fixture, and the fixture was generous in exactly the dimension under test. **The cell it could not see, and it is an ordinary room.** At the 60x10 floor with a tab bar, both chrome rows on screen, `resolveLayoutIn` leaves the compose area one row: `Prompt` is clamped to `Height - rows - 1`, which is 1, while the draft still holds three. Nothing exotic builds that room — a five-seat machine with one vendor not installed, at the smallest terminal council agrees to draw, the moment a gate goes up. **Forty cells of the widened sweep land there**: heights at `MinHeight`, three and five seats, every width from 60 to 240 once `Expanded` forces the tabs tier, both glyph sets. In every one of them the old code flattened the three typed rows into one row of prose joined by spaces, with no marker and, at that width, nothing even elided. **And a flatten is worse than a clip, which is why no count could have fixed it alone.** Clipping drops text and the marker vocabulary describes exactly that. Flattening dropped nothing — it RESTATED the draft, and §7.14's promise is that the string on screen is the string sent. Three typed lines and one long line are different briefs; the room drew the second and the wire sent the first. So the widened sweep asserts two things now, not one: no draft line is dropped without the frame saying how many, and **no two typed rows are ever drawn welded into one**. The weld check is what fails on the old code; the drop check passes there, because at 60 cells the flattened draft still fit. **What the marker shows.** `composerLines` no longer flattens. When the compose area is one row and the draft wants more, the row carries the marker and the TAIL, in that order and at two intensities: `↑ 2 more above third line_`. The words and the count are `moreAbove`, which is the column overflow marker's own spelling (`overflowMarker`, §9.10) and the one the multi-row composer already spent a whole row on — a reader who learned `↑ 36 more above` on a column is not taught a second vocabulary here. The count is rows not drawn, so it agrees with the multi-row path's own arithmetic. The tail stays because that is where the cursor is, elided from the left when the remaining width will not hold it, which is the rule this branch was already built on. The separator is three cells rather than the room's `│`: the composer is a box and its sides are that same glyph (§9.44), so a bar inside it would read as a column rail through the frame's own edge. Three cells is `needsYouGap`'s answer to the same question. Nothing about the frame's HEIGHT changed, deliberately. A floor of two rows on the compose area would have closed the same gap, and it would have moved every golden in the package and taken a body row from a room already at its floor — a room you can type in but not read is the trade §9.44's budget refuses. The marker costs nothing but the row it was already drawing. The goldens are `composer-clipped-to-one-row.txt` and its `-ascii` partner, at 60x10, so the claim is bytes in both glyph sets; they render `PlainStyles`, which is the NO_COLOR half. No existing golden moved — a draft that fits its compose area reaches none of this, and that is every room above the floor. **Amendment, 2026-08-17 — the decisive measurement was taken, and the composer drew THREE rows.** "What is left, stated rather than rounded up" above named one cheap next step: reproduce the paste in a live 5/5 room and record the terminal's rows and columns with it. The operator ran it. A three-line brief went into a live 5-of-5 room in Windows Terminal at a recorded geometry of **170 columns x 54 rows**, and the composer drew **three rows**. The geometry was recorded this time rather than read back off the frame, which is the whole reason the step was specified that way. **So the live render agrees with the sweep (#253), and the t9 one-row observation stands UNREPRODUCED.** The render is now verified twice by two different instruments — 588 synthesized geometries and one live room — and the wire was never in doubt. Against that, one unrepeated eyeball reading. The finding is recorded as a **probable observation error**, and this section already documents the trap that produces one: the echo's row breaks are ambiguous between a newline and a word wrap. That warning was applied to the column echo at the time and the wire was checked against the vendor's transcript because of it. The composer's own row count was the one thing still read by eye, and it is the one thing that did not reproduce. **What this run does NOT establish, because the geometry is the generous kind.** 170x54 sits inside the swept band on width (60 to 240) and ABOVE it on height (`MinHeight` to 40), and above means MORE compose room, never less. It is nowhere near the 60x10 floor where the amendment above found forty cells that used to flatten. At 170x54 the pre-marker code and the marker code draw the same three rows, so this run cannot tell them apart and is not evidence about either. It closes the reported defect and it closes nothing else. The one-row branch is still exercised only by `composer-clipped-to-one-row.txt` and its ascii partner, which is where that claim belongs. **One independent corroboration of the column figure, and only the column figure.** The agy statusline capture taken on the same machine the same evening (§3.8's re-capture block) carried `terminal_width: 170` on all fifteen fires. That is a second instrument reading the same terminal width, which is worth one sentence because the geometry is the evidence here. It says nothing about the row count, which no payload carries. **Amendment, 2026-08-09 — the ergonomic other half: ctrl+u clears the draft.** Paste changed the arithmetic on regret. A draft used to cost at most a typed sentence, so backspace's rune-at-a-time delete was proportionate to any mistake the composer could hold; one wrong paste is now up to 8,192 runes in a single gesture, and 8,192 backspaces is not an editor. `ctrl+u` — readline's own kill-line — empties the composer in one keystroke, in compose mode only. A chord on purpose: `sanitizePaste` drops every control character, so no paste can carry the key into the room, and no stray letter can fire it — the one gesture that can empty a draft is a deliberate hand on ctrl, the same argument that keeps a pasted `\x03` from cancelling. No y/n gate, unlike `c` and `u`, whose confirms price drops nothing can reverse: a cleared draft's ways back are ordinary (the clipboard still holds a paste; a sentence re-types), so the loss is *stated* instead of priced — the notice carries the measured rune count of the string just dropped ("draft cleared — 1204 chars"), in the paste refusal's own unit and spelling, never an estimate (§4a.1). An empty draft clears silently — backspace's own empty-draft behaviour applied at size: after the press the state the key promises is already on screen, and an every-press "nothing to clear" would put noise where dispatch answers land. The key sits below every pending gate in `key()`'s routing, so a stray ctrl+u under a y/n gets that gate's standing stray-key answer (cancel, or the question restated) and never reaches the draft; in view mode it does nothing at all, because esc parked the draft there under the promise "keeping the draft", and a chord that revoked it from the other mode would make esc unsafe in hindsight. Nothing but the draft moves — the routing indicator falls with it only because it describes it. Taught on the help panel's compose-keys row (`ctrl+j/u/esc`, landed inside that row's 114-cell budget; the words "compose" and "the draft" paid), deliberately not on the compose mode line, which is at its own width budget. Tests: `draftclear_test.go` drives `Update` with the real chord through every gate and both modes. ### 9.39 a fifth seat, and the first that reports money (2026-08-09) The ask, in the operator's words: *"grok 30 dollar subscription paid for. create a seat for grok in council."* The seat is built and the invocation is verified end to end; what follows is what had to be measured to earn each claim on the column, including two flags that were refuted and one hazard the seat ships with because nothing in the CLI can close it. **The vendor.** grok 1.0.0 (3cd0d0cbce), signed in against grok.com — the subscription, not an API key — with `grok models` reporting one model, `grok-4.5`. It is a spawn-per-turn batch program like Codex and Antigravity, not a live process like the Cursor seat: `-p/--single` takes a prompt, answers, and exits. So it implements `Vendor` and neither `Persistent` nor `Conversational`, and its column takes the same spawn-per-turn shape those two already have. **The invocation, and what it deliberately does not carry.** `--output-format streaming-json` and nothing else, with the prompt as the value of a trailing `-p`. argv rather than stdin, and unlike Codex that is not a preference — there is no `-` sentinel and no stdin channel for a prompt at all. It is safe for the reason the Antigravity seat's argv transport is safe: the installer drops a native `grok.exe`, not a `.cmd`, so `classify()` never reaches the shim refusal and no `cmd.exe` ever sees a brief. Two flags whose names promise containment were probed, and NEITHER is passed: - **`--permission-mode plan` was refuted, with the write landing.** Asked to create a file under it, the seat called its `write` tool, reported the call `completed`, said so in prose, and the file was on disk afterwards. The control run without the flag wrote its file too, via `search_replace`. The only difference observed between the arms was which write tool the model picked. This is the Antigravity ledger a second time (ADR-008, seventeenth amendment) and it lands the same way. - **`--sandbox` is worse than refuted — it is unobservable.** `grok --sandbox bogus-profile-xyz -p "hi"` does not error, does not warn, and answers normally with exit 0. A flag that silently accepts a profile name that cannot exist gives council no way to tell a real profile from a typo, so asking for one would put a word in the badge backed by a value the CLI may never have read. Feeding a vendor a deliberately INVALID value, rather than a plausible one, is what turned "unverified" into "unobservable" here, and it is the cheapest probe in this document. So the badge is `unsandboxed`, and its detail says both of those things in the vendor's own terms. Also deliberately absent: `--always-approve` and `--permission-mode dontAsk / bypassPermissions`. The default headless mode already writes without asking — measured, above — so an approve-everything flag would buy nothing in exchange for the badge cost ADR-008's fifth and seventh amendments attach to that whole class. **The first seat that reports money, and why that is the honesty rule working rather than bending.** grok's `end` event carries a `total_cost_usd` it computed itself (`0.0407676` on the first captured turn, with a `modelUsage` breakdown beside it). Codex and Antigravity report token counts and no dollar figure, which is why both adapters leave `CostUSD` nil forever — a cost derived from tokens and a remembered price is exactly the invented number section 4a.1 forbids. The constraint was never "council does not show cost"; it was "council does not INVENT cost". This figure is read, so it passes through untouched, and `TestGrokEndCarriesThreadAndTheVendorsOwnCost` asserts the exact captured value so that any future rounding or unit conversion fails there. The pointer matters and is not defensive: a captured turn ends with no `usage` and no cost keys at all (see the slash hazard below), so "reported nothing" and "reported zero" are both real states of this field on this vendor. `TestGrokAbsentCostStaysAbsent` pins it. **Streaming, and the one judgement call.** The deltas are genuinely token-level — `"I'll"`, `" read"`, `"notes"`, `".txt"`. That is finer than the ~80-character chunks section 9.7 flagged as overstating "tokens" on the Claude seat, and finer than the ~95-character ACP chunks section 9.36 measured on the Cursor seat, so this column carries `GranTokens` on the strongest evidence in the room. The judgement call is `thought`, which is dropped. It is the model reasoning rather than answering — 46 lines against 14 of `text` on the first capture, opening "The user wants me to read notes.txt" — and routing it to the column would put private deliberation where the answer goes, in a room built to compare answers. It is the line `codex.go` already draws when it excludes `reasoning` items. The cost is stated rather than hidden: a turn that thinks for a long time before speaking shows an empty column while it thinks. **A hazard that ships, because no invocation closes it.** A brief whose first non-space character is `/` is eaten by grok's own slash-command parser and never reaches the model. The turn is not an error — `available_commands`, then an `end` with no usage, no cost and no text, exit 0. On screen: a column that finishes instantly with nothing in it. Three channels were tried and all three were eaten: `-p "/context"`, `--verbatim -p "/context"` (whose help text reads "Send the prompt exactly as given"), and `--prompt-json` with the text as a content block. The third had a CONTROL — the same `--prompt-json` invocation with a non-slash prompt answered normally — which is what makes this a property of the parser rather than a guess about a flag that might not work. The room mostly protects this seat already, and by accident rather than by design: section 9.31 refuses to dispatch any draft whose first character is a slash, so nothing spawns and nothing is billed. What does NOT hold is that refusal's documented escape hatch. A user who genuinely means a leading slash types one leading SPACE, and the space survives the composer, the parse and the dispatch untouched — and then grok trims it and eats the slash anyway (measured). So the escape hatch reaches four seats and not the fifth. Nothing is rewritten to compensate, and that is the decision rather than an omission. Editing a brief on the way to ONE vendor would make five columns answer different questions while the room claims they answered one, and the room's whole premise is that the seats got the same brief. A blank column is the lesser failure, and it is documented in the file whose column shows the symptom. **The hue argument section 9.28 said would be owed.** That section closed by saying a fifth vendor would have to argue for its colour, and here is the argument. Grok is `14`, bright cyan — the twin of Codex's `6`, so section 9.28's honest weakness is now TWO twinned pairs rather than one. That is forced, not chosen: after 4/5/6/12 the legal set holds only 13 and 14, both twins of a seat already seated. The only real decision was which seat to pair with, and it went to Codex because silence routes to Claude alone — the Claude column is on screen in nearly every room, so keeping magenta unshared protects the seat a reader sees most. `CX` versus `GR` carries the distinction when a scheme renders 6 and 14 close, which is what the two-letter tags are for. Worth stating plainly for whoever adds a sixth: **the legal set is now full.** 13 is the last free index, and after it a new seat cannot have a hue of its own without taking a severity (making a seat read as failed) or abandoning 4-bit indices (council asserting a colour over the user's own scheme). `TestSeatHuesAreExhaustive` fails with that sentence in it. The tag is what scales; the hue was always going to run out. **What is verified, and what is not.** The parser is pinned against captured lines, and the INVOCATION is verified separately by `grok_live_test.go` (`-tags=live`) — because unit tests over captures would still pass if the argv this adapter builds were rejected outright by the CLI, which is the ADR-008 failure mode in its purest form. That test ran: the first turn returned the exact expected text, a session id and a positive cost; the resume turn came back on the SAME session id and recalled its own first answer, which is what distinguishes a real resume from a re-send. Not verified: anything on macOS — this is a Windows measurement, and the Mac's grok is untouched (`PARITY.md`). ~~Not built: an `internal/adapter/grok`, so no HUD row carries this vendor. Council can drive this seat; the gauges cannot yet see it.~~ **Amended 2026-08-11: it was built, and this paragraph went on saying otherwise.** `internal/adapter/grok` landed in PR #183, on the live survey §3.9a records (2026-08-09), so grok sessions render as HUD rows with name, model, workspace, a vendor-REPORTED context percentage and last activity. The gauges now observe this vendor as well as drive it, and the seat's parser and the adapter read the same wire from two sides. One thing stays dropped and it is not an oversight: grok writes a per-turn dollar figure and no session total anywhere, so the last turn's cost reaches the detail pane as a labeled Extra and never the `COST` column (§3.9a). **Fleet guard wiring for Grok under ADR-012.** agent-ops ADR-012 rules that guard wiring, not lane shape, is the control on every vendor — and grok is the fifth vendor seat in the room. ~~To fulfill the ADR-012 guard obligation for Grok, Grok's `PreToolUse` fleet guard is configured by setting `[compat.claude] hooks = true` in `~/.grok/config.toml` (which routes Grok tool calls through the fleet's Claude-compatible `PreToolUse` credential guard wrapper `pre_tool_use_credential_guard.py` / `hooks.json`), or by registering the `PreToolUse` hook script via `grok hooks-add`. This ensures secret stores, published history, and dangerous mutations are screened by the `PreToolUse` guard across all five seated vendors (Claude, Codex, Antigravity, Cursor, and Grok) without restricting lane capabilities.~~ **Corrected 2026-08-11: the struck sentences were instructions, and none of it is wired.** Two things were wrong at once. First the reading: `[compat.claude] hooks` is `false` on this box, and grok's native hook system (`grok hooks-add` / `grok hooks-trust`) has nothing installed into it, so **the seat has no fleet guards wired today** — the struck text described a configuration that does not exist and stated the screening as a fact. `STATE.md` carries the same measurement. Second the shape: this file records what was measured about a vendor, and those sentences told a reader how to configure a machine. A design record is not a runbook, and a runbook here would go stale silently on a box nobody re-measured. What replaces them is the obligation, recorded and routed rather than acted on. Under agent-ops ADR-012 an unwired vendor is an **open obligation on the FLEET**, never a reason to avoid the seat or to route work away from it — the gap closes by building the guard. **That work belongs to agent-ops and not to this repository.** telltale seats the vendor and states each seat's own posture on screen; it does not own the fleet's guard layer, and nothing here wires one. This paragraph exists so the next reader of §9.39 learns the obligation is open, and learns where it is owned. **Amended 2026-08-09, same day, from a live room: the seat could not take a briefed turn at all.** It was merged green and failed on its first real dispatch — `✗ failed 0s`, every seat answering but this one. The whole of it: ``` error: unexpected argument '--- operating context --- You are in a room...' found tip: to pass '...' as a value, use '-- ...' ``` `-p` was passed SEPARATED from its value. `Brief.Apply` prepends a fence to every first turn, so council's real prompt begins with `---`, and clap will not accept a hyphen-leading token as a flag's value unless that flag opts into `allow_hyphen_values` — which this one does not. So grok read the entire brief as an unknown flag and exited 2 before emitting a single event. Exit 2 with an empty stdout is the failure shape this adapter already documents for a bad resume id, so the room reported it correctly; there was simply nothing to report but a dead turn. The fix is one token: `--single=`, attached. Everything after the first `=` is the value, hyphens and newlines included. Verified against the exact failing shape on the first turn, and composed with `--resume`, where the resumed turn recalled a codeword only the first turn carried. The long spelling is deliberate — clap's attached form for a SHORT flag is `-pVALUE`, not `-p=VALUE`, so `--single=` is the unambiguous one. `--resume` keeps the separated form, because a session id is a UUID and cannot begin with a hyphen. **The lesson is about the live test, not about clap.** This seat shipped WITH an end-to-end live test that ran the real argv, and that test passed — because its prompt was `"Reply with exactly: LIVEOK"`, which begins with a letter. The test exercised the transport and never the shape the product actually sends. That is a narrower version of the same mistake §9.39 was written to avoid: a claim verified against a case nobody ships is not verified. `grok_live_test.go` now sends a fenced prompt on both turns, and `TestGrokAttachesThePromptToItsFlag` pins the property offline — no argv element may be a bare `-p`/`--single`, and none may be the naked prompt. Worth stating for the next adapter, because it generalises past this vendor: **a probe prompt should be shaped like a brief, not like a greeting.** Three of this file's captures used friendly one-liners, and the one hazard they could never have surfaced is the one that took the seat down on its first real turn. **Amended 2026-08-14: the drift alarm fired, and the seat was re-measured rather than re-dated.** grok reached **1.0.4 (d846eb93d9)**, four patch versions past the pin every claim above rests on, and nothing in the repository had noticed. PR #174 added version-pinned wire fixtures for exactly this, so this is the mechanism working, not a surprise. The re-measure cost three billed turns and one free one, and it is written up as a measurement because a version bump is not evidence that anything changed — nor evidence that nothing did. **The wire is unchanged, and that is a checked claim rather than an impression.** The 1.0.4 capture and the 1.0.0 one were diffed by SHAPE — frame types, key names, nesting, value types — and they are identical. The single difference in the two files is the KEY of the `modelUsage` map, `grok-4.5-build` → `grok-4.6-build`, which is a model id rather than a schema key, and one this seat's parser never reads. `testdata/wire/grok-1.0.4-turn.jsonl` replaces the 1.0.0 file. **Both containment flags are still dead.** - `--permission-mode plan` is refuted again, with the write landing again: the `write` tool was called, the update reported `completed`, the process exited 0, and `probe-plan.txt` held `WROTE` on disk. `--help` still offers `plan` among six permission modes, which is the whole point — the help text has said the same thing across five builds while the flag has never once been observed to stop a write. - `--sandbox` is still unobservable on Windows: `bogus-profile-xyz` drew no error, no warning and exit 0. **This one was re-probed for free, and the technique generalises.** It was passed alongside `--single=/context` — a prompt this vendor is already known to eat — because profile validation happens at startup, so a refusal surfaces before any model turn. A turn that is never billed still answers the question. macOS still diverges and fails closed (`PARITY.md`). **The flag that earned its re-run is `--resume`.** 1.0.4's help spells it `-r, --resume []` — an OPTIONAL value that also matches session titles, where the pinned build took a required id. An optional-value flag is precisely the clap shape whose SEPARATED form can stop binding, and this seat passes the id separated. Had it stopped binding, every follow-up turn would have quietly become a fresh conversation — the room's worst failure mode, because it looks like success. It did not: the resumed turn echoed the same `sessionId`, recalled the first turn's own word, and reported `input_tokens` 454 against `cache_read_input_tokens` 21504. The conversation was on the vendor's side. **The slash hazard survives, and so does the absent-cost shape.** `--single=/context` still produces `available_commands`, then an `end` with no usage, no cost and no text, at exit 0. So `TotalCostUSD`'s pointer-ness is measured on the CURRENT build rather than inherited from the pinned one. **What grok gained, and what it did not.** `grok models` now offers `grok-4.6` (default) and `grok-4.5`, where the 2026-08-09 survey found one model. The CLI has a `grok trace` subcommand that exports or uploads a session's trace data. **Neither changes a verdict here, and nothing is built on either.** Quota is still structurally absent (§7.16a, #195): a rate/limit/quota sweep over a session directory 1.0.4 itself wrote matches nothing account-level, and `grok trace` moves a transcript rather than reporting a ceiling. The telemetry seam §7.16a already spends is the OTLP push, and it is untouched. **One shape was measured that nobody had looked at before, and it is recorded as new rather than as drift.** A WRITE tool call's `content` element is a diff — `{"type":"diff","path":…,"oldText":"","newText":"WROTE\n"}` — with no nested `content` object, so `grokDetail` reads nothing from it. No write call was captured at 1.0.0, so this is a gap in the original measurement rather than a change under us, and `grok.go` says so at the function. It is left alone: a detail renders only on a FAILED outcome, and composing a sentence out of `oldText`/`newText` would be this package writing the vendor's line for it (§9.6a). **Not re-verified at 1.0.4:** the bad-`--resume`-id error shape, the `--verbatim` and `--prompt-json` slash channels, and anything on macOS. ### 9.40 the room said something was stopped and never said which seat (2026-08-09) The gate (§9.8) blocks a vendor until a key is pressed, and the room already announced that in two places. Neither of them names the seat. - The **card** sits in the blocked column and does not have to name it: its position *is* the seat, which is exactly why `gateCard` passes an empty subject there. - The **mode line** cannot. `gateLabel` prints the oldest request's own text — `GATE Write: internal/council/gate.go (+2 queued)` — which says what is blocked and never who. So in a five-seat room the footer says something is stopped, and the reader finds out which column by going through them one at a time while it stays stopped. On a projector, driving four seats live, that is the stall. **The needs-you strip is one line of room chrome that answers the question the other two cannot:** `⚠ NEEDS YOU 2 Codex 3 Antigravity`. **It is driven by the gate queue and by nothing else, and that is the whole safety property.** `State.Gates` is a structured record of vendors that asked for permission and have not been answered. Every name on this line comes from one of those entries. A seat that has gone quiet, a seat streaming nothing, a seat whose prose happens to end in a question mark — none of them reach it, because none of them is a *measurement* that anyone is blocked (§4a.1). "Needs you" is a claim about a vendor waiting on a keystroke, and the queue is the only thing in this room that knows. `TestTheStripSaysNothingWithoutAPendingGate` builds all three of those look-alikes at once and asserts no strip is drawn. **A seat leaves when the reader goes to it — and that is the only thing besides answering that takes it off.** Derived from `State.Focus` rather than stored as an acknowledged set, on `Gating()`'s own argument: a stored set is a second place for the same fact to live, and the two drift the first time a seat's gate is answered and a *new* one arrives while the old acknowledgement is still in the map. At that point the anti-stall silently omits a seat that is waiting, which is the one failure it exists to prevent. Derived, the worst case is the strip re-listing a seat the reader visited and left — and that is true: it is still stopped and nobody is looking at it any more. `TestGoingToASeatIsWhatClearsIt` pins the three clearings that are refused (the seat going quiet, the seat producing output, time passing) alongside the one that is allowed. **The default focus is a hole in that rule, and it is deliberately left open.** `NewState` seats the keys on column 0 without anyone pressing anything, so a gate on whichever seat happens to be focused never appears here at all. That is the correct outcome rather than a gap: the reader is looking at the column whose card is already spelling the question out in full, and a room-level line naming the seat under their own cursor is the duplication §9.30 spent a section removing. #### Where it sits, and what it does not yield to Directly under the frame's heavy rule — **above** the collapsed-seat notice and above the band. The other two are ordered by subject size (the room outranks the turn); this one is ordered by **urgency**, on `modeLine`'s own precedent: a gate is the only state in this room where something is STOPPED until a key is pressed, which is why `GATE` outranks every other mode word on the footer. Seats that are not on screen and a brief that was just sent are facts a reader can come back to. A blocked vendor is not. It costs a row and, unlike the band, **it does not yield to a short terminal.** The band's whole value is removing a duplication, so retiring it falls back to the frame as it was and loses nothing that was not already on screen. This line is the only place the room says *which* seat is stopped, so a row spent on it is the last row the height budget should reclaim. It is spent in **every tier** for the same reason — at the tabs tier the blocked seat may be the one column not on screen, which is precisely where a reader has no other way to learn it exists. #### The ladder, and one ordering that had to be reasoned about again Longest-first, widest-that-fits-wins — `stripHeader`'s idiom, so this package has one shedding shape rather than three. Three rungs: 1. every seat, by name; 2. every seat, by the two-letter tag §9.25 made permanent; 3. as many tagged seats as fit, and a count of the rest (`+2 more`). **Identity yields before a SEAT does**, which is §9.18's order and the rung boundary worth writing down. A four-seat strip at sixty columns can hold `2 Codex 3 Antigravity +2 more` or it can hold `2 CX 3 AG 4 CU 5 GR`, and the second is better by the only measure this line has: the reader is asking who is stopped, four abbreviations they already learned from the column headers answer it completely, and two names plus a count answer half of it. So the tag rung is tried at *full roster* before any seat is dropped. Nothing is ever clipped at any rung — an entry survives whole or leaves and is counted, because `Ant` is not a shortened `Antigravity`, it is a seat this room does not have (§9.18). The count is never traded away either, on `overflowMarker`'s rule: "there is more" without a number is the marker §9.10 shipped and got reported as a room that could not scroll. Below the width where even one tagged seat fits, `⚠ NEEDS YOU` survives alone — still true, still the signal, and honest about being unable to say who. That floor is unreachable in a real frame (MinWidth leaves the strip 56 cells and the lead costs eleven) and is written out anyway, because the last time a floor was assumed rather than enforced it was wrong by four. #### Two smaller rulings **The seat number is printed only where the key is live.** Digits focus a seat through `viewKey`, and `gateKey` falls through to it, so while the room is gating the number works in both modes — except on a turn page, where `focusSeat` refuses outright because there are no columns to move between (§9.22). A number printed there would be the room naming a key that does nothing, which is §7.8's surprise. The names stay, because *who is stopped* is still true on a page. **A seat folded out of the grid is still named.** Its vendor is blocked whether or not the room drew it a column, and a blocked vendor with no card, no column and no line anywhere is the disappearance §4a.1 forbids. It carries no number, because no key in this room reaches it, and it is therefore the one entry that focus cannot clear — which is honest: nothing the reader can press from here will unblock it. **No new hue, and no new site for the old one.** The line is `Alert` — SevWarn at weight, the gate card's own title style — with the seat numbers `Muted`, which is the chrome/anchor split the column header and the tab bar already use for those same two things. §9.28's list of three places a seat hue is spent stays closed; it was ratified as closed, and a fourth site is a decision for whoever wants to reopen it rather than a side effect of this line. Under `--ascii` and `NO_COLOR` the strip reads exactly the same, which is the property every distinction this UI makes has to have: `NEEDS YOU` is the signal, and the mark and the weight only make it findable. #### 2026-08-16: the Notification hook, measured on three runtime surfaces The strip above reads council's own gate queue and nothing else. A 2026-08-15 research candidate proposed a second source for it: Claude Code's `Notification` hook, with the matcher names `agent_needs_input`, `permission_prompt` and `agent_completed`. Those three names were a vendor claim. No measurement in this repository supported them. §7.21's trap 1 measured that hook firing changes with the runtime surface, so a claim about one surface says nothing about another. This subsection records the survey. It builds nothing. **The rig.** A throwaway workspace holds a `.claude/settings.json` that registers one recorder command against every hook event the survey can name, including three `Notification` entries: one with no matcher, one with `matcher: "permission_prompt"`, and one with `matcher: "agent_needs_input"`. The recorder appends its raw stdin plus an arrival timestamp to a file named for its entry, and exits 0 on every path. **The filesystem is the observable, never the stream**, which is §9.8's rule. A file that does not exist means the hook did not run. The operator's own `~/.claude/settings.json` was never written to, and no credential store was copied anywhere. Claude Code **2.1.233**, Windows 11, `claude-haiku-4-5` on every turn, two trials per arm. The turn asks for one shell command, `install -d probe-marker`, which is the shape §9.8 already measured that no allow rule on this box covers. **The rig lied once, and the record says so.** The recorder took an output directory argument and ignored it, so the first four arms reported "no hook files" when the files were landing in the workspace directory instead. That reading survived three arms and produced a false conclusion, that project settings were not loaded at all. It was caught by running the recorder by hand. The zero a probe reports is a claim about the probe until the probe itself is checked, and this one was wrong. **The source read, at the pinned version.** CLAUDE.md permits a source read at a pinned version as evidence, and the shipped 2.1.233 binary carries four facts the live arms then tested. First, `notification_type` is an enum of eleven values, not three: `permission_prompt`, `idle_prompt`, `auth_success`, `elicitation_dialog`, `agent_needs_input`, `agent_completed`, `elicitation_url_dialog`, `worker_permission_prompt`, `push_notification`, `computer_use_enter` and `computer_use_exit`. Second, the payload builder sets `hook_event_name` to `"Notification"`, copies the type into a `notification_type` field, and passes that same type as the hook's match query. The matcher therefore does select on the notification type, and the three claimed names are real matcher values. Third, the permission notification is armed by a `setTimeout` of **6000 ms** that returns a cancel function, and it is armed immediately before the `can_use_tool` control request is sent. The environment variable `CLAUDE_CODE_DISABLE_PERMISSION_PROMPT_NOTIFY_HOOKS` turns it off. Fourth, `agent_needs_input` and `agent_completed` are emitted from a React effect that watches **background agent** sessions, beside a telemetry event carrying a `jobSessionId`. Their subject is a background job, not the session the hook is installed in. **What fired, per surface.** Every arm denied the tool call, so the `PostToolUse` column is uninformative and is left out: a call that never ran has nothing to report, and this rig therefore neither confirms nor contradicts §7.21's trap 1. | surface | `Notification` (no matcher) | `permission_prompt` | `agent_needs_input` | controls that did fire | trials | |---|---|---|---|---|---| | `claude -p` | no | no | no | `PreToolUse`, `SessionStart`, `Stop`, `UserPromptSubmit` | 2/2 | | `claude -p --output-format stream-json --verbose` | no | no | no | the same four | 2/2 | | control protocol, request answered after 12 s | **yes** | **yes** | no | the same four | 2/2 | | control protocol, request answered at once | no | no | no | the same four | 2/2 | The control-protocol rows replicate council's own gated seat: `baseArgs` plus `gateArgs` plus `--input-format stream-json`, with the recorder supplied through `--settings`. The last two rows are the same rig, and they change one thing, which is how long the probe waits before it answers the `can_use_tool` request. **The six seconds are the finding, and the fast arm is what makes it one.** A request held for 12 s produced the hook on both trials. The notification arrived 6931 ms and 9727 ms after the request appeared on stdout, which is the 6000 ms timer plus the cost of starting the hook process. The same rig answering the same request immediately produced **no `Notification` at all**, on both trials, while all four control hooks ran in every arm and prove the settings file was loaded. So the hook does not report that a vendor is blocked. It reports that a vendor **stayed** blocked for six seconds. Every prompt an operator answers faster than that is invisible to it. **The payload, re-typed with synthesized identifiers.** This is the whole record. The session id, the prompt id and the paths are fake, per the fixtures rule. ```json { "session_id": "11111111-2222-3333-4444-555555555555", "transcript_path": "C:\\Users\\example\\.claude\\projects\\C--probe-ws\\11111111-2222-3333-4444-555555555555.jsonl", "cwd": "C:\\probe-ws", "prompt_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", "hook_event_name": "Notification", "message": "Claude needs your permission to use Bash", "notification_type": "permission_prompt" } ``` **There is no correlation handle in it.** The payload carries no `tool_use_id`, no control-request id, and no tool-name field. The `title` key the builder allows was absent on every trial. The only statement of *what* is blocked is an English sentence, so a reader that wanted the tool name would have to parse `message`. §4a.1 does not permit that to become a displayed value. **The verdict for this strip: refuse the seam, and the refusal is `CapNone`-shaped.** For a seat council spawned, the gate queue already holds the fact, holds it immediately, and holds it structurally, with the vendor, the tool and the request id all present. The `Notification` hook offers the same fact six seconds later, without the tool identity, and only when the operator was slow. It is strictly worse on every axis the strip cares about, and adding it would break the safety property §9.40 is built on, which is that every name on the line comes from a measured pending gate. A second source that is late and lossy would put a seat on the line after the operator already answered it, or leave one off entirely. **Two of the three claimed names never appeared, and the source read says why.** `permission_prompt` exists and fires. `agent_needs_input` and `agent_completed` exist in the enum and fired on no headless surface in any of the ten runs. Their emitter watches background agent sessions from a render effect, so their subject is a background job rather than the current session, and the effect has no render tree to run in outside the interactive terminal. A reader that took the three names as equivalent would have wired two signals that cannot arrive and one that arrives late. **What this does not close.** The interactive surface is **OWED**, and only an operator can drive it. The prepared steps are: start `claude` in a throwaway directory whose `.claude/settings.json` carries the recorder, ask for `install -d probe-marker`, leave the approval prompt untouched for more than ten seconds, then answer it; repeat and answer within two seconds. The open question there is whether the interactive surface adds `idle_prompt`, and whether a real background agent makes `agent_needs_input` fire at all. Three smaller gaps stay open beside it. The operator's own `~/.claude/settings.json` could not be read, because the credential guard refused it and the refusal was not worked around, so the number of user-level hooks in the control column is unknown. `worker_permission_prompt` was never registered and never seen. And the six-second timer was measured only from the request appearing on stdout, which is a few milliseconds after the CLI arms it, so 6931 ms is an upper bound on the delay rather than the delay itself. **A different question this survey does not answer.** Everything above concerns seats council spawns, where the gate queue is the better source. Claude Code sessions telltale merely observes, through the HUD and the statusline, have no gate queue behind them. Whether the `Notification` hook could carry a needs-input signal for **those** is untested here, and it inherits the same six second delay and the same missing tool identity, so it would start from a weak position. ### 9.41 the gate asked about an edit and would not show it (2026-08-09) The approval card (§9.8) names the call — `⚠ waiting on you: Edit: internal/council/gate.go` — and that is the whole of what a user has to decide on. For a `Bash` it is enough: the command *is* the action. For an edit it is not even close. The path says which file is about to change and says nothing about *what changes*, so the only two answers available are "yes, because I trust it" and "no, because I don't" — which is the gate reduced to a mood. **The card now shows the edit itself, as a red/green before/after, when the vendor's payload carried one.** ``` ⚠ waiting on you: Edit: internal/council/gate.go - func gateCard() { - return nil - } + func gateCard() []string { + return lines + } y approve n deny a stop asking ``` #### What it renders is what was measured, and nothing else This is §4a.1 on the one card in the product that guards a write, so the rule is stricter here than anywhere: **council never opens the file, never reconstructs a before from an after, and never shows one half as if it were two.** A preview is drawn only when the vendor's own permission request carried *both* halves. Everything else — every `Bash`, every `Read`, every `Write`, every request from the Cursor seat — renders the card exactly as it rendered before this section existed, and `TestAPayloadWithoutABeforeShowsNoPreview` pins that by comparing against the *existing* `gate-card` golden rather than a new one of its own. Which payloads carry both was measured on 2026-08-09 against **Claude Code 2.1.226** on Windows, driving the gated invocation (`--permission-prompt-tool stdio`, `--permission-mode manual`, `--setting-sources ""`) in a throwaway directory. Two requests from that session, quoted whole on `runner.Gate` and replayed verbatim as tests: | tool | what the payload carries | what the card draws | |---|---|---| | **Edit** | `old_string` **and** `new_string` (plus a `replace_all` the room has no use for) | the before/after | | **Write** | `file_path` and `content` — the after, and no before at all | nothing | | **Cursor / ACP** (`session/request_permission`) | a title, a kind, and an options list (§9.36's capture) | nothing | `content` is deliberately *not* read as a new half. A Write knows what the file will say and says nothing about what it says now, so treating it as an addition would paint a green block against a before council never saw — a plausible value where §4a.1 requires an absent one. The Cursor seat is the same refusal one level up, and a smaller hole than it sounds: that seat does not ask about edits at all (§9.36), so the card it does raise is about a command, where the command is already the whole decision. **The two halves are a pair or neither.** `editHalves` fills both or returns nothing, which makes the renderer's test simply *do they differ* — and that one question folds in three "show nothing" cases honestly: no halves at all, an edit that changes nothing, and an edit whose only difference was a redacted secret. An empty `new_string` beside a non-empty `old_string` is a legal, measured **deletion** and draws all removals and no added-side count; "0 more added lines" would be the card filling a slot rather than answering a question. #### The one thing that crosses the Input boundary, and why it is two strings `runner.Gate.Input` is the vendor's whole argument blob, held only to be echoed back on an approval, and it has always been kept **off** `State` on purpose: for a Write it is the entire file content, one careless line away from the screen. That rule is not repealed here. What crosses is a **projection** — two named strings, `PendingGate.Old` and `.New` — because two fields with one purpose cannot be reached for by accident the way a map can. They are read by the *adapter*, not by council, for the same reason the card's `Text` is composed there: the key names are that vendor's, and the room renders a preview without knowing whose spelling produced it. They are redacted on the way through, and that matters more here than on the argument line the existing redaction was written for. A command is a *likely* place for a token to appear; the body of a file being edited is where one actually lives — and this lands in chrome that does not scroll away. Each half is redacted separately because each is rendered separately. Only trailing newlines are trimmed, never leading whitespace: an indent is content, and a preview that silently unindented the code it is asking about would be showing an edit nobody requested. #### The prefixes carry it; the colour only seconds it `-` and `+` are the entire signal, exactly as they are on §9.37's raw patch lines — which is why the preview reads identically under `--ascii` and `NO_COLOR`, and why the goldens, which render `PlainStyles`, are the *proof* of that rather than an approximation. The styling reuses `Styles.ForDiffLine`, the same classifier those patch lines already go through, so **council adds no hue for this**: green-for-added and red-for-removed is one convention spent twice, not a second vocabulary, and §9.28's closed list is untouched. The marks are patch punctuation and **not** entries in the `Glyphs` alphabet, which is worth saying because the ASCII set already spends `+` on `ActOK` and `-` on the light rule. They do not collide, on `Glyphs.Range`'s own slot argument: those are marks that stand alone in a slot, and these only ever open a line inside this block. #### Bounded, per half, and it says what it dropped The card is **chrome** — it costs body lines, `MaxScroll` derives the ceiling from it, and every row spent here is a row of the reply the user cannot read while deciding. So each half shows at most **three** lines and then counts: `2 more removed lines not shown`. The count is **per half**, not one total, because a long removal would otherwise spend the whole budget and take the additions with it, with no line admitting the additions had ever been there. It carries no glyph of its own: `…` has an ASCII partner of `>`, and `> 2 more removed lines` reads as a comparison rather than as a truncation — a marker that can be misread as a number is worse than a longer sentence. Long *lines* are cut with the ellipsis glyph on the plain text **before** the style is applied — classify, truncate, style, and only then let `fit` pad — because `fit` alone clips silently and would be clipping a string that now genuinely carries ANSI, which is §9.5's trap read backwards. #### What is deliberately not here **No intra-line diff.** The unit is the line, because that is the unit the payload arrives in; highlighting *which words* changed inside a line would be council computing a diff rather than displaying one, and the first thing it would get wrong is the case it was added for. **No scrolling the preview.** A card with its own viewport is a second scroll surface in a room that already has one per column plus a turn page, and the keys to move it would have to be taken from a mode whose whole contract is that `y` and `n` mean one thing each (§7.8). The bound plus an honest count is the trade; the whole edit is in the file the moment it is approved. **No preview on a denial or after the fact.** The trace still says `✗ denied by you` and nothing more. What the seat *would* have written is not what happened, and a room that showed it afterwards would be displaying a file state that never existed. ### 9.42 `telltale doctor`: the one moment probing is allowed, and the three answers it may give (2026-08-09) Council's detection has never run a vendor. `detect.go` says why in its own doc — "council never runs a vendor to find out whether it works: a probe turn costs real quota, and 'is it authenticated?' is a question the first real dispatch answers for free" (ADR-008 §6) — and that rule is correct and stays. What it is a rule *about* is a **turn**: the room may not spend the user's money to answer a question it does not have to ask, and it may not spend it silently, mid-conversation, on a schedule nobody asked for. So the boundary is drawn at **cost and side effect**, not at "running a vendor is forbidden". `telltale doctor` runs ` --version` — a flag that parses argv, prints a string and exits, with no model, no session and no billing anywhere in it. It starts no turn, reads no credential store, makes no network call, and writes nothing at all: it adds no fourth exception to the three writes `CLAUDE.md` lists, because it has none. §9.17 is the frame that makes this the right place for it — *a fact that is true at launch and stays true belongs at launch*, and what binary is on this disk and what it calls itself is exactly that shape. The inverse of the same rule is why there is no `/doctor`: nothing this reports changes while the room is open. **Three states, and the whole design is keeping them three.** | | what it means | what it may carry | |---|---|---| | `ok` | the check ran and passed | the measured text — the path that was stat'd, the line the vendor printed | | `FAILED` | the check ran and did not pass | the reason | | `not checked` | the check did not run | **why** it did not, and never a value | **Auth and network are `not checked` on every seat, always.** A binary that exists and answers `--version` establishes nothing whatever about a login or about reachability, and a report that let those two ride along on the good news would be believed on the one day it was wrong — the §4a.1 collapse wearing a preflight's clothes. The reason is printed rather than shrugged: the only thing that establishes a login is a real turn, and a turn costs quota. A seat that is installed and signed out reports its own auth failure on its column the first time you dispatch to it, which is where that fact was always going to come from. **`NotChecked` is the zero value**, for §9.17's `GateOff` reason exactly: a safety property whose default is the reassuring answer is the wrong way round however carefully the constructor sets it. Every `Check` a test types by hand, and every field a future change forgets to fill in, reads as an honest blank instead of a silent pass. `TestNotCheckedIsTheZeroStatus` pins it. **The measurement that would have been a plausible lie.** One seat's resolved binary is not the program it names. Detection steps over `cursor-agent.cmd` to the bundled `node.exe` its launcher would have run (§9.33), so the obvious `--version` answers **`v24.5.0`** — node's version, printed under a row labelled `cursor`, real and about the wrong program. Handed the bundle first, `node.exe index.js --version`, the same install answers **`2026.08.04-aaa8809`**. Both measured here, one after the other. The version argv is therefore per-seat data taken through `vendors.CursorNodeBundle`, not a `--version` constant and not a second `filepath.Join` — that function's own doc names two copies of one join as the agreement that silently stops holding. The other four take the bare flag, verified by running them: claude `2.1.226 (Claude Code)`, codex `codex-cli 0.147.0`, agy `1.1.11`, grok `grok 1.0.0 (3cd0d0cbce) [stable]`. **Found and undrivable is a third thing, and it stays one.** `detect.go` already refuses to collapse "not installed" with "installed somewhere council will not drive from", because the fix differs. The preflight carries that through as two checks rather than one verdict: `binary` passes and `drivable` fails, with the shim note attached. The version probe still runs on such a seat — what is installed is worth knowing where the room cannot use it, and a fixed `--version` carries none of the prompt text that made it undrivable. **A fifth mode, not a council flag, and no colour in it.** What it prints goes somewhere else: council's output is a full-screen room, so a preflight rendered inside one is unreadable at exactly the moment it is wanted — before the room opens, piped into a file, pasted into an issue. Every distinction it makes is carried by a **word** (`ok`, `FAILED`, `not checked`), so `--ascii` and `NO_COLOR` have nothing to switch off and neither is a flag here; a flag that does nothing is a promise that something was configurable. `Render` is pure over its `Report` for council's reason, with the probe durations measured in `Run` and arriving as data. Each seat gets its own `--timeout` (15s default) rather than the run sharing one: a wedged vendor must cost its own deadline and a failed check, never the report. And the command exits **0** whatever it finds — a failed check is this mode working, and exiting non-zero would make "I have four of five seats" indistinguishable from "doctor itself broke". **Capabilities are printed and are explicitly not checks.** How a seat streams and whether it can be asked to ask first were measured once, against live runs (§9.7, §9.33, §9.36, §9.39, and `canGate`), and written down; nothing re-measures them on this machine now. They render on their own labelled line — *"council declares, and did not check here"* — outside the status column, because putting them in it would give them a fourth state and imply a check that never happened. The words are read off `granularityFor` and `canGate` rather than restated, so the preflight cannot drift from the room's own badges. Worth the line at all because "installed" and "will stream to you" are different promises, and the seat a user is most likely to think is broken is the one that is working and silent until the end of the turn (§9.14). #### 2026-08-16: the survey pin gets its maintainer loop, and it lives here **The gap.** Every adapter is pinned to the vendor build its field map was surveyed at, and §3.10 is the inventory of those pins. `internal/adapter/drift` watches the on-disk SHAPE and tells the USER when a row goes quiet. Nothing told the MAINTAINER that the survey behind a row is now older than the vendor on this disk. `agy` and `grok` self-update, so that pin goes stale in silence: no check fails, no row blanks, and a set of displayed fields rests on a survey of a build nobody runs any more. **CI can never close this.** CI installs no vendors, so every version probe there resolves to nothing and every comparison is vacuous. The loop has to live where the vendors live. The owner ruled on 2026-08-16 that the place is this mode, and it is the obvious host: `doctor` already resolves each seat's binary and already asks it its version. The comparison costs one string compare on a probe that was happening anyway. **It is a staleness fact, and it is not one of the three states.** A drifted pin says nothing is wrong with the reader's machine. The seat works, the binary answered, and every check that ran passed. So the notice renders on its own labelled line — *"telltale's field map for this vendor was measured at"* — beside the capability line and outside the status column, for the neighbouring reason: both are claims this repository measured once and wrote down, and neither was re-measured here. Making it `FAILED` would redden a working vendor, move the tally, drop `Ready`, and put a red cross in front of a user whose install is perfect. Making it `Passed` would claim a check ran on their machine when what actually happened is that this repository compared its own homework. The tally, `Ready` and the exit code are all untouched, and `TestDriftIsNotAFailedCheck` pins that by running the same seat with and without a pin and requiring the three counts to be identical. **Four outcomes, and the two silent ones are silent on purpose.** | | what the line says | |---|---| | the versions differ | the pinned build, the installed build, and the §3 section to re-measure | | they match | the pinned build, and that it is what is installed | | the version was not read | the pinned build, and that nothing is claimed in either direction | | the two cannot be compared | the pinned build, and why no comparison is possible | Rows three and four are the honesty of the feature. A seat whose `--version` probe failed hands the comparison an empty string, and an empty string gets **no verdict** — claiming a match there would be as invented as claiming drift, and both would rest on nothing measured. Row four is the Cursor seat specifically: its pin names the Cursor **application** the store was surveyed inside (§3.9, `Cursor 3.14.7`), and what this mode probes is **cursor-agent**, which answers a date-stamped build of its own (`2026.08.04-aaa8809`, §9.42's own version table). Comparing `3.14.7` with `2026.08.04` would manufacture a permanent drift notice out of two unrelated numbering schemes — a notice that fired forever on a correct install, which is exactly the report `internal/adapter/drift` rules out as one nobody reads. **The comparison is equality, never ordering.** It extracts the first dotted-numeric run from each string, because the pin and the probe agree on a number and on nothing else: the adapter wrote `Claude Code 2.1.219` and the binary answered `2.1.226 (Claude Code)` when this was measured (that pin has since moved to `2.1.233`, §3.1's re-measure); the adapter writes `grok 1.0.4 (d846eb93d9)` and the binary answers `grok 1.0.0 (3cd0d0cbce) [stable]`. A commit hash carries digits but no dot, so it cannot be mistaken for the version beside it. No line claims "newer" or "older": a direction needs per-vendor precedence rules this program has no business inventing, and a downgrade means the same thing an upgrade does — the survey was measured somewhere else. `TestTheComparisonNeverClaimsADirection` holds that. **This does not make a version comparison a canary, and nothing on the read path changed.** `internal/adapter/drift`'s doc rules a version out as a TRIGGER, because a report that fires on every release is a report nobody reads. That ruling stands untouched: no adapter degrades a field, writes a diagnostic, or reads anything new. The version comparison is admissible only because it is asked once, by hand, before the room opens, by an operator who ran a preflight — and it is answered to a different reader, about this repository rather than about their machine. The two are separated in the report itself: the closing note says a format that actually moved is reported on the HUD row and does not wait for this. **One source for the pins, enforced by the compiler.** `internal/adapter/pins` is the table, and every `VerifiedAgainst` in it **is the adapter's own constant** — which is why those six constants are now exported. No pin string is copied. `Section` and `DocLabel` are the only facts the table adds, and they have no other home: an adapter knows what it was verified against, not which prose section carries the evidence, and that pointer is what turns "this is stale" into an instruction. `internal/doctor` stays stdlib-only and holds no inventory; `council.DoctorSeats` attaches the pin, exactly as it already attaches the capability declaration, and for the same stated reason. **The doc guard reads the cell, not the document.** §3.8 records that `TestTheCanaryInventoryMatchesThisAdapter` substring-matches a pin *anywhere* in this file, so a dated paragraph quoting a new pin turns the guard green while the table it guards stays wrong — which is what let the Antigravity cell read `agy 1.1.9` for a release after the adapter moved to `1.1.13`. §3.8 named the fix, "scoping that assertion to the table", and left it unowned. `pins_doc_test.go` parses §3.10's table and compares it cell by cell, in both directions: a pin quoted in prose elsewhere cannot satisfy it, and a doc row no adapter claims fails it too. This was verified by mutation — restoring the stale `agy 1.1.9` cell fails the new guard while `internal/adapter/antigravity`'s own test stays green, reproducing the 2026-08-03 miss exactly. **The six adapter tests are deliberately left as they are.** They assert a different thing (that a pin is documented at all), nothing here weakens them, and rewriting them is a separate concern. **Measured, not assumed.** Run on this Windows box on 2026-08-16, the report drifts `claude` (pinned `2.1.219`, installed `2.1.233`) and `codex` (pinned `0.146.0`, installed `0.147.0`), matches `agy 1.1.13` and `grok 1.0.4`, and declines to compare `cursor`. That is the feature finding two genuinely stale surveys on its first live run. **Re-measuring those two surveys is not part of this change** — the loop reports staleness; a re-survey is its own work, with its own live corpus. #### 2026-08-17: the preflight states each seat's posture, from the room's own claim **The gap.** Every council column carries a sandbox badge, and the help panel's posture page carries the measured argument under each one (§9.13, §9.2's ruling that a claim you cannot see is not a claim). Both of those are read **inside the room**, which is after the decision they inform. A user picks a workspace and a posture *before* the room opens, and the only surface that runs before the room opens said nothing about either. §9.17 settles that the fact belongs here: what a vendor's own flags buy on this machine is true at launch and stays true, and it is a property of the vendor and the OS rather than of a turn. **One source, two surfaces.** Nothing in the block is written in `internal/doctor`. `council.DoctorSeats` builds it from `postureClaim` — the same function the room's own columns are built from — and hands over the badge word off `SandboxClaim.Badge()` and an evidence class off the claim's `Level`. That routing is the whole design: a preflight with a per-vendor posture table of its own would agree with the badges on the day it was written and diverge the day a level moved, and a reader looking at two disagreeing surfaces has no way to tell which is lying. The capability declaration and the survey pin are attached at the same seam, for the same stated reason. `TestThePreflightPostureIsTheRoomsOwnBadge` pins it through *different* construction paths on each side — `DoctorSeats` against the columns `stateWith` builds — because comparing `doctorPosture` with `postureClaim` would be comparing a call with itself. **The badge says what the posture IS; the evidence class says what it RESTS ON.** They are different questions, and `unsandboxed` is the case that proves it: two seats reach that badge because a live run **refuted** the flags and because **no flag was ever passed**, and a reader deciding whether to point council at a worktree needs the second sentence. §4a.1's rule that two kinds of nothing must not render alike is the same rule one level up. | badge | evidence class | |---|---| | `ro:tools` | enforced by **construction** — the write and shell tools are absent from the session | | `ro:enforced` | enforced by an **operating system** — the vendor's own sandbox | | `ro:requested` | **asked for**, and never observed on this machine — weaker than either above, and says so | | `unsandboxed` | **measured** not to restrict — refuted by a live run, not merely unestablished | | `WRITES` | nothing was asked for at all | | `gated` | **your keystroke** — the seat asks before every tool call that changes anything | `evidenceClass` is a table keyed by level, so `TestEveryPostureLevelHasAnEvidenceClass` can walk the type and fail the build the day a sixth level renders a badge with nothing to classify it — the guard `helpBadgeGloss` already carries inside the room. `TestNoEvidenceClassSoftensItsBadge` holds the other half on `TestThePostureLegendDoesNotSoftenAnyClaim`'s terms exactly: these sentences classify evidence and never weaken it, and none of them may call a posture read-only, safe or unable to write. **The rows are the `--read` room, and the argv is on the block.** The room **WRITES by default** and `--read` is the opt-out (`cmd/telltale`, and the legend inside the room was already corrected once for crediting the retired `--write` flag). The rows report the `--read` posture because that is the only one that is a fact about the *machine* — the default room's badge is a property of an argv the reader has not typed yet, and five cells all reading `WRITES` would carry nothing per seat. So the header names `telltale council --read` in its first clause, and **one closing declaration** states the default: the room writes, *n* of *m* seats can be asked to ask first, and what contains a writing room is the workspace, not any of these words. The gating half is **counted off `canGate`**, never written down — that measurement has already moved once, when the Cursor seat became a live process that can be asked and still does not ask about edits. **It is a claim, and it is not one of the three states.** A posture was measured once against a live run and written into this repository; nothing re-measures it on the reader's machine. So it renders outside the status column, beside the capability line and the survey pin, and it is wrong in *both* directions as a check: a `FAILED` would redden a working install over a vendor's own design decision, and an `ok` would claim this preflight established a containment property it never probed and could not probe without spending a turn. `TestAPostureIsNotACheck` pins that the way `TestDriftIsNotAFailedCheck` does — the same seat with and without the data, the three counts required to be identical — and `TestThePostureBlockCostsNoProbe` pins the other half: the block adds **no probe, no network call and no login check**, because every string in it arrives with the seat. The exit code is untouched. **A seat council states no posture for gets `no claim`, not a missing row.** A seat absent from a posture table reads as a seat with nothing to declare; this one has an unanswered question. The word deliberately is not shaped like `not checked` — the three state words are spoken for, and a fourth column borrowing one would put this block back inside the block it is outside of. **Measured on the reference Mac, 2026-08-17.** The five rows read `claude ro:tools`, `codex ro:enforced`, `agy unsandboxed`, `cursor ro:requested`, `grok unsandboxed`, and the declaration reads *1 of the 5 seats above can be asked to ask first: claude*. The `codex` row is the platform branch working: at that date the same block on Windows read `unsandboxed` there, because council passed `-s danger-full-access` on that OS (ADR-008's twelfth amendment; since §9.2's 2026-08-29 amendment the Windows row reads `ro:enforced` too, with its own dated detail) — and it reads it from `postureClaim`, not from a second platform test in the preflight. ### 9.43 the agy seat stops pretending a lost thread resumed (2026-08-09) `STATE.md` carried this as an unowned gap for as long as it took to write the entry. **The agy seat could not tell a lost thread from a resumed one, and the stream would not say.** Measured 2026-08-09 against agy 1.1.11 during the wire-fixture capture (PR #174; the record is `internal/council/vendors/testdata/wire/README.md`, under *what could NOT be captured, and why*): handed a `--conversation` id it does not hold, that CLI **does not error**. It opens a NEW conversation, answers the brief normally, and reports `status: "SUCCESS"`, exit 0. That is the whole difficulty in one sentence. Every other seat resolves the question for the room: Claude Code returns a `result` frame whose `errors` array says *"No conversation found with session ID: …"* (PR #178); codex writes `no rollout found` to stderr and exits 1; grok does the same with a 404; the Cursor seat answers `session/load` with -32602 and opens a fresh one in the same process (§9.36). agy claims success either way — so a room reading status and exit code, which is every honest thing the seat did before this, rendered **a continued conversation over a reply that had no history behind it.** Not a crash and not an empty column: a plausible answer under a `restored` mark that was no longer true. **The tell, and it is the only one the capture surfaced: the `conversation_id` that comes back is not the one that was asked for.** Nothing read it. Reading it is this change. **The comparison lives in council, not in the adapter, because neither half is where the other is.** `ParseEvent` sees one line at a time and never learns which id the turn requested; dispatch knows the request and never sees the stream. So `specFor` now *returns* the id it asked to resume — the one fact a caller cannot re-derive, since it is buried in a vendor-specific argv position — dispatch records it, and `adoptSession` compares it against the id the vendor names on its own session event. **The vendor gate is structural, and that is the honesty argument rather than a taste in plumbing.** The arithmetic (`asked != returned`) is vendor-neutral; the *conclusion* is not. A CLI that re-keys a resumed thread while keeping its history would look identical on the wire, so a room that compared ids for everybody would announce lost threads it had never measured — §4a.1's inference, wearing a comparison's clothes. The claim is therefore made by the seat, through `vendors.SilentResumeFork`, whose one method returns **the build the fork was measured against** rather than being a bare marker: a seat cannot make the claim without naming its evidence, and a vendor bump that fixes the behaviour leaves a version string that no longer matches the fixture beside it. Only agy implements it. `TestOnlyAMeasuredVendorArmsTheForkComparison` pins that it stays alone until somebody captures a second case, and the gate is applied at *dispatch* — a seat that never enters `forkWatch` cannot raise the card at all. **Three rulings on what the room then does, each of which had a plausible alternative.** | | what happens | the alternative, and why not | |---|---|---| | the reply | **renders, untouched** | failing the column would throw away an answer the user paid for, to punish a bookkeeping mismatch. The turn succeeded; what was false was only the claim that it was informed by everything before it | | the new id | **adopted as this seat's thread** | discarding it orphans a real turn — the reply happened *inside* that conversation — and leaves the room rebuilding the same forking invocation on every later turn | | the card | **the calm lost-thread card already in use** | a second card would say the same fact in different words. The outcome is identical to a refused reattach — this seat is starting fresh — and only the body differs, because there the turn failed and the id was let go, here it succeeded in a thread nobody asked for | So the column reads *"thread not restored — starting fresh"*, quietly, with the mechanics demoted underneath it: the seat asked to resume its saved thread, the vendor answered in a new conversation instead and reported success, the reply below is real and the history behind it is not, and the next brief continues from this turn. **No warning mark**, for the reason `settleRestoredThread` states: this is the same fact `reattachCard` says calmly at idle when no thread came back, discovered a turn later, and spending the ⚠ on it blunts the mark that carries real failures. The seat's probation ends here too — the restored id is gone by evidence rather than by a turn's outcome, so `settleRestoredThread` has nothing left to decide and a later failure on the *new* thread cannot be blamed on a reattach that was already reported. **The fixture is derived, and it is labelled derived.** `testdata/agy-forked-conversation.jsonl` is the real 1.1.11 capture with one textual substitution — every `conversation_id` value moved from `2222…` to `3333…`, nothing else, not a key and not a token count. It sits **outside** `testdata/wire/`, whose contract is real captures only: the forked turn itself was measured, but that probe's stream was not kept, and a hand-edited file among the captures would silently restate a measurement nobody re-ran. What it proves is the narrow thing it can: a turn that looks entirely successful still delivers the mismatched id to the room. **What was not done, and is not claimed: no live agy turn was driven for this change.** The behaviour rests on the 2026-08-09 capture, and the code rests on that capture's fixture. A re-measurement against a later build is what would retire `SilentResumeForkMeasuredAt`, and until somebody runs one, this seat's claim names 1.1.11 and no other build. #### 2026-08-16: the agy tail, measured — and why this seat still names no end of turn §9.33's amendments settled the codex seat on `turn.completed` and left this seat alone. The second of them said why in one sentence: **"this vendor's linger is not measured at all"**. That sentence is now false. This block is the measurement that replaces it, and the verdict is a measured **no**: the marker exists, the tail does not, and `vendors/agy.go` keeps its behaviour. **Instrument and version.** `agy 1.1.13`, read from `agy --version` at run time rather than from this file. agy self-updates, so a version quoted from a document is a version nobody checked; §3.8's re-verification records the same discipline. The probe ran this seat's own argv (`--output-format stream-json --disable-slash-commands --print-timeout 30m -p `) on a brief-shaped prompt, in a throwaway directory outside any repository. It recorded the arrival time of every stdout and stderr line, then the process exit. It polled `%USERPROFILE%\.gemini\antigravity-cli` every 250ms for size and mtime changes. It read no file content, and every prompt told the model to use no tools. **Three trials, not two.** Trials 1 and 2 are one-sentence replies. Trial 3 asks for about 600 words, because a tail that scales with the reply or with the conversation database would not show itself in two short turns. | | trial 1 | trial 2 | trial 3 (long reply) | |---|---|---|---| | `init` | 12.756s | 4.380s | 3.403s | | first `agent_response` delta | 14.181s | 6.005s | 9.487s | | `agent_response` DONE | 14.181s | 6.005s | 10.091s | | `checkpoint` | 14.603s | 6.606s | 10.694s | | **`result` (the answer)** | **14.603s** | **6.606s** | **10.694s** | | process exit | 14.917s | 6.655s | 10.829s | | **tail** | **0.314s** | **0.049s** | **0.135s** | **The marker half of the question is yes.** On all three trials `result` is the LAST line on stdout, stderr stays empty throughout, and the line carries both the full `response` text and the `status`. The adapter already parses it. A reliable answer-complete marker exists on this seat. **The tail half is what fails.** 0.049s, 0.135s and 0.314s, against the 4.06s and 4.25s §9.33 measured on codex the same day. `EndsTurn` would move two things and neither survives those numbers. The turn clock would stop at `result` rather than at the exit, which corrects at most 0.314s on a clock that renders whole seconds. The column would settle at `result` and read `exiting` until the exit, which puts a status word on screen for a third of a second at most. §9.33's settle exists to remove a false `streaming` that ran for four seconds. It does not exist to add a true `exiting` that nobody can read. So this seat keeps the process exit as its end-of-turn signal, and the decision is now written down rather than left as an unexamined default. **Nothing rides this tail either.** The vendor's own state writes were polled rather than diffed, so this is an observation of when writes stopped. On all three trials the final write to the conversation database, to the four transcript files under `brain//.system_generated/logs/`, and to the CLI log all land within one 250ms poll of the exit. There is no gap between the last receipt and the death, because there is no gap to hold one. **What is NOT claimed.** - **The failing turn is unmeasured.** All three trials ended `status: "SUCCESS"`. §9.33's second amendment settles a *failed* agy turn on its `status: "ERROR"` line, and that path's tail is not covered here. Producing a failure on purpose needs either a posture flag this adapter no longer passes (ADR-008's seventeenth amendment) or a lost thread this CLI answers with success (§9.43), so no probe reached one. - **The tool-using turn is unmeasured**, for §9.33's own reason: a probe gets no write access. - **The size is not a constant.** Three trials, one box, one build, one day. - **The boot is not the tail.** `init` lands 3.4s to 12.8s after the spawn. That is the operator's wait and the seat bills it honestly. It is named here only so a later reader does not mistake a slow start for a linger. **The same capture refutes a claim in the adapter, and the claim is corrected rather than left.** `ParseEvent`'s doc comment said a whole `agent_response` arrives as ONE delta when the step turns ACTIVE, plus a trailing newline on DONE — therefore a `PhaseWaiting` case and never a streaming one. At 1.1.13 both halves are wrong. A short reply sends no ACTIVE line at all, and one DONE step carries the whole text. A long reply sends true incremental deltas about 200ms apart (trial 3: 9.487s, 9.687s, 9.889s, with the final chunk on DONE), and each delta continues where the last one stopped, so nothing duplicates. **No code changes for this.** The seat declares `GranFinalOnly` and `applyEvents` promotes the phase to `streaming` on the first chunk, which is the modest-claim rule working exactly as it was written. Only the comment was stale. **Spend:** three billed turns, all trivial prompts, no tools. #### 2026-08-16: what agy reports when a turn needs the operator, and why that is `CapNone` §9.40's needs-you strip reads council's own gate queue and nothing else. The 2026-08-15 research asked whether agy's vendor-REPORTED `agent_state` (§2.1) could become a second source for it. This block is the measurement. The verdict is a `CapNone` refusal with **two independent reasons**, and each reason is sufficient on its own. **Instrument and version.** `agy 1.1.13`, read from `agy --version` at run time. agy self-updates, so a version quoted from a document is a version nobody checked; §3.8's re-verification records the same discipline. Windows 11. Six billed turns ran through this seat's own argv (`--output-format stream-json --disable-slash-commands --print-timeout -p `), from a throwaway directory outside any repository. The probe timestamped every stdout and stderr line, then the process exit. It then read the resulting transcripts for structural fields only: `type`, `status`, `step_index` and `created_at`. It copied no credential, and no prompt content from this machine enters this document. **Three shapes, two trials each.** Shape A is a normal completing turn and sets the baseline. Shape B drives agy to its `ask_question` tool. That is the pure needs-input case: the agent asks the operator a question and writes nothing. Shape C asks for a file write, which is the permission case. **The first reason: the waiting state never occurs in print mode.** agy answers on the operator's behalf, and it says so in its own words on both paths. - **A question is skipped.** In both shape B trials the agent asked, received no answer, and continued inside the same turn. Its own next message reads: "I've presented the prompt to select a file to rename, but it looks like you skipped it." Both turns ended `status: "SUCCESS"`, at 6.9s and 5.2s. Neither one waited. - **A tool permission is auto-approved.** Shape C trial 2 recorded a `SYSTEM_MESSAGE` step carrying this text verbatim: `stop hook blocked termination due to reason: The user has automatically approved the artifact through their review policy. Proceed to execution.` The write landed. `init.permission_mode` reads `request-review` on all six turns, so that value does not mean the operator is asked. - **`ask_permission` has never been reached.** It is one of the 56 tools in the `init` tools array. Across 90 conversations and 4,344 transcript records there are **zero** `ASK_PERMISSION` records. `ASK_QUESTION` has five. A gauge fed from this seam would therefore have nothing to report. `--print-timeout` is not the ceiling `agy.go`'s comment describes either. No probe turn reached it, because no probe turn waited. **The second reason: the print-mode stream cannot name the state it does emit.** The disk keeps the record type. The stream discards it. The two surfaces cross-walk by `step_index` exactly: | shape | idx | stream `step_type` / `state` | `tool_name` | `duration_seconds` | disk `type` / `status` | |---|---|---|---|---|---| | A (baseline) | 1 | `unknown` / DONE | absent | 0.0015, 0.0011 | `CONVERSATION_HISTORY` / DONE | | B (needs input) | 1 | `unknown` / DONE | absent | 0.0010, 0.0010 | `CONVERSATION_HISTORY` / DONE | | B (needs input) | 3 | `unknown` / DONE | absent | 0.5977, 0.5945 | **`ASK_QUESTION`** / DONE | | C (write) | 3 | `tool` / ACTIVE then DONE | `write_to_file` | 13.5304, 0.5917 | `CODE_ACTION` / DONE | **The needs-input step and the every-turn preamble step are byte-identical in every field this adapter reads.** Both arrive as `step_type: "unknown"`, state DONE, with no `tool_name` and no `tool_info`. Only `duration_seconds` separates them, and that is a continuous measurement rather than a marker. A threshold over it would be the invented vocabulary §4a.1 forbids. This is the `grok` problem the research named, in its worst form: waiting and working do not merely share bytes, because the vendor resolved the wait before it wrote the line. **No liveness signal pairs with it.** The status vocabulary across the whole corpus is exactly two values, `DONE` (4,180) and `RUNNING` (164). §3.8's re-verification already refused `RUNNING` as liveness, because its oldest rows sit in conversations nothing has touched for days. All five `ASK_QUESTION` records carry `DONE`. Three of those five predate this probe and come from real interactive sessions on 2026-08-09 and 2026-08-15. No record has ever been observed in a state that means an ask is outstanding. The `RUNNING` count is also unchanged from the 2026-08-15 re-read, which is the independent check that six completed probe turns leave no `RUNNING` residue. **`agent_state` is not on this surface at all.** It is a statusline-payload field (§2.1), and that payload exists only inside an interactive agy session. Council drives this vendor in print mode and never sees it. The candidate seam and the strip that would consume it sit on different surfaces, which is the structural half of the refusal. **One vendor flag behaves differently than assumed, and it changes nothing today.** Shape C trial 1 passed `--mode plan` beside the seat's `--disable-slash-commands`, and agy answered on stderr: `warning: --mode plan has no effect while slash command expansion is disabled.` The seat does not pass `--mode` (ADR-008's seventeenth amendment), so no behaviour moves. It does sharpen that amendment: the read-posture flag it dropped was inert twice over on this argv. Trial 2 re-ran with plan mode live and the write still landed, which corroborates `PARITY.md`'s Antigravity row at 1.1.13. **A claim in `agyPlumbing` is refuted by this capture, and the code was deliberately not touched.** That comment argues `unknown` is a fixed preamble slot agy declines to name, on the evidence that every observed one sits at `step_index` 1, carries no tool name, and lasts under 5ms. Shape B produced a second `unknown`, at index 3, carrying a real act. The suppression still reaches the right outcome today, because agy skipped the question before the line arrived and there is nothing actionable to draw. The stated reason is now wrong. This lane measures rather than changes code, so the correction is recorded here and is owed in `internal/council/vendors/agy.go`. **What is NOT claimed.** - **The interactive seam is unmeasured.** §3.8 observed `agent_state` transitioning, and `tool_confirmation_pending: true`, on agy 1.1.9 statusline payloads. A re-capture at 1.1.13 needs a TUI session and a statusline capture command, and this probe changed no agy configuration. Whether an interactive `ASK_QUESTION` sits `RUNNING` while it waits is the open question this block leaves behind. It is the one measurement that could reopen the seam. - **Six turns, one box, one build, one day.** - **The model picked its own write path in shape C trial 1.** It ignored the process cwd and wrote under `~/.gemini/antigravity-cli/`. Both files were removed after the run. That is model behaviour rather than a seam property, and it is named so a later probe expects it. **Spend:** six billed turns, all short prompts. ### 9.44 the composer was a gap under a rule, and the room's state floated below it (2026-08-09) **Inspiration is named because it should be: Grok's CLI.** Its input is a rounded, clearly bordered box, and the bottom border carries a right-anchored legend — `Grok 4.5 (high) · always-approve` — laid *on* the line rather than under it, with the remaining key hint (`Shift+Tab:mode`) on a muted line below. The thing that reads well there is not the corners. It is that the box says **where you act**, and its own frame says **what you are acting under**, which leaves the line below free to be nothing but keys. Council had neither. §9.26 closed the frame with two full-bleed heavy rules, and the composer sat in the gap between the lower one and the mode line — a prompt glyph on an unbounded row, with no mark anywhere saying that this strip of the screen is the one place typing does anything. Every other region in the room is a reading area. The one region that is an *input* was the only region with no shape of its own. **So the composer is a box, and it is the only bordered element on screen.** Rounded corners (`╭ ─ ╮ │ ╰ ╯`, and `+ - |` in the reduced set), a side and a cell of air on each row, and the bottom border carrying the legend. Bordering exactly one thing is the whole design: a second box anywhere would make this one a decoration instead of a signal, the same scarcity argument §9.26 made for a second rule weight and §9.28 made for the seat hue. **The lower heavy rule is gone, and that is a correction to §9.26 rather than a cost of this change.** §9.26's claim was that the two full-bleed rules were the only *closed shape* on screen and that a closed shape earns the second weight. They were never closed — two horizontal lines with nothing joining their ends is a pair of lines, and the weight was doing the work a shape should have done. The box actually closes, by corners and sides, so it draws **light**: closure moved from ink to geometry, which is a carrier that survives `NO_COLOR` outright. What the heavy rule now says is narrower and true — *the chrome stops here and the seats begin* — and there is exactly one line in a grid where that holds. **The legend is split by lifetime, not by topic.** On the border go the facts that stay true until a key changes them: the mode word (`VIEW` / `COMPOSE` / `GATE` / the page label) and, when the guard is off, `a not asking`. Under the box go the keys, which change with the mode, the draft and the turn. Nothing appears in both places — `statusLine`'s left-hand slot is now empty and is deliberately not backfilled, because a slot that survives its content is how a footer becomes the wall §9.11 spent a whole pass taking apart. The cadence cell keeps its **key** on the way up, not just its words. `a not asking` was added (§9.24-era footer work) to close the §9.17 defect where a permanently ungated room documented the way back nowhere on screen; moving the words without the key would have reopened it one release later. It is on the border, whole, and unsheddable. **What it costs, stated plainly.** One row of body — `promptChrome` goes from 2 to 3, since the box replaces the rule with a top border and adds a bottom one — and six cells of composer width, which is `boxChrome`: a side glyph plus `gutter` cells of air, on each side. The air is `gutter` rather than one cell because **the room spells its separator one way** — two cells each side of every `│` it draws, which `TestTheRoomSpellsItsSeparatorOneWay` already holds for the header, the key line and the column rails. A box welding its prose to its own sides would be a second grammar for the room's only vertical mark, on the element the eye is meant to read as the frame. **The help panel paid the row, the way it always pays.** Its budget was 17 lines to the pinned `?` and is now 16, which pushed `ctrl+c / q` off page one and the WORKSPACE sentence — the load-bearing line — off page two at the reference machine's 24-row room. Neither was allowed to fall: both pages spend the blank row directly under their title instead, on §9.11's own ranking that **a rule outranks a blank**. The title *is* a labelled rule, so the blank beneath it was the one row on each page restating a boundary the row above already drew. `MinWidth` (60) and `MinHeight` (10) are unchanged and still resolve: at the floor the room is 2 header rows, 3 of footer chrome, a one-row composer and four rows of reading area. When the frame is too narrow to lay the legend on the border with air each side and a rule cell outboard of it, the border **closes bare** rather than truncating — a legend cut in half is a claim cut in half. **One golden moved further than the box did, and it is a gain rather than a surprise.** Emptying `statusLine`'s left slot gives the key line back four cells and its gap, and at the tabbed tier that is enough for `1-N seat` to stop shedding — `empty-tabs.txt` now names a key that always worked and had been dropped for width since §9.29. Nothing was un-shed by hand; the ladder in `hint.shed` is unchanged and simply has more room to not use. **No new hues.** The border and the sides are `Rule()`, i.e. muted chrome; the legend keeps the exact styles the mode word already had on the mode line, gate included. This change spends *shape*, which the palette does not pay for. ### 9.45 the turn clock counted the operator's reading time as the vendor's work (2026-08-15) **A gated column said `⋮ streaming 5m` while nothing was streaming.** The room was stopped on an approval card, the vendor was blocked waiting to be told yes or no, and the five minutes on the header were five minutes of a person reading a diff. The number was real wall clock and it was still a false reading, because of the word it sat under: `streaming` is a claim that output is arriving, and this column had a stopped process behind it. That is the same failure `TestWaitingIsNotStreaming` was written for, arriving by a different route. That test guards the WORD — a seat with nothing to show must not render like a seat that is showing something. Nothing guarded the FIGURE under the word. A stopped seat wearing a moving seat's clock is the honest-gauge rule broken in the one place §4a.1 cares about most: the displayed value no longer describes the thing it is labelled with. **So the turn clock splits in two, and both halves are measured.** The column header states the VENDOR's own time — wall clock minus whatever of it the operator held — and the turn's separator states the operator's share beside the number it came out of: ``` ▸ 1 CC Claude Code ⠋ streaming 12s ro:tools tokens ⚠ waiting on you: Write: internal/council/clock.go y approve n deny a stop asking … turn 1 ───────────────── you 4m48s ``` Twelve seconds of vendor, four minutes and forty-eight seconds of operator, five minutes of wall clock — and the two figures add up to it, which is the property that makes the split worth drawing rather than merely correct. **The stopwatch runs over the SEAT, never over the card.** One assistant message can raise a parallel batch of requests and each one blocks separately, so a seat can have three cards up at once — and it is stopped ONCE. Summing the cards would bill one person's one wait three times. So `PendingGate.StoppedAt` is when the seat stopped, not when the card was raised: the first card of a stretch carries its own moment, every later card in the same stretch inherits that stamp, and the stretch closes when the LAST card goes. Stamping each card with its own moment was the first cut and it was wrong twice — the figure on screen jumped backwards when the first of two cards was answered, and the charge lost every second before the newest card, because the stretch's start left the queue with the card that owned it. **Stretches accumulate across one turn, and reset with it.** `Column.GateWait` is the operator's share of the turn in flight, `TurnRecord.GateWait` is the same fact for a turn in the transcript, and `startTurn` files one and clears the other. A turn that asked three times reports all three waits as one figure, because what it claims is the operator's share of THAT turn. **Auto-approved calls contribute nothing.** `queueGate` answers three ways without ever drawing a card — `autoApproveRoutine` (this shell command is routine), `isReadOnlyTool` (this tool changes nothing), and `!Asking` (the user said stop asking) — and all three return before the stamp. Nobody was asked, so nobody waited, and charging the operator for a decision they never saw would be the room inventing a measurement. **Zero and absent stay different, as they must.** The figure is a `runner.Span` rather than a duration, for the argument already written on that type: a turn that raised no card is UNMEASURED and renders nothing at all, while a card answered inside a second is a measured zero and renders `you 0s`. `gate-clock.txt` pins the three states side by side — an open card counting, an absent one, a measured zero — because each is only legible against the others. **Two spellings, one fact, and the surface picks.** `waiting on you 4m48s` is the room's own phrase, already on the approval card and on §9.40's `NEEDS YOU` strip, and it is twenty cells. A three-up room at 120 columns gives each column thirty-six, where the long form is more than half the width and pushes the separator's own clock and cost off the line. So the grid sheds the LABEL and keeps the fact — `you 4m48s` — which is §9.18's order applied to a phrase instead of a name, and the by-turn page, which is the full frame wide, says it whole. It is a width TIER rather than a per-line measurement, the same way `stripHeader` and `stripBadges` choose a form: one rule a reader can learn, instead of a cell that rewords itself when a neighbouring number grows a digit. **Why the separator and not the chrome.** The live turn's separator carried its number and nothing else, on the rule that a turn's clock and cost are in the header and the badge line and repeating them a row later would say one thing twice. The operator's share is the exception, and it is there because the chrome has no room for it: `▸ 1 CC Claude Code` and `⠋ streaming 12s` already spend thirty-three of thirty-six cells, and the badge line's right edge belongs to the cost. The separator is also where the figure lives for every turn already in the transcript (`historyMeta`), so the live turn and the filed one state it in one place and one spelling. What it costs is that the figure scrolls with the transcript — and what stays pinned in the chrome is the card itself, which says the room is stopped on you for as long as that is true. **Render stays pure.** The room stamps and the renderer subtracts, exactly as `Reattach.SavedAt` does: `queueGate` stamps when the card goes up, `decideGate` and `dropGates` charge when it comes down, and Render turns a stamp into an age against `State.Now`. An open card with no stamp — every State a test types out by hand — adds nothing and does not make the span measured, because a duration arrived at by arithmetic over an absence is the invented figure §4a.1 puts at the top of the rejected list. `TestGateClockIsPureOverState` pins it, and no existing golden moved. **The `--trace` line is deliberately unchanged.** `runner.TurnClock` already says a seat blocked on an approval card is inside its `Stream` span, "because from the process's side that is exactly what it is". That stays true: the runner cannot see the decision. A gate decision reaches a live seat through `Session.SendAside`, which is documented as carrying "an interrupt, a gate decision, a protocol reply" — undifferentiated bytes — so a gate span in the trace would need the runner to be told what it was writing, which is a change to that package's shape rather than to this figure. The room knows, and the room is where the split is drawn. Adding the span to the trace is a separate measurement with a separate seam to build. **Wall clock is still reported where wall clock is the claim.** `turnElapsed` — the by-turn page's figure for how long the whole turn took — is unchanged and still selects the longest seat's raw elapsed. That number is labelled as the turn's duration rather than as any vendor's, and the operator's own reading time really is part of how long the turn took. The per-seat rule under it carries the split, so the page states both without either one contradicting the other. **Amended 2026-08-29: the word was left behind, and the split alone could not fix the reading.** This section opens by naming the defect as a WORD problem — "`streaming` is a claim that output is arriving, and this column had a stopped process behind it" — and then corrects only the number. The result was `⠋ streaming 12s` on a seat with a blocked process behind it. Twelve seconds is the honest figure and `streaming` is still a false claim, so a reader who scanned the header still read a working seat. A corrected number under a wrong word is a wrong reading. **While a card is up, the header says `needs you` and states no clock.** The word is `needsYouWord`, which is `needsYouLead` in lower case rather than a second spelling of it. One state now has one vocabulary in three registers: `NEEDS YOU` on §9.40's strip, `waiting on you` on the card and on the long form of the operator's own figure, and `needs you` in the column header, where every state word is lower case. The mark is `Warn`, the card's own glyph. The spinner is this room's only moving cell and it means a turn in flight (§7.1 rule 4), so a spinner over a stopped process makes the same false claim the word did. Colour is `SevWarn` and carries nothing extra: the phrase survives `--ascii` and `NO_COLOR` on its own. **The clock goes because neither figure is time spent in this state.** The vendor's twelve seconds are frozen for as long as the card is up, so the number is not moving and it does not describe what the seat is doing. The operator's four minutes are moving, and they already have one home — the turn's own separator, where `historyMeta` states them for every filed turn as well. Printing them in the chrome too would put one fact in two places on one screen. What the header loses is a figure that had stopped; it returns the moment the card is answered, and `TestWaitingOnYouIsNotStreaming` asserts that arm rather than trusting it. **Nine cells, so no layout moves.** `needs you` costs exactly what `streaming` and `cancelled` cost, which is the width `stripColumn`'s floor is derived from (`layout.go`). The strip header takes the same substitution, because a folded seat can hold an unanswered card and a strip is the one width where the reader has no card beside the word to read it against. **Three conditions, and the queue is the only source.** `stoppedOnYou` requires an installed seat, a turn in flight, and a card in `State.Gates` for that vendor. `State.gateStopped` answers the last one, and it deliberately differs from `gateStoppedAt`: that function skips an unstamped card because a duration derived from an absence is the invented figure §4a.1 rejects, while this one reports the seat stopped, because the card's existence establishes that on its own. Every fixture in the package is unstamped and every one of them draws the card, so a stamp-sensitive predicate would have put the header back in contradiction with the card two rows under it. **The by-turn page takes the same word.** `seatMeta` states one seat's turn on one rule, and a page saying `streaming` while the grid says `needs you` would be two surfaces disagreeing over one queue. It applies to a LIVE entry only. A filed record's phase is how that turn ended, and the queue only ever describes now. **`gated-vs-streaming.txt` is a new golden rather than an extension of an existing one.** It puts a blocked seat and a working seat on one frame at the same wall clock, because the reader's question is never "what does a blocked seat look like" — it is "which of these two is running". `waiting-vs-streaming.txt` pins two claims about a VENDOR; this pins the same shape of claim about the operator, and neither case is an edge of the other. ### 9.46 cursor hooks report a blocked action, but never report a request to a human (2026-08-16) **Environment and evidence class.** The survey drove `cursor-agent` **2026.08.11-e8db854** on Windows 11, from a PowerShell parent. This build is NEWER than the build every other cursor record in this document pins (`2026.08.04-aaa8809`), so read the tables below as the current build and the older records as the older build. The event catalogue and the configuration paths come from a source read of the installed bundle. Every claim about what fires comes from a live run. The rig put a recorder on each event. The recorder appended the raw stdin payload, a UTC timestamp and the leading bytes to a per-event log. **Why the survey ran now.** §1 recorded the needs-input seam as a watch item and named Hooks as the supported surface for it. §8's roadmap items 3 and 5 carry the same item. Nobody had asked the narrower question: does any hook event mark NEEDS INPUT, as opposed to marking the absence of completion? The answer decides whether the future needs-you strip can source this vendor. **The catalogue, from the installed build.** `index.js` holds the event registry as a single map. It carries **21** events: ``` beforeShellExecution beforeMCPExecution afterShellExecution afterMCPExecution beforeReadFile afterFileEdit beforeTabFileRead afterTabFileEdit stop beforeSubmitPrompt afterAgentResponse afterAgentThought sessionStart sessionEnd preCompact subagentStart subagentStop preToolUse postToolUse postToolUseFailure workspaceOpen ``` Before this survey the repository knew three of these names, plus `preCompact` named but never used. **The Claude compatibility map has exactly two holes, and they are the two that matter.** The same file maps Claude Code's hook events onto cursor's own, for the imported-configuration path. Eight map across. Two map to `null`, and the bundle lists them together as the unsupported pair: | Claude Code event | cursor event | |---|---| | `PreToolUse` | `preToolUse` | | `PostToolUse` | `postToolUse` | | `UserPromptSubmit` | `beforeSubmitPrompt` | | `Stop` | `stop` | | `SubagentStop` | `subagentStop` | | `SessionStart` | `sessionStart` | | `SessionEnd` | `sessionEnd` | | `PreCompact` | `preCompact` | | **`PermissionRequest`** | **`null`** | | **`Notification`** | **`null`** | `Notification` is the Claude event that fires when the agent waits for a person. `PermissionRequest` is the other one. Cursor's own catalogue has no equivalent of either, which is why the import drops them rather than renaming them. This is the verdict in one line, read off the vendor's own table. **Where cursor-agent reads hook configuration.** The loader reads seven paths. Four are cursor's own and three are Claude's: | scope | path on Windows | |---|---| | enterprise | `C:\ProgramData\Cursor\hooks.json` | | team | `\.cursor\managed\active-team-hooks\hooks.json` | | user | `~\.cursor\hooks.json` | | project | `\.cursor\hooks.json` | | claude user | `~\.claude\settings.json` | | claude project | `\.claude\settings.json` | | claude project local | `\.claude\settings.local.json` | The three project-scoped entries sit behind a boolean in the loader. The loader also refuses any config path that contains a symlink. The project scope is what let this survey run without changing the operator's own configuration for most of its arms. **What fires on which path.** Two trials per arm unless the table says otherwise. A dash means the event did not fire on any trial. | event | print mode (`-p --trust`) | ACP, project scope | ACP, user scope | |---|---|---|---| | `workspaceOpen` | fires | — | — | | `sessionStart` | fires | — | — | | `preToolUse` | fires | — | fires | | `beforeShellExecution` | fires | — | fires | | `afterShellExecution` | fires | — | fires | | `postToolUse` | fires | — | fires | | `postToolUseFailure` | fires, on a hook denial only | — | not observed | | `sessionEnd` | fires | — | — | | `afterAgentThought` | fires, and kills the turn | — | not tested | | `beforeSubmitPrompt` | — | — | — | | `afterAgentResponse` | — | — | — | | `stop` | — | — | — | | the other nine | — | — | — | **Two results in that table are new, and one confirms an older record.** First, **ACP honours the user scope and ignores the project scope.** A project-scoped config fired nothing at all on ACP, over two trials, not even `sessionStart`. The same file fired eight events in print mode. This agrees with the loader's gate and with `PARITY.md`'s row that the ACP protocol has no workspace-trust step. Second, **ACP fires no lifecycle event.** `sessionStart`, `workspaceOpen` and `sessionEnd` are print-mode only, so a needs-input consumer on ACP gets tool events and nothing that brackets the session. Third, `afterAgentResponse` still does not fire on ACP, which confirms §7.16's 2026-08-15 amendment at this newer build. **What a blocked moment looks like in bytes.** A hook that returns `permission: "deny"` replaces `afterShellExecution` and `postToolUse` with one `postToolUseFailure`. That payload is the only affirmative block marker on the whole seam: ```json {"tool_name":"Shell","error_message":"Command execution was blocked by a hook: telltale-seam-deny …","failure_type":"permission_denied","duration":0,"tool_use_id":"…","is_interrupt":false, "hook_event_name":"postToolUseFailure","cursor_version":"2026.08.11-e8db854"} ``` `failure_type` is a free string, not an enum, and the hook path writes only two values into it: `permission_denied` when a hook denies, and `error` when a fail-closed hook errors. Both describe a refusal that already happened. Neither describes a wait. **The awaiting-human moment is where the seam collapses.** Two arms produced a real human decision point. In print mode a hook returned `permission: "ask"` with no person present. On ACP the client answered `session/request_permission` with `reject-once`. **Both produced the success-shaped event sequence** — `preToolUse`, `beforeShellExecution`, `afterShellExecution`, `postToolUse` — with no `postToolUseFailure` anywhere. The only difference from an allowed command is one field: | arm | `afterShellExecution.output` | |---|---| | allowed | `"[ERROR] - (starship::print): Under a 'dumb' terminal (TERM=dumb).\n\r\nseam-probe\r\n"` | | `ask`, nobody to ask | `""` | | human rejected over ACP | `""` | An empty `output` is also what a silent successful command writes. So the seam encodes "a person was asked and said no" and "the command printed nothing" with the same bytes. That is §4a.1's zero-versus-absent collapse, arriving from the vendor rather than from a render path, and it is the reason a needs-you strip cannot be built on these events. **The affirmative signal exists, on the other seam.** The ACP wire carries it plainly. It is a blocking JSON-RPC request from the agent to the client, and the turn stops until the client answers: ```json {"jsonrpc":"2.0","id":0,"method":"session/request_permission","params":{"sessionId":"…", "toolCall":{"toolCallId":"…","title":"`echo seam-probe`","kind":"execute","status":"pending", "content":[{"type":"content","content":{"type":"text","text":"Not in allowlist: echo"}}]}, "options":[{"optionId":"allow-once","name":"Allow once","kind":"allow_once"}, {"optionId":"allow-always","name":"Allow always","kind":"allow_always"}, {"optionId":"reject-once","name":"Reject","kind":"reject_once"}]}} ``` This names the pending tool call, gives the reason, and lists the choices. It is everything a needs-you strip wants. It is **not a hook**, and only the process that drives the ACP session receives it. The council seat already reads it (`acpPermission` in `cursoracp.go`). A passive gauge cannot, because there is no file and no second reader. After the client rejects, the wire reports `tool_call_update` with `"status": "completed"` and no `rawOutput`. That is the fourth shape §7.16's amendment predicted and could not attribute; this survey attributes it, because the rejection here was the client's own and was known in advance. **The verdict.** **No cursor hook event marks NEEDS INPUT.** The seam reports a completed refusal (`postToolUseFailure`, `failure_type: permission_denied`) and it reports completion. It never reports a wait. The vendor's own Claude-compatibility table says the same thing by mapping `Notification` and `PermissionRequest` to nothing. For this vendor the needs-input signal lives on the ACP wire, in `session/request_permission`, and it is available only to a seat that drives the session. §8's roadmap item 3 should be read against that: cursor's entry belongs under the council seat, not under the hook relay. **Two traps for whoever builds on this seam.** - **The BOM is on every event.** Every payload captured here began `EF BB BF 7B` — a UTF-8 BOM, then `{`. `cursorhook.Parse` does not strip it, and does not need to today, because `afterAgentResponse` is the one event that does not fire on the paths telltale drives. Any new event routed into that parser fails on the first byte. `internal/cursorstatus/stdin.go` holds the working strip. - **A command hook on `afterAgentThought` kills the turn.** Three turns registered it and all three died with `RetriableError: WritableIterable is closed` after the tool call. Six turns without it completed. A compiled recorder in place of a PowerShell one changed nothing, so this is not hook latency; the event fires inside the response stream and registering it breaks that stream. **What is NOT claimed.** - **The interactive TUI is unmeasured here.** This survey drove print mode and ACP only. §7.16's per-surface table already records that the interactive console behaves differently, so the dashes above are measured absences on two paths, not a claim about a third. - **`stop` never fired on either path, and its interactive behaviour is unknown.** It is the natural turn-complete event and it stayed silent, which is worth knowing before anyone designs against it. - **The ACP rejection arm has one clean trial, not two.** The second trial lost its connection mid-turn, a vendor-side failure this survey saw on other arms too. The wire shape matched on both; the full event set was captured once. - **No subagent, MCP, compaction or tab event was exercised.** They are listed above because the build declares them, not because anything drove them. **The operator's own configuration was restored.** Most arms used a project-scoped `\.cursor\hooks.json` in a throwaway directory. The two ACP user-scope arms needed `~\.cursor\hooks.json`, because ACP ignores the project scope. That file was copied first, the copy was verified at SHA256 `3A9F05582D99DFEEB95E705559789F3B41D01DF1292F811D1A94834A54DFCB3C` (146 bytes), the test configuration preserved telltale's own `afterAgentResponse` entry, and the original was restored and re-verified at the same hash and length. No credential store was read or copied at any point. ### 9.47 the room raced fourteen times and could not say who won (2026-08-29) `/arena` has been building a record since it shipped, and nothing could read it. Every race leaves an `arena/t/` branch per seat and every adoption leaves an `adopt/t-`; both outlive the room by design (§9.37's kept-until-deleted ruling), and `arenaRaceNumber` already reads the first namespace to number the next race — "the refs are the one record that shares the leftovers' lifetime". What the ROOM kept of a race was `Column.Arena`, a per-turn fact the next dispatch clears, and `TurnRecord` never carried it. So a repository holding fourteen races and nine adoptions could not answer *which seat do I actually take*, and the operator answered it by reading `git branch` in a second terminal — which is the one thing §9.17 says a command surface exists to remove. **`/arena record` reads the refs and states the standing, one line per seat.** It is a verb inside `/arena` rather than a new room word, and that is a budget decision with a name: `refuseUnknownCommand` prints the whole vocabulary on one line against a hard width, and that line's own comment records `/adopt` as "the last cheap one — the next verb has to find its characters somewhere else". `/arena drop` had already established the shape (a sub-verb, the exact form only, anything longer races as prose), so the record costs the refusal nothing and the help panel nothing. It is taught the way `x`, `/adopt` and `/arena drop` are: by this section, and by the notice the command itself prints. #### Derived from the refs, never stored **Nothing is written and no new file exists.** The obvious build was a counts file under `~/.telltale`. It was rejected before it was written: `CLAUDE.md` enumerates the writes the gauges are ratified to make — three relays and the event sink, each with a test pinning its serialized form — and a fourth exception is an owner-level edit to that contract, not a feature's side effect. A tally over refs the repository already holds needs no such grant, and it cannot go stale against them, because it IS them. Two `git for-each-ref` scans, over the two namespaces `freeAdoptBranch` and `arenaRaceNumber` already scan, through the same `gitOut` argv. **The read happens in the command handler and the page renders from State**, exactly as `ArenaResult` is computed in `finishColumn`. A body whose content came from a subprocess inside `Render` would make every golden depend on the repository the tests happen to run in, which is the purity rule `TestRenderIsPure` exists to hold. **The verb is not refused mid-turn**, and that is what separates it from the other two arena verbs. `/adopt` and `/arena drop` mutate worktrees a race is writing; this one reads refs. A record read during a race is a measurement of a moment already past, which is what every other reading in this room is. #### What the refs can say, and the three things they cannot This is the honesty boundary of the feature, and the page states both halves rather than implying the first. - **They CAN say who entered a race and whom the operator adopted from it.** An adopt branch exists only on an adoption that LANDED — `undoAdoptBranch` deletes the branch a failed one cut — so a surviving `adopt/t-` is a merge that happened, on the operator's own `y`. - **They CANNOT say a rank, a phase word, or that a seat was cut with `x`.** Those are turn-scoped and die with the room. **So this surface never claims a LOSS.** A seat that entered a race the operator never decided is counted as UNDECIDED and reported beside the rate, never inside it. A race with a give-up is exactly that case, and the give-up is the most probable ending a five-seat race has (§9.37's 2026-08-17 amendment) — folding it into a denominator would be the room scoring a seat for a race nobody judged, which is §9.22's refused "cross-seat agreement mark" arriving through a side door. **The narrower case is stated rather than solved: a seat CUT with `x` in a race the operator then decided for somebody else counts as one that was not adopted.** The refs cannot tell a stall from a worse answer, and no honest reading of them can. What answers it is the WORD on the page. It is `adopted`, never `won`, and that is the whole mitigation: `0 of 4 adopted` is literally true of a seat that stalled four times, and it makes no claim about why. A column headed `won` would make one, on evidence that does not exist. - **A dropped branch leaves the record.** `/arena drop` deletes an arena ref by design, and `adoptSeat`'s own notice offers that drop as the next command — so the winner's arena branch is the one most likely to be gone. Two consequences, both handled: an adopt ref is treated as evidence the seat ENTERED that race, which is what keeps a rate off the far side of 100%; and the page says it counts over the branches the repository STILL HOLDS, in a line of its own. **The window sentence is a LINE, not the rule's meta.** `labelRuleIn` drops a rule's meta whole when the width will not take it — correct for a count, wrong for the sentence that bounds the claim, which would then vanish exactly where the room has least room to make it. That is the act ledger's own ruling on its retention line (§9.22, amended 2026-08-17), and this page's claim is bounded the same way. The clipboard document carries it too, and there it matters more: a table pasted into a review a week later has nothing else saying these counts were ever bounded. **Only the refs this room minted are counted.** `arena/t/` and `adopt/t-` with `freeAdoptBranch`'s numeric collision suffixes, parsed back against the same two functions that write them. A hand-cut receipt is refused — the real one is `adopt/t9-claude-helpers`, which the first live adoption left behind before the 2026-08-11 ruling gave the verb a spelling of its own. That is an UNDERCOUNT, said out loud here rather than papered over, and it is the honest direction to be wrong in: a looser parse would credit a seat for a branch somebody merely named after it. It is `dropRacer`'s judgement about paths — "no state this room's arena created can have that name" — applied to refs. #### The three renders, and the one figure this page is entitled to compute - **A seat with no ref at all is ABSENT: `never raced`.** Not 0%. §4a.1's founding rule, on the surface where a zero would be read as a verdict about a vendor rather than as a count. - **A seat whose races were all undecided has NO RATE**, and its races are reported as undecided. Inventing a denominator out of them is the same error one step down. - **A seat the operator decided against is a MEASURED zero and prints one: `0 of 4 adopted 0%`.** The distinction between that line and the one above it is the whole feature. **The rate never appears without its count, and the count comes first.** The two counts are what was measured; the percentage is arithmetic over them, which §7.12 names as the one kind of computed figure this product may show — "telltale's measurement of telltale's own observations" — and that carve-out is conditional on the reader being able to check the division. A bare `67%` would be the claim with its evidence removed, which is the same defect as a total printed without its window (§7.15). The seat name takes that seat's own hue and nothing else on the line does (§9.28's ratified exception), for the turn page's reason: this is a stack of seats in one column, so position answers nothing about who is being described. Every distinction the page makes is carried by a WORD first — `never raced`, `no decided race`, `undecided`, `adopted` — so `--ascii` and `NO_COLOR` lose nothing. #### The keys, and what the page deliberately does not get `t` closes it, because `t` is already what this room means by "give me the grid back" from a full-frame body. A key of its own would be a second thing to remember for one act, and no key at all would be a body reached by a typed command with no keyed way out — the help panel's missing `?` with a whole surface behind it. The mode word is `RECORD`, against `TURN` and `ACTS`: §7.8's always-on statement of what is on screen has to tell the room's three full-frame bodies apart, and this is the one whose subject is neither a turn nor the keymap. It carries no coordinate because it has none. `y` takes the page, on §9.22's own argument — a copy key that took the focused column's reply from behind a body the reader is looking at would break the one claim that earns it a footer cell. `Y` is the same document for the same reason the page's is: there is no per-seat focus for a narrower key to address. **No scroll cell, and no scroll keys.** The record is one short line per seat, so the only geometry that can clip it is the height floor, and there `recordCell` draws the overflow marker instead. A footer naming arrows that move nothing is the false promise §7.8 forbids — the same reason `f` and `tab` are absent and are swallowed rather than left to change the grid invisibly. **Deliberately not built, and each one is a ruling rather than a backlog item:** - **Elo, or any rating.** The sweep that proposed this feature said it plainly and it is right: Elo is overkill for one operator's vote volume, and a rating is a number with no measurement under it. A plain adopted-of-decided tally is the honest shape. - **A blind-review mode**, the other half of the sweep's candidate. §9.34 states that the PERSON is never blinded — "columns stay labelled by vendor; the blind applies to what the models read" — so an opt-in user-blind arena inverts a stated position and is an owner's ruling, not a builder's. It is not built here and nothing here assumes it. - **Anything cross-repository.** The record is one repository's refs, because that is what a room is pointed at and what `/cd` moves. A tally across every repo the operator has ever raced in would need a store, which is the write this whole section refuses. - **A rank or a phase in the tally.** They are not in the refs. Carrying them would mean filing `ArenaResult` into `TurnRecord` and persisting it, which buys a richer number by taking on the store — and the number it buys is one the operator's own adopt decision already summarises. Verified offline. `record_test.go` pins the tally's arithmetic against hand-written ref lists (one race counting once however many refs it left, an adopt ref proving entry, the numeric collision suffix not double-counting), the refs this room did not mint being ignored, zero and absent staying apart, no rate ever printing without its count, an undecided race never reaching a denominator, the window sentence surviving the narrow width, an unreadable record rendering as unavailable rather than as an empty one, `t` giving the grid back without opening the turn page, the clipboard document agreeing with the screen and carrying the window, and the verb taking only its own word — measured in a read-only room, where a longer draft is refused by name and therefore provably reached the race path. `arena-record.txt` and its `--ascii` twin are the frame. The one test that touches git builds its refs with `arenaBranch` and `adoptBranch` in a temp repository, so the parser is asserted against the functions that write the names rather than against the strings the test types. No test here spawns a vendor. **Nothing in this section is a claim about vendor behaviour, so no live vendor run is owed.** What IS owed is one live open against the reference box's own leftovers — 27 `arena/t` branches and the `adopt/*` refs beside them are recorded in §9.37 — to confirm the counts a real pile of refs produces and that the page reads at the room's own geometry. Stated here rather than implied paid. ### 9.48 the race said what changed and never whether it worked (2026-08-29) `/arena` measures everything about an attempt except the one thing an operator adopts on. It reports what each racer CHANGED — the live stat, the settled `git diff --stat`, the full patch, the commit receipt (§9.37) — and `/arena record` reports which seat the operator TOOK (§9.47). Neither says whether the attempt WORKS. The room's own founding note admits it: rank is arrival order, and the only clock that ranks a race is the room's. So the operator answered the question by hand, once per seat, in a second terminal — `/cd` into each kept worktree and run the same command — which is the act §9.17 says a command surface exists to remove. **`/arena check ` names one command; every racer runs it in its own worktree, and each attempt's block says PASS or FAIL from that run's real exit code.** #### The grammar, and the two shapes it is not It is a sub-verb inside `/arena`, for §9.47's budget reason: `refuseUnknownCommand` prints the whole room vocabulary against a hard width, and that line's own comment records `/adopt` as "the last cheap one — the next verb has to find its characters somewhere else". `/arena drop` and `/arena record` had already established the shape, so a third costs the refusal nothing and the help panel nothing. It is taught the way they are: by this section, and by the notice the command itself prints. **It is NOT a file in the repository, and that is a ruling rather than a preference.** The obvious build was `.worktreeinclude`'s sibling — a `.arenacheck` the racer trees inherit — and arena.go's own seeding doc already refuses exactly that shape: agent-deck pairs seeding with repo-carried setup scripts, and council took "copy only, never execute", because a repository that can run a command on the machine by merely CONTAINING a file is a different product with a different threat model. A command a person typed into their own room is that person's act. A command a clone brought with it is not. The parked byte-level trust question is untouched by this feature, which is the point of not touching it. **It is NOT a new room word.** `/check` would have cost the refusal line a re-wording and, worse, a second meaning for a word this codebase already spends: the write gate, the gate cards, `gatehook.go`. Two facts cannot wear one word (§9.13) — which is also why the verdict is spelled `FAIL` and never `failed`. The room already spends `failed` on a phase, and a seat that finished cleanly while its check exited 2 is a different fact from a seat whose process died. **The cost of taking free text, stated the way `parseArenaDrop` states its own.** This is the one `/arena` sub-verb that cannot close its grammar with a length cap, so a brief opening with the word `check` is at risk. What protects it is a PATH lookup on the first word: a draft whose first word after `check` is not a program this machine can run is refused by name, handed back to the composer, and neither raced nor set — nothing spawns and nothing is billed. The narrow case that survives is a brief opening `check `; the notice names exactly what was set, and `/arena check off` takes it back in one line. A path-bearing first word (`./scripts/check.sh`) skips the lookup on purpose, because it is resolved against the RACER's tree rather than against the room's own directory. **The command is room state and survives `/cd`**, which is `/write`'s rule rather than an oversight: posture is room state and moving the workspace does not quieten it either. The mitigation is that every result names the command it ran, so a command left over from another repository is visible in the verdict rather than assumed behind it. **It does NOT survive the room.** `room.json` holds session ids and a workspace — keys and numbers, never content (`CLAUDE.md`'s read/write boundary) — and a command is neither. The consequence is deliberate rather than reluctant: a saved command would run in a session whose operator never typed it, which is the one property the "a person typed it into their own room" argument above rests on. Re-naming it is one line. #### Four rulings, and each is a line the code may not cross - **The exit code is the only source.** PASS is exit 0; FAIL is any other code; both are read back from the process. Nothing here infers a verdict from output, from the diff, from a duration, or from a model's opinion. **An LLM judge is refused by ruling twice over** — §9.2's refusal of "a ranking stage, a chairman, or any synthesis hop" and §9.44's declined cross-seat quality mark — and it stays refused even wearing an estimate's `~`, because `~` marks a figure telltale COMPUTED and an opinion is not a computation. A measured exit code is the opposite case, and it complies with ADR-001 exactly as the diff does. - **A command that could not run is its own state.** A missing binary, a tree that could not be entered, a run the deadline stopped: none of them is a FAIL. They render as `check unavailable: `, because "this attempt failed the check" and "nothing measured this attempt" are the degraded-vs-zero distinction §4a.1 exists to keep apart. `Exited` is a field of its own and is the ONLY gate on a verdict, so a check record's zero value can never read as a pass. - **No command named is ABSENT, and absence draws nothing at all.** Not a dash, not a 0, not a pending word. It is nil `Seed`'s rule on a second field: a room that was never asked for a check has no check to report. Absence is a whole-race property rather than a per-seat one — the command is the room's, so either every racer with a tree carries a check or none does. - **The room captures no output.** The exit code is the whole claim. A failing command's stdout would put unredacted subprocess text on a screen whose vendor streams all pass through a `Redactor` first, and reading it would need the scroll surface the gate cards were already refused. The block names the command and the worktree; the operator re-runs it there. #### Where it runs in the finish line, and what that costs **The check runs LAST, after the diff is read and after the attempt is committed, and the order is a ruling.** A check that ran first would park its own build output on the arena branch wearing the racer's name — the false receipt, §4a.1 pointed at a write. So nothing a check writes can reach the stat, the patch or the commit. What it CAN reach is a later `/adopt`, which commits a dirty attempt before merging it. That is said rather than swept up: the run is bracketed by two `git status --porcelain` reads, and a check that found a clean tree and left a dirty one prints one sentence saying so. The room does not reset a tree the operator did not ask it to reset, and `/adopt`'s own y/n card names every command it will run before it runs one. A tree that was ALREADY dirty, and a state that could not be read, both claim nothing — the field reports a measured change, never the absence of one. **It runs off the render loop**, for the reason `arenaSetup` moved off it (§9.37's 2026-08-17 amendment): a `go test` inside `Update` is a room that draws no frame and reads no key for minutes. `finishColumn` queues the run, `Update` drains the queue on the event batch and again on the spinner tick, and the result arrives as a message — `arenalive.go`'s pattern, with one difference that matters. **The check's lifetime is the ROOM's, not the turn's.** The last racer landing ends the turn, and a check that started at that moment must still be allowed to finish and report onto a column that is still on screen — so the run hangs off `roomCtx`, and teardown kills it for the same reason quitting kills every other child this room started. A stale run is dropped by comparison (the vendor and the turn number), never by hoping the timing worked out. **The run is a process TREE, and it is contained like one.** `runner.RunContained` is a new one-shot sibling of `runner.Start` — no streaming, no parsing, no clock record, because a check is not a turn — and what it takes from `Start` is the Windows job object and the unix process group. That is not belt and braces: `proc_windows.go` exists because `codex` resolves to an npm `.cmd` shim, and `npm test` has exactly that shape, so a deadline that killed only the direct child would leave the real work running two processes down with nothing on screen to say so. `contained_test.go` asserts it the way the vendor path is asserted — a helper that spawns a grandchild writing to a file, and a cancel that has to stop the FILE growing. Both of `Start`'s own limits carry over unchanged and are recorded there: the microsecond window before the group is assigned, and unix's need to be killed on the way out. **A cancelled turn runs no check.** ctrl+c is the operator saying stop, and a room that answered it by starting a subprocess per seat would be ignoring the one act it exists to obey. A seat cut on its own with `x` is NOT that case: §9.37's give-up ruling says a given-up seat lands like any other finisher, and its partial work is as worth checking as its diff is worth reading. **The bound is ten minutes, and the anchor is this repository's own suite.** `CLAUDE.md` records `go test ./...` at ~455s locally and ~4m22s in CI, which is about the slowest command an operator would name here; ten minutes is past twice that. The margin is the decision rather than the figure — what this bound must never do is kill a run that would have finished, since a killed run yields no exit code and therefore no verdict at all. A run the clock stops is reported as unavailable with the bound named. **Checks may overlap, and that is stated rather than serialised.** Seats land at different moments, so in practice they stagger; two that land together run together. The worktree adds are serial because they contend for the repository's own refs (§9.37), and nothing analogous is true here — each check runs in its own tree. Serialising them would make one slow check hold every other seat's verdict, and the room measures nothing about the machine's capacity, so the bound would be chosen from nothing. #### Deliberately not built - **Any LLM judge**, per the ruling above. It is the half of the sweep candidate that fails §9.2 and §9.44, and a mark does not soften it. - **Repeat sampling** (N attempts of one seat), the candidate's other half. §9.37 rules every attempt a fresh session across all seated vendors and rules a one-seat race an ordinary turn in a worktree; N-of-one contradicts that as written and collides with the `arena/t/` identity scheme `lifecycle.go` re-mints from refs. It needs an owner's ruling rather than a builder's, and nothing here assumes one. - **Per-attempt cost and diff-size columns.** They already ship — the stat is §9.37's, and the vendor-reported cost is `turnview.go`'s, absent where the vendor reports none. - **A check that gates anything.** It reports; it does not refuse an `/adopt` and it does not re-order a rank. The founding posture is offer-never-take, and a room that blocked an adoption on its own reading of a test would be taking. - **A check on an ordinary turn.** The question is whether an ATTEMPT works, and an ordinary turn writes into the operator's own tree rather than into an attempt. Verified offline. `arenacheck_test.go` pins the grammar (set, report, clear, and the refusal that hands a brief back), one run per racer in that racer's own tree with the named argv, absence drawing nothing, a cancelled turn queueing nothing while a given-up seat still runs, the stale-message drops, the ordering — a check that writes into the tree reaching neither the stat nor the commit — and the four renders, with `arena-check.txt` and its `--ascii` twin as the frame. `contained_test.go` pins the run itself: an exit code coming back as itself, a missing binary carrying no code at all, and a cancelled run stopping a GRANDCHILD. Colour is asserted separately against the room's existing severity pair, per the goldens' own split. No test here spawns a vendor: `countSpawns` stubs the check as its fourth spawn var, and `TestMain` panics on any model-path run whose command this machine could actually resolve. **One test runs a real process, and the process is this test binary.** `TestPassAndFailComeFromARealExitCode` calls `runCheck` directly rather than through the guarded var, with `os.Args[0]` and a helper `-test.run` as the command, and asserts that three exit codes come back as themselves, that a missing binary comes back with no verdict at all, and that a run which writes a file is measured as having written one. Without it the feature would rest on a stub returning the answer it was handed, which pins the render and nothing about the claim. **The live half is owed.** No live race has run under a check. What is unmeasured is a real vendor's attempt meeting a real suite in a real worktree, and specifically whether a check that runs for minutes reads well on a column whose turn has already ended. One `/arena check` and one `/arena` against a brief that changes files pays it. Recorded in `STATE.md` rather than implied paid. ### 9.49 the patch was on screen and there was no way to say anything about it (2026-08-29) `d` has flipped a racer's arena block between the diffstat and the whole patch since §9.37, and the block prints the branch and the worktree path under it. What the operator could do with any of that was read it. To say *this hunk is the problem* they retyped the hunk into the composer by hand, and to open the tree they selected an abbreviated path off a fixed-width column and repaired it in another terminal. The help panel, meanwhile, named none of it: `d` shipped and no page ever said the key existed. **Three keys close that, and the surface they make is the patch view itself.** A cursor points at one hunk of the drawn patch, `D` quotes that hunk into the composer draft, and `o` hands the operator the worktree — started in their own editor, or copied to the clipboard. #### The comment is a DRAFT, and that is the whole design The lane this was taken from (Vibe Kanban, Parallel Code, Sculptor) routes an inline comment straight back to the agent as its next instruction. **This room cannot do that honestly.** A race attempt is one-shot in a worktree: the seat that wrote the patch has finished, and §9.37 rules every attempt a fresh session, so there is no conversation for a comment to resume. A "reply to this attempt" control would therefore have to invent one — a new race, a new turn dressed as a continuation, or a queue. So the quote lands in the **live composer draft**, and nothing else happens. It is visible on the next frame, it is editable with every key the composer already has, and `enter` — a person, pressing a key — stays the only thing in this room that spends a quota. That is the standing rule about queued drafts (one visible, editable draft, never auto-send) reaching a feature that was invented to break it, and it is asserted as a spawn count and a turn counter rather than as prose: `TestDQuotesIntoTheDraftAndSpawnsNothing`. Two consequences are stated rather than hidden. The next brief runs in the **room's** workspace, not in the racer's worktree — so the fence names the worktree path, which is the difference between a comment a seat can act on and one it can only agree with. And an empty draft is seeded with the racer's own `@mention`, because silence routes to claude (§9.9) and a comment about codex's attempt reaching claude would be the room misdelivering the operator's words. The seed is one token, the footer's route cell resolves it on the very next frame, and a draft that already says something is left alone — it may carry a route the operator chose, and a second mention would be the room editing a line being written. #### The quote is fenced as DATA, because that is what §9.15 settled The fence is `quote.go`'s, for `quote.go`'s reason: this text goes from one program's output into another language model's input, which is a prompt-injection path whatever wrote it. The wording differs in exactly one way, and it matters — it names the material as **measured `git diff` output** rather than as another participant's answer, because that is what it is. A fence that mislabelled its contents would be the room narrating over its own evidence. The **whole hunk** crosses even when the drawn frame cut it off partway. The cursor points at a hunk and the hunk is the unit git itself framed; sending half of one because a render cap fell inside it would hand a seat an incomplete measurement under a fence claiming to carry git's output. What the frame decides is where the cursor may GO, not what a quote CONTAINS. A quote that would push the draft past the composer's own cap is **refused whole**, never truncated — `paste.go`'s atomicity rule reached by a different key, and for `paste.go`'s reason: the Antigravity seat takes its prompt on argv, so a draft past that size is one no seat could be handed. #### The cursor never scrolls, and that is a refusal rather than a limit The drawn patch is capped at `arenaDiffScreenLines` and does not scroll. A cursor free to walk past that cutoff would need a second scroll surface inside a column — the device this product already refused for the gate card's preview — so **the cursor points inside the frame the reader can see and refuses to leave it by name.** Both ends say which end they are, and the far end says *the cursor does not scroll* and names the two routes to the rest of the patch the cutoff line already carries (`y` copies the whole diff, the worktree holds it). The anchor is a **hunk**, not a rendered line, and the reason is that a rendered line has no name anything can resolve: "line 412" is a position in a capped render of a one-megabyte diff. A hunk header is git's own anchor, printed in the patch, and it survives the cap, the yank and the quote. The parser reads git's framing and no heuristic. Inside a hunk every line is prefixed by git (space, `+`, `-`, `\`), so a line that is not prefixed is not in the hunk — which is what makes it safe against its own input. A patch is whatever five language models wrote into five worktrees, and this repository's own diffs routinely ADD lines beginning `+++`, `---` and `diff --git`; a parser that scanned for those anywhere would cut a hunk in half at the first of them. #### The keys, and why none of them is new vocabulary - **`D`** is SHIFT on the key whose surface it acts on, which is the spelling `T` already earns beside `t` and `Y` beside `y`. It costs the help panel no row of its own for `T`'s reason: it is taught on the row that teaches `d`. - **`[` and `]`** step the cursor. These keys have always meant *step one unit of whatever the body is showing* — a turn in the grid, a page in the by-turn projection (§9.22) — and a patch is a body with a unit of its own. A third reading of one motion, not a fourth binding. The footer's cell says `[ ] hunk` while a patch is open, because §7.8's contract is that the mode line announces what every key means on every frame. - **`o`** raises the worktree card. A key rather than a room command for `c`'s reason: no vocabulary leaves the composer. The cursor arms with the patch and costs no key at all: `d` opens the patch and puts the cursor on the first hunk, so the patch view IS the review surface rather than a mode a reader has to discover before the feature exists for them. Its mark is the room's own `▸` — reuse, not a collision, because that glyph already means *this is what the keys address*, which is exactly what it means here. It is drawn only on the FOCUSED column, since that is the column `D` reaches. **Both keys refuse over a full-frame body** — the turn page, the act ledger, the arena record. `y` can follow a page because a page IS a document (§9.22, §9.47); these two cannot, because a page has no hunk and no worktree, and a projection whose whole point is that the turn is the unit deliberately has no per-seat focus for a narrower key to aim at. `D`'s claim is *the hunk under `▸`*, and a claim a reader cannot check against the screen is the one thing this room does not print. #### `o`: the one process council starts that no vendor asked for Spawning an editor is council's loud exception and it is taken deliberately. The gauges spawn nothing; council spawns vendor CLIs already, and this is a **stricter** trigger than any vendor spawn in the package has: two keys, in front of a card that names the program. - **$VISUAL, then $EDITOR, then nothing.** An unset pair renders as unset and the card offers the copy instead. Falling back to `notepad` or `vi` would be the room inventing a value the operator never gave it — §4a.1 on a setting rather than on a gauge. - **The whole variable is tried as a program first.** `C:\Program Files\…\code.cmd` is one program with a space in it, and splitting on whitespace before asking the operating system would turn the operator's own setting into a program named `C:\Program`. Only when the whole string does not resolve is it read as `program arg arg`. - **The card names the program and `y` starts THAT program**, resolved when the card armed rather than when the key was pressed — `adoptOnto`'s contract, applied to a spawn. - **No stdio is wired**, because the room owns the terminal. The honest consequence is that a terminal editor opens where the operator cannot see it, and the notice says so rather than letting them discover it. What is REPORTED is exactly what `Start` measured: the process began, or the operating system said why it could not. Whether the editor then drew a window is not observable from here, and the notice does not claim it — `yank`'s rule, one keystroke over. - **`c` copies the FULL path**, through the same native-helper-then-OSC-52 path `y` uses, and a cancel writes nothing at all: an empty OSC 52 write is how a clipboard is CLEARED, so "nothing happened" must not be spelled the same way as "your clipboard is now empty". `startEditor` is a package var for the reason the other three spawn vars are: `TestMain` makes it fail closed for the whole suite, and `countSpawns` stubs it with the vendors. A test that pressed `o` then `y` on a box where `$EDITOR` resolves would otherwise open a window on the desktop of whoever ran `go test`. #### What the help panel had to give up The panel's budget is hard (16 rows) and the arena keys had nowhere to come from, so the line-wise and screenful scroll rows were merged into one. That merge is the category rather than a saving — the bar this panel's own budget comment sets: they are one act at two scales, which is the argument `f / t / T`, `y / Y` and `g / G` are each already merged on, and the two rows had come to restate each other's "in compose too" clause besides. The row it freed pays for `d / D / o`, and `d` is documented for the first time. Verified offline. `review_test.go` pins the parser against a patch whose hunk body carries three lines that are file headers in every other context, a deletion's `/dev/null` never rendering as a filename, absent staying apart from hunk zero, the cursor stopping at the drawn frame with the reason named, `[` and `]` yielding the turn hop when no patch is open, the mark appearing only on the focused column and surviving `--ascii`, the footer cell following the body, `D` spawning nothing and starting no turn, the quote following the cursor rather than the column, the route seeded only into an empty draft, an over-large quote refused whole, four refusals each saying their own reason, both keys refusing over each of the three full-frame bodies, the card copying the path without spawning, `y` starting the program the card named in the racer's own tree, an absent `$EDITOR` said rather than guessed, a stray key touching neither clipboard nor process table, and the three keys named above the help panel's fold. `arena-review.txt` and its `--ascii` twin are the frame. No test here spawns a vendor or an editor. **Nothing in this section is a claim about vendor behaviour, so no live vendor run is owed.** What IS owed is one live open: `o` then `y` against a real `$EDITOR` on the reference box, and `D` on a real race's patch through to a dispatched turn. Stated here rather than implied paid. ### 9.50 the codex seat gets a second protocol, and the sandbox re-measurement that kept it unseated (2026-08-29) §9.33 drove `codex app-server` on 2026-08-15, recorded a warm turn at **1.44 s** against `codex exec --json`'s 5.33 s, and stopped: *"This is measurement only. It authorises no seat change."* `STATE.md` carried the standing caution beside it — the Windows sandbox finding does not port to that path, and **any seat move re-measures rather than inherits**. This is that re-measurement, and the build it authorised. **The verdict is split, and both halves are measured.** The protocol ships: a second `runner.Protocol` beside the ACP one, parsed, tested, and pinned to a real capture. The SEAT does not move: the room still dispatches codex through `codex exec --json`. What decided that is not caution — it is a liveness defect measured on the new path that #311 had just cleared off the old one. #### Version pinned first, and the surface was driven before it was believed Everything below is **codex-cli 0.149.1** on Windows 11 (`codex --version`), 2026-08-29 — the same build #311 re-measured the `exec` seat on the same day, so nothing here is confounded by a version difference between the two surfaces. `codex app-server generate-json-schema --out ` wrote the whole protocol at this build; every shape below was then either captured live or is labelled a schema read. **Eight billed turns**, one to three per arm, in throwaway directories, with files checked on disk rather than read out of the model's reply. Prompts are **brief-shaped**: they open with `brief.go`'s own `--- operating context …---` fence, because §9.39 records a seat that shipped broken for a day behind a green live test whose prompt began with a letter. #### The handshake, and what a warm thread actually costs | stamped line | arm A | arm F | |---|---|---| | spawn returned | 0.035 s | 0.151 s | | `initialize` response | 1.091 s | 0.673 s | | `thread/start` response | 1.986 s | 1.032 s | Two round trips, and the spread is this box rather than the vendor: five MCP servers start on the same path and a `sessionStart` hook runs, both of them the operator's own configuration. §9.33's warning stands unchanged — **that figure must never be quoted as a property of the vendor.** Then the two turns of arm A, one thread, one process, read-only posture: | turn | sent | `turn/completed` | total | |---|---|---|---| | 1 (three shell attempts, hooks, MCP startup) | 1.986 s | 34.771 s | 32.785 s | | 2 (a question only turn 1's history answers) | 34.771 s | 36.257 s | **1.486 s** | Turn 2 was asked what file turn 1 had been told to create and answered `wrote-ro.txt`, from the same pid. **1.486 s reproduces §9.33's 1.44 s at a newer build**, so the prize is real and it is now measured twice, fourteen days apart, on two builds. **The linger question re-opens and closes in the vendor's favour.** `codex exec` prints its last line and holds the process open ~4 s (§9.33's 2026-08-16 amendment). Here there is nothing to settle around: turn 2's `turn/start` went out on the same millisecond `turn/completed` landed. The shutdown fact runs the other way and is the one worth carrying: **closing stdin does not reliably stop this server.** Four runs exited in 1.5–3.3 s; one was still alive 15 s later and had to be killed. The caller owns the kill. #### The sandbox, re-measured on its own surface `thread/start` takes a `sandbox` parameter whose enum is spelled identically to `exec`'s `-s` values. That spelling is a coincidence of naming, not a shared mechanism, and the results are not the same. | probe | result | |---|---| | `read-only`, write via **direct `cmd.exe`** | **DENIED** — `Access is denied.`, exit 1, `status:"failed"`, no file on disk | | `read-only`, write via the router's default `pwsh.exe` | **NO PROCESS** — `CreateProcessAsUserW failed: 5`, no `commandExecution` item at all | | `workspace-write`, write inside the workspace | landed, exit 0, file on disk | | `workspace-write`, write to `.git` (two independent turns) | **DENIED** both times, exit 1, no file | | `windowsSandbox/readiness` (a free request, no turn) | `{"status":"ready"}` | **The read posture enforces when a process starts.** The denied write is the strong arm, and its own failure mode confirms it: the second call in that turn came back with cmd.exe's *own* `'…' is not recognized as an internal or external command`, which a process that never started could not have produced. So cmd.exe ran INSIDE the sandbox and obeyed it. **And the seat still cannot be trusted to read, which is why it is unseated.** This protocol's tool router wraps a shell command in `"…\WindowsApps\pwsh.exe" -NoProfile -Command "…"` unless the model names a program itself, and pwsh cannot start under the Windows sandbox on this box — the identical `CreateProcessAsUserW failed: 5` that 0.146.0 failed *everything* with. #311 measured `codex exec` retrying through cmd.exe and obeying the sandbox. On this path the retry is the model's own move and **it did not always make it**: in two of three read-posture arms the model abandoned the turn and reported it could not inspect. A read seat that cannot list a directory is the 0.146.0-class defect arriving on a new path, and it is the exact failure the `ro:enforced` badge would be sold against. **Two things are NOT measured, and they are stated rather than inferred from the `exec` path's answers — the whole reason this section exists is that those answers do not port:** - **Writing outside the workspace.** The one arm that tried it wrote into a directory under `%TEMP%`, which `workspaceWrite` permits by default (`excludeTmpdirEnvVar` defaults false). The write landed and the arm proves **nothing**. Recorded as void, never as a permit. The probe could not be re-aimed: this lane's scratch boundary is inside the temp root. - **The `writableRoots` override on `.git`.** A per-turn `sandboxPolicy` naming `.git` was sent and the model's own shell quoting broke the call before the sandbox saw it. So whether this path can buy `.git` back — which `codex exec` measurably cannot on Windows (#311) — is open, and it is the measurement the write posture's shape depends on. #### What this protocol carries that `codex exec --json` hides §9.33 named these and built nothing on them. They are now parsed, and still rendered nowhere. - **`thread/tokenUsage/updated`** — `total` and `last` counts, plus `modelContextWindow` (258400 on this account's model). That denominator is the one §3.2 records as absent from every other codex surface; it is what would make a context percentage a read rather than an invention. - **`account/rateLimits/updated`** — `usedPercent`, `windowDurationMins`, `resetsAt` and `planType`, per window, live on the socket. These are **quota** in §7.15's vocabulary — a share of a window that resets — and never spend, and nothing here converts one into the other. - **`hook/started` / `hook/completed`** — the hook's id, event, source path, source and `durationMs`. Dropped by the adapter, and deliberately: the one hook captured was the operator's own `sessionStart` script, and rendering somebody's local configuration as council activity is the machine-specific claim §9.33 already warned this figure must never become. - **Typed shell items** — `commandExecution` carries `command`, `cwd`, `processId`, `status`, `exitCode`, `durationMs` and `aggregatedOutput`. Richer than the `exec` stream's item, and `status:"failed"` is load-bearing here in a way it is not there: it was captured on a command that never started, where there is no exit code to read. **A free read surface, which is new and worth naming.** `account/rateLimits/read`, `hooks/list`, `permissionProfile/list` and `windowsSandbox/readiness` are ordinary requests that answer **without starting a turn** — driven live, no model spend. Nothing consumes them yet. #### The trap this wire carries, measured rather than assumed away **The deltas and the completed item are the same text.** Across four `agentMessage` items over two turns, the concatenated `item/agentMessage/delta` payloads equal `item/completed`'s `text` byte for byte, every time. This is §9.6c's whole-message repeat, present on a surface nobody had parsed. A reader that took both would print every answer twice. The adapter consumes the deltas — streaming is what the column is for — and spends the completed item as a **separator only**, per item rather than per turn, because one turn carries several complete messages and a per-turn guard would pass the first and fail the second. An item that never streamed still prints its text, which is the safety net the ACP seat explicitly does not have. #### The fork, taken conservatively §9.36 ruled the Cursor seat's equivalent fork **WHOLESALE** under one-attempt probation, and the numbers there justified it: the old path was ~13 s per turn against 1.1–1.8 s, with nothing left to fall back to that anyone would want. **This fork is not that one**, and the difference is on the safety side rather than the speed side. A cutover here would move the codex seat onto a path where, at this build, a read-posture turn was measured **failing to inspect at all** in two of three arms. The old path does not have that defect — #311 measured it away the same day and paid for the badge. Trading a 5.33 s turn for a 1.49 s turn that sometimes cannot run a command is not the trade §9.36 made; it is the inverse of it. And the two measurements the write posture would need — the outside-workspace boundary and the `.git` override — are the two this lane could not close. So: **dual, with the default seat unmoved.** The protocol, its parser, its refusals and a version-pinned real capture all ship. `vendors.CodexAppServer` implements `Conversational` and is deliberately absent from `Registry()`, with `TestTheRoomStillSeatsTheExecCodexAdapter` pinning that absence as a decision rather than an oversight — it is the test that fails, and is read, on the day somebody registers the seat. **What the flip needs, named so the follow-up does not rediscover it.** One line maps `model.VendorCodex` to this type; the seat then drives through `StartRPCSession` exactly as the Cursor seat does. What must be re-measured FIRST is the read posture's **liveness** — a turn that lists a directory and reads a file, under the sandbox the badge claims — because that is the property the badge sells and the one this path was measured failing. The write posture needs the two open sandbox arms above. Until then the badge rule is unchanged, and `unsandboxed` stays the floor wherever restriction is unverifiable. #### Verification Fixture replay through the real protocol driver, with no process anywhere near a test. `codexappserver_test.go` pins the handshake sequencing and the held brief it releases, the posture riding `thread/start` rather than argv, the write posture NOT inheriting the `exec` seat's Windows flag, the whole-message repeat read once and two messages not running together, a message that never streamed still landing, command outcomes on both branches with the vendor's own failure line, a `completed` status with no exit code staying Unknown, the brief and the model's thinking both dropped, the turn ending exactly once with no cost, an unseen stop status not rendered as an answer, a refused handshake being terminal, a refused resume costing two round trips rather than the turn, the interrupt naming both required ids, every server request being answered and an unrequested approval refused into the trace, `Decide` never claiming to have answered anything, usage and limits captured while producing no events, and garbage on the stream not losing the turn. `wire_test.go` replays the real capture, `vendors/testdata/wire/codex-app-server-0.149.1-turn.jsonl` — the sandbox arm itself, sanitized, so a denial arriving as a success or the vendor's `Access is denied.` going missing fails there. `go vet ./...` clean, full `go test ./... -count=1` green. **Spend:** eight billed turns, all cheap prompts in throwaway directories. One of them bought nothing: the `writableRoots` arm broke on the model's own quoting and is recorded above as owed rather than as an answer. **Not verified here: macOS.** Every arm ran on Windows 11. Whether `codex app-server`'s sandbox behaves differently there is unmeasured, and `PARITY.md` is where that belongs. ### 9.51 the columns were panes, and the operator could not size one (2026-08-31) The room drew a row of columns and gave the operator no control over how wide any of them was. Width came from three places, and none of them was a key. `resolveLayoutIn` divided the usable cells evenly. `FrameOwners` narrowed the frame from the ROUTE, which the operator sets by addressing a turn and not by pressing anything. `Expanded` — the `f` key — gave the focused column the whole frame and was a boolean with no middle position. A reader who wanted one seat wide and the other seats still on screen had no way to ask for it. This section names what the grid already is, then adds the one control it never had. #### The vocabulary **A pane is one drawn seat column, and the room is a single row of panes.** The word is new; the thing is not. `VisibleColumns` decides which seats hold a pane, `Layout.widthAt` gives each pane its width, and `columnsBody` paints the row. Nothing about the roster changes here. The operator now owns two facts about that row: - **which pane owns the reading width** — the SPLIT. - **where a boundary between two panes sits** — the SIZE. Everything else was already built, and this section restates it rather than rebuilds it. Focus moves between panes with `tab`, `shift+tab`, `h`, `l` and the seat numbers (§9.12, §9.29). `f` zooms the focused pane to the whole frame (§9.11). A second key for either act would give the room two spellings of one thing, which is the defect §9.31 records under its own name. #### The keys, and why they sit behind a prefix `^w` arms the pane prefix. The next key is a pane key: | Key | Act | |---|---| | `^w` `s` | split: the focused pane owns the reading width, and the other panes hold at `stripColumn` | | `^w` `>` | grow: move the focused pane's boundary right by one step | | `^w` `<` | shrink: move the focused pane's boundary left by one step | | `^w` `e` | even: every pane gets the same width again, and the split clears | | `^w` `c` | compare (added 2026-09-16): the focused pane joins the split's owner, and the two share the reading width. With no split in force it is a split. On a third seat it replaces the peer. `docs/room-identity.md` carries the rule | | any other key | cancels the prefix, and is swallowed | **An unrecognised key is swallowed and NOT re-dispatched.** A prefix that let the second key fall through would read as tolerant, and it is dangerous in this keymap: `^w` then `q` would quit the room, and `^w` then `ctrl+c` would cancel a turn. Those are two irreversible acts reached by a chord the operator has already shown they did not finish. The footer says `any cancel` for the same reason. A footer that named `esc` alone would imply that the other keys still mean what they mean, and one of them is `q`. A prefix, and not four more chords, because the keymap has almost no surface left. §9.49 states the pressure in its own words: none of the three keys it added was new vocabulary, and each one had to be taught on a help row that already existed. `[` and `]` there take a third meaning rather than a fourth binding, and `h`, `j`, `k` and `l` are already focus and scroll. Four top-level bindings would spend the last of that surface on one feature. One prefix spends one binding and leaves the rest open. **The prefix is view mode only.** In compose mode every printable character is draft text, which is the contract `q`, `f`, `c` and `t` already keep. A prefix armed there would change what the next letter does while the operator types a brief, and the operator would learn it by losing a character. **An armed prefix says so.** The composer box's bottom border reads `PANES` while the room waits for the second key, at the rank `GATE` and `COMPOSE` already take (§9.44), and the footer names the pane keys (five since `c` joined on 2026-09-16). §7.8 forbids a mode that changes what an unmodified key means without saying so. This is such a mode, for exactly one keystroke. #### The arithmetic, and the two invariants it may not break The split reuses `weightedWidths`. `framePrimary` already marks which seats own the frame, and `State.PaneOwner` joins `State.FrameOwners` as a second source for that mark. **The operator outranks the route.** A split the operator asked for is a request. `FrameOwners` is an inference from where a turn went. When the two disagree the request wins, and the next dispatch does not clear it. The split is pinned to a SEAT, by vendor, and not to whichever pane holds focus. A split that followed focus would reflow the whole grid on every `tab` press. `tab` moves a marker today, and §7.1 rule 4 does not budget for a keystroke that re-wraps two columns of prose. The size is a per-seat bias in cells, `State.PaneGrow`, applied over whatever the base apportionment is. One press moves one boundary: the focused pane gains a step and the pane to its right loses the same step. The bias therefore sums to zero, and the row still fills the terminal exactly. The rightmost pane takes its step from the pane to its left, because it has no right neighbour to take it from. Two invariants hold over every frame. `TestColumnsExactlyFillTheWidth` and `TestPaneWidthsHoldTheirFloors` assert them: 1. **The panes plus their separators fill the terminal exactly.** A short row leaves a ragged edge. A long row wraps, and the grid shears. 2. **No pane goes below `stripColumn`.** 18 cells is the width at which a column stops being a seat and renders as a strip, and §9.18 measured that a strip below it cannot say the two things a strip exists to say. `normalizeBias` repairs a bias that no longer sums to zero. That occurs when a seat folds out of the grid between the keystroke and the frame. The repair is deterministic, and it runs inside `resolveLayoutIn`. A `State` that a test types out by hand therefore cannot produce a torn frame. #### The step is the separator's own width `paneStep` is `1 + 2*gutter`, which is five cells: one rail and its two gutters. The value is derived and not tuned. One press moves a boundary by exactly the gap the reader sees between two panes, so the move is visible on the first press. A one-cell step re-wraps nothing on most lines, and it reads as a key that did not work. #### The floors, and the ladder below them `minColumn` keeps its old job and gets no new one. It is the width below which the whole tier drops to tabs, and `tierFor` still tests it against the EVEN width, before any operator bias. **The operator therefore cannot size the room out of the columns tier.** A terminal too narrow for a grid is too narrow before a pane key is pressed, and it stays that way after. Below the columns tier the pane controls do nothing, and the room refuses to offer them: `^w` does not arm at the tabs tier, over a turn page, over an arena record, or in a zoomed frame. That refusal is the reason the footer needs no permanent pane cell at all — the pane keys are named only while they are live, so there is never a frame that promises a key which does nothing. It is §9.11's rule reached by a different route: that section drops `f` and `tab` outright in a one-seat room, and this one never offers the keys in the first place. The composer border is also silent below the columns tier, even when the split and the bias are still stored. A legend describing a boundary the reader is not looking at would be the room describing someone else's frame. The stored split and the stored bias survive the narrow frame and return when the operator widens the terminal, exactly as `Expanded` does. A control that forgot on a resize would punish the operator for dragging a window. The full ladder, widest first: even panes, then a split pane with strips beside it, then one zoomed pane with a tab bar, then the tab tier, then `floorMessage`. Every rung was already built except the second one, and the second one is this section. #### Colour carries none of it WORDS on the composer box's bottom border carry the pane state: `^w e panes split`, `^w e panes sized`, or `^w e panes split, sized`. The key comes first and the state follows, which is the shape `a not asking` already has (§9.24). The border names the state and the key that reverses it, so the operator does not have to remember which press undoes a split. `^w` is ASCII, the words are ASCII, and the feature spends no glyph and no hue. The legend appears only while the operator has split or sized something, so the ordinary room's frame is byte for byte what it was. #### What this does NOT build - **A second axis.** Panes split left to right, and never top to bottom. A seat's column is one transcript, and a horizontal boundary through it would cut one document into two viewports. Nothing measured says a reader wants either half on its own. - **A pane that holds something other than a seat.** A pane is a seat's column. A pane that held a file, a diff or a shell is a different product, and the arena review surface (§9.49) already answers the one case that came up. - **A key that adds or removes a pane.** `/seat` and `/unseat` do that (§9.31), and they do it in the roster, where the effect on DISPATCH is visible. A `^w` key that hid a seat while the seat kept answering would put a live vendor off screen with nothing anywhere to say so, which is the failure `CollapsedColumns` and the notice row exist to prevent. - **Mouse drag on a boundary.** §9.10 refused the mouse wheel, and its reasoning transfers whole. The enum half of that section is MEASURED: there is no wheel-only mouse mode, so a program cannot ask for a drag without claiming button reporting. The cost half is stated there as INFERRED rather than measured — that button reporting suppresses the terminal's own text selection — and it is repeated here at that same strength. Nothing new was measured for this section, and a boundary drag would buy an input convenience with the room's output, which is the trade §9.10 already refused on this surface. #### Verification `layout_test.go` sweeps every width from `columnsBreak` to 220 and every pane count from 2 to 4, over the biases the keys can produce, and it asserts the two invariants above. It sweeps a bias far larger than any keystroke writes, on purpose: `Render` is pure over `State`, so `State` is an input this package does not control, and an invariant that held only because `paneResize` was careful is one a hand-typed test could break by accident. `panes_test.go` drives the keys through `Model.key`. It asserts that the prefix arms, that every branch clears it, that an unknown key is swallowed rather than quitting the room, that no pane key reaches the draft in compose mode, that the split does not follow focus, that one press moves one boundary, that the boundary stops at the floor rather than appearing to move, and that the arrangement is legible with `PlainStyles` and the ASCII glyph set — which is the test that would catch a pane feature readable only in colour. Four goldens are new: `panes-split.txt`, `panes-split-ascii.txt`, `panes-sized.txt` and `panes-keys.txt`. One golden changed, `help.txt`, by one row, because the panel now names the prefix. ### 9.52 the room ended every agent on the way out and never said so, and the room that came back implied they had lived (2026-09-01) Two defects, one sentence apart, and they are the same defect seen from each end of a quit. **On the way out, the room says nothing.** `q` and `ctrl+c` reach `teardown`, which kills every seat process, and that is correct: an agent that outlives the window that shows it is the invisible state this product refuses. Nothing on screen says it happened. The operator learns the contract by noticing, later, that a conversation is cold. **On the way back in, the room says the wrong thing.** The reattach notice reports `2/3 seats restored`, and the seat card reports `this seat's thread came back`. Both sentences are true about the THREAD. Neither says one word about the PROCESS, and the process is the half that died. A room that opens on four columns of restored threads reads as a room that was left running. It was not. This section rules both halves and it introduces one word. #### The word is `rebuild`, and `reattach` and `rejoin` are not touched `reattach` is taken. It means *resume the vendor session ids from `room.json`*, in `Reattachment` (`resume.go`), in the goldens (`testdata/golden/reattached.txt`) and in the demo script (`STATE.md`). It keeps that meaning exactly. `rejoin` is reserved and deliberately unspent. A later host lane needs a verb for *a client reconnects to a live process*, and spending it here on something that is not that would leave that lane renaming a shipped word. So the new verb is **rebuild**, and the three words name three different facts: | word | what it is a fact about | when it is true | |---|---|---| | **reattach** | the FILE | the room read `room.json` and holds the saved ids | | **rebuild** | the PROCESS | the room launched a NEW vendor process on a saved id | | **rejoin** | reserved | *(a client reaches a process that was already running — nothing does this today)* | The room performs the first two. It has never performed the third, and until something does, no surface may use the word. #### Rung 0 — the quit path states the contract it already keeps Nothing changes about what quitting does. What changes is that quitting says it. The room prints a closing line on stdout after the alternate screen is released. It is stdout rather than a card because there is no longer a frame to draw a card in — the same reasoning that already puts a failed save on stderr at that point. The line reports three measured facts and infers none of them: 1. **How many vendor processes were ended.** Counted in `teardown`, from the seat registry it is already ranging over. A room that spawned nothing reports a measured zero in its own words — `no vendor process was running` — rather than `0 vendor processes ended`, because the two sentences answer different questions and only the first one is true here. 2. **What survived on disk, and what did not.** The session ids and the turn number survived, at the named path. The conversation did not. Saying only the first would let `room.json` be read as a transcript, which it has never been (`resume.go`'s doc comment). 3. **What reopening costs.** Named as a rebuild, in rung 2's vocabulary, so the two ends of one quit use one word. The line is not a warning and carries no mark. Ending the seats is the room working. #### Rung 2 — the rebuild happens at room open, and it says which of the two things it is Today the saved session ids are spent on the FIRST DISPATCH: `seatProcess` launches the seat's process with the saved id when the operator's first brief arrives. Everything that startup costs is therefore charged to the first brief, and the operator waits for it while looking at a room that appeared instantly. Rung 2 moves the launch to room open. The room rebuilds every restorable seat as soon as it opens, in parallel, and reports each one's progress in its own column. **What that buys is latency, and it is worth stating exactly which latency.** `runner/session.go` records the measurement this rests on: a one-word turn cost about 25 seconds and about $0.23, *nearly all of it startup*. Moving the launch earlier moves the seconds and does not move the dollars — a process that has started has run no model turn, so nothing is billed until the first brief. So the claim is split, and the room states both halves: - **The ~25 seconds are spent at room open instead of on the first brief.** That is what the operator gets. - **The ~$0.23 a seat is still billed by the first brief.** Rung 2 does not spend it early and must not appear to. Both figures carry a leading `~` and both name what was measured: one one-word turn, once. Four seats is an extrapolation from one measurement, and the room says so rather than printing a four-seat total as though somebody had counted it. **Per-seat progress is measured at every step, and there are four outcomes.** They are carried on the column's existing note, so this rung adds no field and no render path: | state | what was measured | what the seat says | |---|---|---| | **rebuilding** | the spawn returned with no error | a new process is loading the saved thread | | **rebuilt** | the vendor announced a session id, and it is the saved one | the thread came back, on a NEW process — and the one you left was ended | | **forked** | the vendor announced a DIFFERENT session id | §9.43's existing sentence, unchanged | | **failed** | the spawn failed, or the process exited before it announced anything | the vendor's own line, and the next brief opens a new session | #### Where each half of the news lives, and why it is not the notice The first build put the whole statement in the room notice, and the notice **is one line and it is truncated, not wrapped**. At 120 columns the reattach sentence already fills most of it, so the settled rebuild rendered as `… 2/4 seats rebuilt in 0s — NE…`. A cost clause that disappears at a hundred columns is not a stated cost, and the clause being cut was the exact one the rung exists to say. The columns are the opposite shape: every note wraps, every seat has one, and `noteCard` already draws a muted detail block under its title. So the split follows the shape of each surface, and it lands where `reattachCard`'s own rule already points — the room fact in the notice once, the seat fact in the seat. - **The COLUMN carries the sentence that must never be lost**, and the measured cost under it as detail. The cost is a *per-seat* fact (`$0.23 a seat`), so per-seat is the correct home and not merely the roomy one. - **The NOTICE carries the room fact.** `rebuilding 2 seats`, joined to the reattach sentence while the rebuild runs; then `2/4 seats rebuilt in 24s — NEW processes, not the ones you left` once it settles. **The settled sentence replaces the reattach sentence rather than joining it**, and that is the one thing this rung spends. Joined, it is cut. The reattach sentence is not lost by the swap: it was the entire notice from room open until the moment the rebuild settled, so its once-only clauses have been on screen for the whole window — and `telltale council ls` (§7.27) can print them again at any time. **`rebuilding` and `rebuilt` are two states and they must not collapse.** A launched process is not a proven thread. `persistent.go` already refuses to claim otherwise on this exact path — "Deliberately NOT reattached to the saved thread. Nothing has come back yet" — and the rebuild keeps that rule. What promotes a seat from `rebuilding` to `rebuilt` is the vendor's own init line arriving with a session id, which is a statement from the vendor and not a timer. **The fork case is not new machinery.** A vendor that is asked to resume and answers in a fresh conversation is §9.43's finding, and `adoptSession` already detects it through `forkWatch` and already prints the honest correction. The rebuild arms `forkWatch` with the saved id exactly as a dispatch does, so a fork at room open and a fork at turn 5 report identically. One mechanism, one sentence. **The seat card and the note say different halves, on purpose.** The existing reattach card says `this seat's thread came back`, which is a claim about the thread and stays true. The rebuild note under it says the process is new. Together they state what came back and what did not. Neither alone would. #### What rung 2 deliberately does not do **It persists nothing new.** No transcript, no scrollback, no vendor output. `resume.go` ruled this for the same data, and the ruling does not change because a process moved: duplicating any of it would be a second copy of a private conversation in a place the user did not choose. `room.json` stays session ids, a workspace, and numbers. **It starts no process the room would not have started anyway.** Every seat it launches is a seat the first brief was going to launch. The rebuild changes WHEN, never WHETHER. A room that opened and spawned a seat the operator had not seated would be spending on a roster nobody typed. **It launches nothing that is not restorable.** A seat with no saved id, a vendor this machine cannot run, and a seat that is not driven as a live process are all skipped. Each is skipped for a measured reason, and none of the three is reported as a failure. **It does not survive anything.** This is the sentence the whole rung is built around: the agents did not live through the quit. They were ended, and new ones were started on the ids they left behind. A room that rendered a rebuild as a continuation would be the most expensive lie this surface could tell, because the operator would trust a history that no process holds. #### Verification The rebuild is exercised with the spawn vars stubbed (`countSpawns`), which is the package's standing rule: a council test never starts a vendor. The kickoff is fired from `Init`, which no test calls, so a model a test builds directly launches nothing at all. Goldens pin the two rendered states apart — `rebuilding.txt` and `rebuilt.txt` — because `rebuilt` and `survived` rendering alike is the regression this section exists to prevent. Both take their strings from the model itself rather than from text typed into the test, so a wording change moves the golden instead of quietly passing a stale assertion. **No footer hint was added**, so no existing golden moved: every golden's last line is the footer's key hints, and one new hint there rewrites about eighty-nine files at once. ### 9.53 a live seat shows Claude Code's own screen, and measures nothing from it (2026-09-01) Every seat in this room is the same kind of thing: a column whose body is text council parsed out of a structured stream. That is what makes the gauges honest, and it is also the whole limit. The operator cannot see what the agent's own terminal shows. The permission prompt it draws, the spinner it spins, the box it puts around a diff — none of that is in `--output-format stream-json`, so none of it is in the room. This section adds a second KIND of seat. A LIVE seat runs the vendor in its own interactive mode, under a pseudoconsole, and draws that program's real screen inside the pane the seat already owns. Everything else about the seat does not move. #### The display-only contract **The live screen contributes NO measured field.** Not a cost, not a token count, not a context percentage, not a quota window, not a posture, not a phase, not a turn clock. Every gauge, every badge and every number on a live seat still comes from the structured adapter path, or renders absent under [§4a.1](#s4a-1)'s rule. This is not caution. It is the honest-gauge rule ([ADR-001](#adr-001)) applied to the one input that would break it. A screen of ANSI is a picture of a program, and reading a number off a picture is inference. `claude.exe`'s own status row prints a dollar cost and a weekly-limit sentence — both were captured in the spike, both are exactly the figures this room renders elsewhere, and both are exactly the figures the room must NOT take from here. `CapNone` exists to refuse inference. A field scraped off a repaint is a field this product spent its whole life declining to invent. So the seam is drawn at the type. The emulator lives on `Model`. The only thing that crosses onto `State` is `State.Live.Grid`, a slice of already-decoded plain rows, and `Render` draws those rows and nothing else. `applyPTY` writes `State.Live` and no other field of `State`. `TestLiveSeatMeasuresNothing` asserts that as a property rather than as a review note: it feeds the emulator a script carrying a cost, a quota sentence, a posture word and an elapsed time, then compares the whole `State` before and after and demands the only difference is `Live`. **The live pane is a SECOND process, and it doubles that seat's spend.** The structured session keeps running; the live child is another `claude` on the same account. The room says so where the pane is, in a word, because a surface that quietly billed twice would be the most expensive thing this section could ship. #### Only Claude Code can take the seat `vendors.Persistent` has exactly one implementer (`vendors/claude.go`). Codex, Antigravity and Grok are batch programs that exit every turn, so a pseudoconsole on one of them is a pane that is empty between turns. Cursor speaks ACP JSON-RPC and draws no terminal UI at all; seating it live would mean a different invocation that loses every measured field the ACP path supplies, which is the display path bought with the structured one — the exact trade this section refuses. The live invocation is not council's invocation. Council runs `claude --input-format stream-json --output-format stream-json ...`. The live child runs `claude` with no arguments, which is a different program mode with a different contract, and that is why it is a separate child rather than a second reader of the first. #### `os/exec` cannot spawn a ConPTY child Go's `syscall.SysProcAttr` on Windows has no field for a `PROC_THREAD_ATTRIBUTE_LIST`, so `EXTENDED_STARTUPINFO_PRESENT` cannot be honoured through `exec.Command`, and `PROC_THREAD_ATTRIBUTE_PSEUDOCONSOLE` is the whole mechanism by which a child attaches to a pseudoconsole. The spawn half is therefore a direct `windows.CreateProcess`, and it is the only place in this repository that starts a process without `os/exec`. The containment half is NOT rewritten. `windowsGroup`'s job object was measured working unchanged on a ConPTY child: `AssignProcessToJobObject` succeeded, and closing the job handle killed the child immediately where `ClosePseudoConsole` alone had left a `claude.exe` REPL alive more than three seconds. `attach` is refactored to take a pid, `attach(*exec.Cmd)` calls it, and the PTY calls the same function. One job-object implementation, two spawn shapes. `golang.org/x/sys/windows` v0.47.0 already exposes `CreatePseudoConsole`, `ResizePseudoConsole`, `ClosePseudoConsole`, `PROC_THREAD_ATTRIBUTE_PSEUDOCONSOLE` and `NewProcThreadAttributeList`, so the pseudoconsole itself costs no new module. One non-obvious requirement, measured: the child does NOT attach to the pseudoconsole unless `STARTF_USESTDHANDLES` is set with all three std handles left at zero. Without it the child keeps the PARENT's std handles, even with `bInheritHandles` false and even when the parent owns no console at all — a `cmd /c echo` printed on the parent's own stdout and the pty stream stayed empty. That one flag is the difference. #### `CREATE_NO_WINDOW` is the trap, and it fails silently `proc_windows.go` sets `CREATE_NEW_PROCESS_GROUP | CREATE_NO_WINDOW` and `HideWindow` on every child today, and the comment there is correct for every child it was written for. It is fatal here. Measured on 2026-08-31, a ConPTY child created with `CREATE_NO_WINDOW`: - emits ZERO bytes — not even conhost's own preamble, - accepts no input, so a write to its stdin does nothing, - exits 0, and `CreateProcess` returns a NIL error. There is no log line, no error and no output. The pane is simply blank. `DETACHED_PROCESS` fails the same way and additionally exits 1. `CREATE_NEW_PROCESS_GROUP` and `HideWindow` are both compatible and both measured working. An implementer reaches this bug by copying the existing spawn helper, which is the most likely thing to do, so the refusal is written down at the call site and the flag is named in `pty_windows.go`'s doc comment rather than left to this file. **There is no console flash to suppress.** That was measured with a differential `EnumWindows` scan carrying a positive control, across every flag combination and on runs up to twenty seconds: zero new visible windows. The reason is structural — a child holding `PROC_THREAD_ATTRIBUTE_PSEUDOCONSOLE` attaches to the pseudoconsole's own headless conhost and never asks Windows for a console, so `CREATE_NO_WINDOW` has nothing to suppress and its only effect here is to break the attach. One caveat for whoever writes the regression test: ConPTY's conhost DOES create a window of class `PseudoConsoleWindow`, `IsWindowVisible` reports it as visible, and it never paints. A flash test must match `ConsoleWindowClass` and must not treat `PseudoConsoleWindow` as a flash. #### `fit` is not sufficient, and no width helper can be `fit` (§9.5) is ANSI-aware for SGR, which is the `padRight` trap it was written for. It does not solve this problem and cannot. `lipgloss.Width` measures display cells, and a cursor-move, an erase-in-line, an absolute cursor position and a scroll-region set all measure as ZERO cells — so `MaxWidth` passes them through untouched into telltale's frame, where they execute against the host screen. These were all found in the real captures, and each is a different kind of damage: | Sequence | Damage if forwarded | |---|---| | `ESC[8;20;60t` | resizes the operator's REAL terminal window | | `ESC[>0q` | the real terminal answers into telltale's own stdin — a reply to a question the host never asked | | `ESC[?9001h` | switches the host into win32-input-mode, and Bubble Tea then decodes keys wrong | | `ESC[?1004h` | enables focus reporting on the host, injecting focus events into its input | | `ESC[?2004h`, `ESC[?2031h` | change how the HOST receives pastes and theme changes | | `ESC[2J`, `ESC[13;1H` | absolute addressing against the pane's origin, painting over telltale's frame | | `ESC]0;...BEL` | sets the host window title to the guest's path | The escapes must be CONSUMED, not measured. That is an emulator, and it is mandatory. #### Why `x/vt`, and not a hand-rolled parser `github.com/charmbracelet/x/vt` at `v0.0.0-20260830003929-9f48cc723c1c` (MIT) is taken as a direct dependency. Three reasons, in the order that decided it. **A hand-rolled parser is not a smaller problem, it is the same problem with a smaller test suite.** The table above is not the list of sequences to handle; it is the list found in twenty seconds of two guests. Correctness here is not "parse the ones we saw" — a sequence this room does not model and forwards by accident is a host-terminal corruption the goldens cannot see, because `PlainStyles` neutralises telltale's own escapes and is blind to a vendor's. The failure mode of an incomplete parser is invisible to every test this repo has. **The dependency cost is exactly two modules.** Measured by diffing `go list -m all`: `x/vt` and `github.com/charmbracelet/x/exp/ordered` (v0.1.0, MIT, a small generics helper). Everything else it needs — `ultraviolet`, `x/ansi`, `x/term`, `colorprofile`, `displaywidth`, `uax29`, `go-colorful`, `go-runewidth`, `cancelreader`, `uniseg`, `terminfo`, `x/sys`, `x/sync` — is already in the graph because Bubble Tea v2 pulls it. `golang.org/x/exp` is NOT pulled in. **It links against this repo's exact pins, proven rather than assumed.** A compatibility program built against `bubbletea/v2 v2.0.9`, `lipgloss/v2 v2.0.6` and `ultraviolet v0.0.0-20260811164956-006e29f97886` compiled and ran, and one type satisfied both `tea.Model` and `uv.Drawable`. `x/vt`'s own `go.mod` asks for an OLDER `ultraviolet` pseudo-version; minimal version selection promotes it to this repo's newer one and it still works. That promotion is the standing version risk and it is measured green today. The alternatives were surveyed and rejected on evidence: `hinshun/vt10x` has no tagged release and no commit since 2022-03-01, and `liamg/darktile` is a terminal application whose emulator is not importable. The counted cost of taking it: `x/vt` has no tagged release either. It is a pseudo-version off `main` from a vendor moving fast, so its API can change with no semver signal. That is a real cost and it is accepted, because the alternative is owning a VT parser in a repository whose test suite cannot see the bug class it would introduce. #### The update path, which is dictated by `Render` purity `Render` is pure over `State` — `TestRenderIsPure` renders twice and demands byte equality, and `TestElapsedIsPureOverState` sleeps between the two renders and demands it again. A pane that asked an emulator for "the current screen" from inside `Render` fails both. So the screen is materialised in `Update` and `Render` draws finished rows: ``` ConPTY output pipe -> reader goroutine, chunks into a BOUNDED channel (a full channel stalls the child) -> waitPTY() tea.Cmd, shaped like waitEvents (block on one, drain up to a cap) -> Update, case ptyBatchMsg emulator.Write(chunk) -- the emulator is on Model, never on State st.Live.Grid = snapshot() -- decoded rows cross onto State re-arm waitPTY() -> Render draws st.Live.Grid ``` The emulator sits on `Model` for the same stated reason `gateInputs` does: it holds the whole raw stream, and `State` is what the renderer can reach. Only finished rows cross. Coalescing falls out of the shape. One `Update` applies every chunk it drained but takes ONE snapshot, so a flood costs one grid copy per frame rather than one per chunk. The emulator holds the state, so nothing is lost by coalescing — which is the same argument `drainMax` already makes for text. **Resize is issued from `Update`.** `layoutFor` is pure and callable there, so the room computes the pane's cell rectangle, and when it differs from the pseudoconsole's current size it calls `ResizePseudoConsole` and `Emulator.Resize` together. Doing one without the other leaves the grid and the pty disagreeing about the width. Never from `Render`. #### What the pane draws, and what it drops The pane draws the emulator's TEXT grid. It does NOT carry the guest's colours, and that is a decision rather than an omission. Three reasons. A golden may not embed ANSI (§9.5), and `PlainStyles` neutralises telltale's own escapes and not a vendor's, so a coloured grid could not be pinned by a golden at all. Colour in this room is always a SECOND signal for a distinction telltale is making, and the guest's colours are not telltale's claims. And carrying SGR through a clip re-opens the bleed hazard `fit` exists to close: a row clipped mid-run leaves a colour open into the next pane. The cost is real and worth naming: a guest that distinguishes something by colour alone loses that distinction here. Claude Code does not — its own output carries words and glyphs — but that is a measurement of one guest, not a rule about guests. The guest's own GLYPHS are kept verbatim, including under `--ascii`. `--ascii` governs telltale's glyph set, and the live pane's chrome respects it. The grid is another program's screen, which is quoted material in the sense CLAUDE.md already carves out for error text: a room that rewrote what a vendor printed would be showing something the vendor did not print. The pane does not scroll. A live terminal viewport is the emulator's current screen, and the emulator is sized to the pane, so there is no hidden region for a scroll key to reach. The guest's own scrollback stays inside the guest. #### Honest degradation, and the build floor **Only Windows build 10.0.26200 was measured.** ConPTY's documented floor is Windows 10 1809 (build 17763), and that floor is DOCUMENTATION, not a measurement made here. Older builds are also reported to repaint more heavily, so the 97.3%-payload result may not hold below Windows 11. **This is the single largest unverified claim behind this section.** The seat is therefore gated rather than trusted. `StartPTY` reads the running build with `RtlGetVersion` and refuses below 17763 with a sentence naming the build it found and the build it needs. A refusal renders as an `unavailable` card in the pane, which is the shape §4a.1 already rules for a field that could not be read: the room says what it could not do, and does not draw a blank pane that looks like an agent with nothing to say. **Non-Windows compiles and refuses.** `pty_other.go` carries the `!windows` build tag and returns `unavailable on this OS`. A build tag rather than a runtime check, so a Unix build does not carry a pseudoconsole path it can never take, and so the refusal is a fact about the binary rather than a guess made at the moment somebody presses a key. The other measured gaps, recorded so a later reader does not mistake silence for coverage: no streaming agent turn was driven through a pane, no alternate-screen guest was exercised (`x/vt` has two screen buffers and should handle the alternate-screen mode set, and that path is untested here), the longest spike run was twenty seconds so nothing is known about handle or memory growth over a long session, and the documented ConPTY close deadlock did not reproduce on this build — a negative result on one machine, not a guarantee. #### The spawn guard grows a sixth var, in the same change `startPTYSession` is declared beside the other spawn vars in `persistent.go` and takes a `runner.Spec`, so `refuseRealVendor` needs no fork — a second refuser would be a second thing to keep honest. `TestMain` wraps it, and `countSpawns` stubs it with the restore in the existing `t.Cleanup`. This is not optional and it is not a follow-up. A council test never spawns a vendor, and a PTY child is a vendor process like any other: the fact that its output is DISPLAY ONLY changes nothing about whose account it runs on. A spawn that escaped the count would let "nothing was spawned" pass over a vendor launched with a pseudoconsole attached, which is the exact assertion those tests exist to make. #### The seat is seated by a flag, and by nothing else `--live claude` names the seat. Parsing lives in `ParseLive` in `internal/council` and `Run` calls it before the alternate screen, so `cmd/telltale` hands over one field and learns nothing about vendors — and an impossible seat is a line on stderr rather than a card behind a TUI the user has to quit to read, which is the discipline `--brief` and `--trace` already follow. There is no key. A key would have to be taught on the help panel and would be a second way to spend a second account that the operator has not asked for; the flag is a decision made once, at the point where the room is opened, and that is the right place for a control that doubles a bill. **Nothing types into the pane either.** `PTYSession.Write` exists because the pseudoconsole's input pipe exists whether or not anybody writes to it, and no council code calls it. So the first cut is a WINDOW, not a terminal: the operator watches the agent's own screen and answers it, when it asks something, in the seat's structured column beside it. Sending keystrokes needs a key encoding, a focus rule and an answer to what `q` means while a pane has the keyboard, and none of those is a display question. Owed, and named here rather than half-built. #### Verification `go vet ./...` clean, `go build`, and `go test ./internal/council -timeout 20m`. The live seat's tests drive the emulator and the render path directly, never `startPTYSession`, which is the same shape `arenacheck_test.go` uses for the check var. Two goldens are new — `live-seat.txt` and `live-seat-ascii.txt` — and no existing golden moved. The feature is entirely opt-in: `State.Live`'s zero value is off, no footer hint was added, and no chrome row was spent, so every existing frame renders byte-identically. The goldens pin the seat's chrome and a FIXED emulator grid produced by feeding a fixed byte script through `x/vt` in the test. They pin what a golden can pin: that escapes were consumed, that the grid is exactly the pane's rectangle, and that the display-only marker is a word rather than a colour. **Not verified here: a live drive.** No `claude` interactive session was run through a pane by the session that wrote this. The spawn guard makes that impossible from inside the suite by design, and it is the same class of debt as the host's first live turn ([STATE.md](../STATE.md)): an operator-driven check, owed and named rather than implied. ### 9.54 the room was a committee, and a crew's seats are busy one at a time (2026-09-02) `dispatch()` opened with one line — `if m.turn != nil { "a turn is already in flight — ctrl+c cancels it" }` — and that line was the whole difference between the room that existed and the room the owner asked for. Seats were concurrent inside a turn and the room was serial across turns: `@all` fanned one brief out to five processes at once, and the moment any of them was still answering, a second brief to a different seat was refused. You could not hand codex a refactor and, while it ran, hand grok the docs. The owner's ruling is that council is a **crew**, not a committee answering one question at a time, and a crew's seats are busy or idle **one at a time**. #### What a turn is now A turn is a fact about a SEAT. `Model.turn` is gone; `Model.turns` maps each seat in flight to the dispatch it is answering, and `turnState` is the record of ONE DISPATCH — the seats one press of enter sent a brief to, and what those seats share while they answer it: the turn number, the route the header names, the arena's all-or-nothing bookkeeping. What moved down to the seat is everything the operator can now do to one seat while its neighbours work: its process handle, its own child context (so cancelling it kills its child and nobody else's), the cancellation word, the give-up. **The turn number stays one room-wide sequence**, and that is a ruling rather than a leftover. A per-seat count was considered — "codex's turn 4, grok's turn 4" — and refused, because every surface that already prints a turn number reads one coordinate: the separators, the by-turn page (`PageTurns`), `/retry`, `room.json`, the reattach card. Two seats both on "their turn 4" would leave the page able to open only one of them. So a turn number is a **dispatch** number: turn 5 is the fifth brief the room sent, whoever it went to, and a seat's own history is the subset of those numbers it took part in. `Column.TurnN` already carried exactly this. #### What the operator sees - **A brief to a busy seat is refused for THAT seat, by name and turn.** `@codex …` while codex is mid-answer: `a turn is in flight on codex (turn 4) — ctrl+c on its column cancels that turn, or address another seat`, and the draft stays put. `@all` with codex busy goes to the idle seats and says so: `sent to grok, agy — skipped: codex (turn 4), still on a turn; ctrl+c on its column cancels it`. Both halves are measured — the seats in the new dispatch's live set, the seats whose own turn refused them. The refusal is what stands between a persistent seat and a second prompt written into a process mid-turn, which is the failure the room-wide wall was in front of. - **The header names the newest dispatch and counts the rest.** `turn 5 → codex · 3 in flight`. The count is `State.SeatsInFlight()`, measured over the columns, and it is printed only when some seat in flight is on a turn OTHER than the one the cell names (`inFlightBeyond`, read off `Column.TurnN`) — an `@all` turn with its three seats streaming reads `turn 3 → everyone` as it always did, because `· 3 in flight` beside it would be the route restated as a number. A live column with no turn number at all is not counted as another dispatch: zero means never dispatched, and a count that read it as elsewhere would be inferring. The route retires when ITS dispatch lands (`ts.n == st.Turn`), not when the room goes quiet. - **ctrl+c has three meanings and the footer says which is live.** The focused seat's turn if it has one (`ctrl+c cancel codex`); everything in flight when the focused seat is idle (`ctrl+c cancel all`); quit when nothing is. With one seat in flight the label is the plain `cancel` every earlier frame carried. The gate line's `cancel the turn` became the same label, because the key reaches `viewKey` through `gateKey`'s fall-through and means there what it means here. `x` is unchanged: the per-seat give-up with its card. - **The room-wide refusals name the seats.** `q`, `/cd`, `/seat`, `/unseat`, `/read`, `/write`, `/retry`, `/adopt` and `/arena drop` still need the whole room idle — each changes something a busy seat was dispatched against — and each now says `a turn is in flight on codex (turn 4), grok (turn 5) — …`. `c` and `u` became per seat: the thread or the tree they touch is the focused seat's alone. - **A race is the one turn that still owns the room.** `/arena` refuses while any seat is busy (a race that skipped a seat is not a comparison), and every brief is refused while a race runs (its racers are writing into worktrees a room brief would cut across). `race()` is the read. - **A `/flow` hop waits on its own seat.** Hop N+1 dispatches the moment hop N's seat lands, whatever the rest of the room is doing; the chain's death-on-teardown runs only when the hop's own dispatch ends (`turnState.flow`). A hop whose seat is busy with an unrelated brief stops the chain by name — `flow stopped at hop 2/3: @codex is still on turn 4 — …` — rather than queueing behind the seat, because a chain that dispatched itself later, when a seat happened to free up, is the room acting on its own at a moment nobody chose (§9.16's argument). - **A rebuttal quotes what a busy neighbour last FINISHED saying.** The snapshot is taken per seat at its dispatch; a seat mid-answer contributes its last filed turn (`settledReply`) and nothing if it has never filed one. Quoting the half it has streamed would put half an argument in front of another model as though it were whole. #### The inbox §9.40's strip named only the seats stopped on a gate. A crew has a second stall of the same shape: an answer lands in a column the reader is not on, the room knows, and the reader has to go looking. The strip now lists **seats whose turn ended since the reader last had the keys on them** — `⚠ NEEDS YOU 2 Codex 3 Grok done 4 Cursor failed` — with the terminal phase word, the same word the column header speaks, as the whole distinction between the two kinds of entry (no glyph, no colour: it reads the same under `--ascii` and `NO_COLOR`). *Amended 2026-09-03 (`LEDGER.md`):* the lead is `NEEDS YOU` only while a listed seat is blocked on a gate. A strip of landed replies and nothing pending opens `UNREAD`, with no mark (`unreadLead` in `internal/council/needsyou.go`); the entries and their phase words are unchanged. Every entry is a measurement, on §9.40's own rule. A landing is two stamps the Model took itself: `Column.Ended`, written in `finishColumn` while the seat still holds a turn (so the second retirement a persistent seat or an ACP racer goes through cannot re-stamp it), and `Column.LastFocus`, written by `setFocus` on BOTH the seat left and the seat entered, so it marks the end of the reader's last look. The strip is `Ended.After(LastFocus)` over terminal columns. Nothing stores what was acknowledged — the comparison cannot drift, which is the argument §9.40 made for deriving the gate half from `Focus`. The default-focus hole stays open for §9.40's reason and closes itself: the focused seat is never listed, and the reader's first departure stamps it. `.` is the strip's key: the next listed seat after the focus, wrapping. The footer names it only while the strip has an entry (§7.8: never a key that does nothing); in compose it is a full stop. #### Mechanics worth knowing - **One event reader.** Every dispatch and every batch re-arms the pump, and two goroutines reading one channel would deliver batches to `Update` out of order — an exit before the text it followed. `Model.eventsArmed` makes `waitEvents` hand out one reader; `Update` clears it on the batch. `sendTurn`'s Cmd can therefore be nil for a dispatch that DID start, so `applyArenaSetup` reads `race()` rather than the Cmd. - **The give-up and the cancel outlive the seat's turn.** `givenUp` and `cancelling` moved from the turn to the Model, keyed by seat, and are cleared at the seat's next dispatch. A cut seat's turn ends the instant its column lands, while its process is still draining — and on the persistent seat still answering the interrupt with a failed `result` — and those echoes met a seat with no turn, where `applyEvents` would have written the abort error over the give-up's own note. - **Teardown walks `dispatches()`** — every distinct record in the map — and reaps each one's one-shot handles, racer handles, ephemeral sessions and context. `TestTeardownReapsEveryDispatch` pins two dispatches; `teardown_test.go` still pins the racer. - **Geometry.** `frameOwnersFor` counts a column still in flight as an owner, read off the column's own phase, so a brief to grok while codex streams does not narrow codex's prose under the reader. The no-mid-stream-reflow rule is about the room moving because a VENDOR did something; a dispatch is the operator's act. #### What this section does NOT change `room.json` gains no field: sessions and the turn count are room-level and the save runs at each dispatch's end. `--trace` clocks are per process and unaffected. Session resume per vendor is unchanged — a seat is refused a second prompt while busy, which is the only new rule a resume could meet. The spawn guard is untouched: every crew test dispatches through `countSpawns`. #### Verification `gofmt`, `go vet ./...`, `go test ./... -count=1`, `go build ./...`, and the windows/amd64 and darwin/arm64 cross-builds, all clean; `go test -race ./internal/council` once. Five goldens are new — `two-in-flight`, `busy-seat-refused`, `inbox-landed`, `inbox-landed-ascii` — and the existing goldens that moved moved on ONE cell each: the footer's cancel label, on frames where two or more seats are in flight, which is the key's new meaning drawn honestly. No header of an existing frame moved. **Not verified here: a live crew.** No two vendors were run concurrently by the session that wrote this. What the suite pins is the room's bookkeeping over stubbed processes; what only a live run can show is a persistent seat taking its NEXT brief cleanly after a per-seat cancel, two one-shot seats' event streams interleaving through one reader without a stall, and a `/flow` hop landing beside an unrelated seat mid-answer. Owed and named rather than implied, on §9.53's rule. ### 9.55 a crew's writers share one tree, and the room becomes the integrator (2026-09-02) §9.54 made the seats concurrent and left them where they were: every ordinary turn ran in `State.Workspace`, and a worktree existed only inside `/arena`. The moment two writing seats could answer at once, two writers were in one checkout — which §9.37 had already ruled out for a race in one sentence, "four writers in one shared tree are not four answers, they are one trampled tree". And the containment could not be a card: only the stream-json seat can be asked before a write (`canGate`), the other four act unasked. So the crew's containment is the race's, made structural and made permanent: **a writing seat gets its own worktree, cut once from the room's HEAD and reused for every writing turn after it, and the room is the integrator** — nothing a seat writes reaches the room's tree except by `/adopt`. #### What changed - **One worktree per writing seat, by default.** In write posture the first brief to a seat cuts `-seat-` beside the workspace on `seat/`, from HEAD, and the seat's process runs there (`seatDir` is the one read every spawn path — `specFor`, `seatProcess`, the flow receipt, `/hand` — goes through). Read posture and a `/flow` read hop keep the shared tree: a tree cut for a hop that cannot write is containment for a hazard that does not exist. `--shared-tree` is the opt-out, a FLAG rather than a room word because it decides where five processes write, the same class of fact as the workspace. The cut runs off the render loop under arenasetup.go's deadline, serially, with the step named on the footer (`worktree: preparing worktree for codex…`) and ctrl+c stopping the setup; a tree already cut is kept. A tree an earlier room left is found by name and reused, and the notice says so; a directory at that name that is not the seat's worktree is refused by name. Nothing is seeded into a seat's tree and no brief file is written into it — those are affordances for a fresh throwaway tree, and a seat's tree is reused. - **The badge says which holds** (`Column.Containment`, stamped at dispatch from the same read the spawn uses): `wt: seat/codex`, `shared tree`, or `⚠ shared tree · ` for a fallback the room could not avoid, with git's own sentence in the notice when it happened and the whole reason on the `?` postures page. At a three-seat column's width the reason sheds before the word and the mark stays, because a clipped reason is not a reason (§9.11); the granularity word sheds first, on stripBadges' own order. A seat never dispatched wears no badge: no process, no directory, no claim. `--ascii` keeps the mark as `!` and the separator as `-`. - **`/adopt ` merges a seat's branch, the race's way** (`adoptSource`). The card names `git merge --no-ff seat/codex` onto `adopt/seat-codex`, refuses a dirty room and an empty branch exactly as before, and after the merge **resets the seat's tree and branch onto the new HEAD** — a reset that fails degrades the notice, never the adoption, and says what moves the tree by hand. Resolution order is a ruling: the arena branch when the seat's CURRENT turn is a race attempt (that is the block on the column), else the seat branch, else the race receipt. Hybrids work across two seat branches on `adopt/seat-+`; a hybrid across a race attempt and a seat branch is refused, because one receipt names one kind of branch. The donor is not reset — only its named paths were taken. - **`/arena record` counts seat adopts apart** (`ArenaRecord.SeatAdopts`, from `adopt/seat-[-k]` refs): their own sentence under the window, inside no rate. A race adopt is a verdict among competing attempts at one brief; a seat adopt is the room taking one seat's ordinary work with nobody else in the running, and folding it into a rate would score a seat for races it never entered. - **`/flow` fans with `&`** (`FlowStep.Stage`): hops joined by `& @seat` are one stage and one dispatch, each seat handed its own task; the stage after `->` dispatches when the stage's last seat lands (`StageDone`), carrying every predecessor's reply as its own labelled fence, in landing order. The header names the stage (`hop 1/2 @codex & @grok`). Two refusals a fan adds, at parse: a seat named twice in one stage, and a stage mixing write and read hops — a stage runs at ONE posture, because a fan is one dispatch and §9.16's table gives a dispatch one posture. One `y` releases a writing stage and the card names every target. A busy seat stops the whole stage by name, on §9.54's rule. An ampersand not followed by a mention stays prose. - **`/hand `** (`handcmd.go`) puts ``'s stat and patch — against the point its branch parted from the room's HEAD, read with `add -N` so created files count — into the draft addressed `@`, fenced as measured git output with the tree and branch named. A draft, on §9.49's whole argument; `enter` is the only thing that spends. The cap is the composer's: a patch past it is cut at a hunk boundary and the closing fence states the cut and the way to the whole (`y`). #### What this section does NOT change `room.json` gains no field: a seat's tree is rediscovered on disk by name. The arena's own worktrees, seeding, brief file, ranks, `x`, `u` and `/arena drop` are untouched — a racer's badge now reads `wt: arena/t/`, which is the tree it already stood in. The spawn guard is untouched: every test here dispatches through `countSpawns`, and the git that runs is against a temp repository. ### 9.56 the gauges route, and a real run can be shown without a vendor (2026-09-02) Two changes, one premise: §9.54 made the seats a crew, and a crew changes what two existing things are FOR. The quota readings (§9.21) used to say "this turn may not land"; on a crew they answer "which seat has room for this brief". And the room's event stream, which every column already applies, is the one thing that could show the room to someone who has not paid for five seats — if it were kept, and if what was kept could be shown honestly. #### The readings as routing The routing cell qualifies the route. `→ codex · 5h 94% used` when the draft addresses a seat whose relayed window is at or above `--headroom-warn` (default 90) and under a hundred; `→ everyone · codex 5h 94% used` on a route that names a set, so the seat is named. `@auto` is a route word (`Route.Auto`) that resolves against State: among seated, idle seats with a measured reading, the most headroom in the SHORTEST window — the window that resets soonest, since that is the one the next brief spends from — ties to seating order. The cell states the pick before enter, `→ auto: grok (5h 12% used)`; enter rewrites the draft to `@grok …`, dispatches, and the notice repeats the choice. The header and room.json record a turn to grok, which is what happened, and no surface learns a fifth route shape. Every rule is §4a.1's, restated for a rank: - **Nothing is invented and nothing is aggregated.** The hint copies one window's label and percentage. `@auto` ranks headroom (a hundred minus the vendor's own figure) and prints the figure it ranked. No total, no average, no dollar. - **An absent seat is never ranked.** Cursor and grok have no reading anywhere (§7.17), an unrelayed Claude has none today, and a rank needs a number. With no measured seat in the room the cell says `→ auto: no measured reading` and enter refuses — `@auto needs a measured reading; none of the seated seats has one` — rather than falling back to the default route. A brief handed to the readings must not go quietly to Claude because the readings were empty. - **A stale reading is absent, not a number.** A window whose reset has passed is dropped by quotacache on read and by `measuredWindows` between reads, when the room's own clock passes it. A reading past `quotaAgeWarn` is the alarm's to name as stale (§9.21) and is absent here. - **The hint stops one short of a hundred.** A full window is the alarm's cell, with the warning mark, and one fact in two cells on one line is the drift vendorTag's comment warns of. - **A busy seat is never picked.** §9.54's refusal would meet the pick a keystroke later. The threshold is council's own pick, which is why it is a flag: the default is a boundary no vendor published, so the operator can move it, and the cell prints the vendor's figure beside whatever threshold applied. `quotaFullPercent` is unchanged and still the only threshold this package did not choose. #### The recording, and the boundary it is written under `--record ` writes the room's event stream as it happens: the room line (seats, postures, posture, workspace as the header drew it), each dispatch as its seats hold it (turn, route, each seat's brief as `startTurn` echoed it, whether it ran on a persistent process), every `runner.Event` in every batch `applyEvents` receives, and each gate decision the operator made. JSON lines, millisecond offsets from the monotonic clock, one flat record type so a fixture can be typed by hand. `runner.Gate.Input` is the one field dropped: never rendered, a Write's whole file content, and a replay has no vendor to hand it back to. **This is content, and it is not argued as one of the numbers-and-keys writes.** CLAUDE.md's read/write boundary holds room.json, the quota relay and the token relay to numbers and keys. The recording is the event sink's kind of exception — "different in kind", contained by scope rather than redaction — and it is contained three ways: an explicit path the operator typed (a path under `~/.telltale` is refused before a byte is written, and an existing file is refused rather than overwritten or extended); an explicit flag typed at the door (no key in the room starts one); and off by default (the recorder is nil, every hook on it is a no-op, no test builds one). `recording.go` carries the argument at length; this paragraph is the decision. **No redaction, on purpose.** The recorder writes what `applyEvents` saw. A recording that differed from the run would be a second truth; the replay puts the same bytes through the same redactor at the same choke point, so the frames match and the file underneath is the raw stream. The review the README requires of every frame is given a tool instead: `telltale council replay-check ` lists the workspace, the seats, every session id, every tool line and gate card, and the size of the prose — and says it did not read the prose. #### The replay `--replay ` opens a room over the file. `Run` returns to `runReplay` before `LoadRoom`, the brief, the trace, the relay or a host is consulted. The State is built from the room line — seats, labels, postures, granularities, the abbreviated workspace with `Home` left empty — so a replay on a laptop with nothing installed draws the recorded room. Each dispatch line puts the recorded seats on the recorded turn through `startTurn` and `holdTurn` (a `turnState` with no handle and a no-op cancel); each event line goes through `applyEvents`; each gate line takes the card down the way `decideGate` does, with the wait charged from the recording's clock. **The clock is the recording's, and `Render` stays pure.** `State.Now` on a tick is the recording's start plus the wall time elapsed since the replay opened, times `--replay-speed`; each record stamps its own offset on the same clock and never moves it backwards. The two wall reads the live retirement path makes — `finishColumn`'s Elapsed and the charged gate wait — are stamped from the recording before the event lands, using the guards those paths already have. A fixture played through `Update` therefore renders the same bytes every time (`replay`, `replay-gate`, `replay-ascii` goldens), and two plays of one file are one frame. **Labelled on every frame, in words.** `⚠ REPLAY` in the header where WRITE/READ sits; `REPLAY` first on every column's badge row (the whole row at strip width); `⚠ REPLAY nothing here is live` where the compose footer's routing cell and `enter dispatch` would be; `ctrl+c quit` on the view line. A replay is a real run played back — not §8's invented recording, because somebody ran it and every event was measured — and the label is what the honesty rule asks in return: a replayed frame must never pass for a live one. **Three refusals, and nothing else.** Enter, the card's `y`/`n`/`a`, and the per-seat verbs say `this room is a replay; nothing here is live`. `ctrl+c` and `q` quit in every mode. `saveRoom` returns on the replay flag, so a finished turn and a teardown write nothing; `TestAReplayNeverTouchesRoomJSON` plants a sentinel room under the suite's sandbox home and reads it back unchanged. No spawn var is reached: `Init` starts the feed and nothing else, which is the spawn guard's own "no test calls Init" argument met from the other side. #### What a recording does not hold The operator's cancels and give-ups (a cancelled column replays as the vendor's own exit), focus and scrolling, the `--brief` file's text, and `--trace`'s clocks. A recording with a long idle gap replays the gap at the recorded pace; `--replay-speed` is the remedy, not a cap. #### Verification `gofmt`, `go vet ./...`, `go test ./... -count=1`, `go build ./...`, the windows/amd64 and darwin/arm64 cross-builds, and `go test -race ./internal/council`, all clean. Goldens new: `containment-badges`, `containment-badges-ascii`, `containment-badges-expanded`, `flow-fan`, `arena-record-seats`, `arena-record-seats-ascii`. Goldens moved: the two `slash-refusal` frames, on one line — the refusal's vocabulary gained `/hand` and paid for it with "room" and an article. No other existing golden moved: a column dispatched before this section carries no containment claim, so no frame built before it changed. **Not verified here: a live crew in its trees.** No vendor was run by the session that wrote this. What only a live run can show: a persistent seat resuming its thread after the respawn that moves it from the workspace into its tree (the same `--resume` composition `/cd` measured, in a directory that did not exist a second earlier); a one-shot seat honouring the tree as cwd on a resume (`codex resume` rejects `--cd`, and the tree is passed as `Dir`); two seats writing at once into two trees and `/adopt` folding each in turn; the fanned `/flow` stage's two artifacts landing beside each other and the join reading both; and the `⚠ shared tree` badge at real widths on a real fallback — SEEN 2026-09-03 on every seat of a five-seat room opened in `~` (not a git repo) at about 180 columns, with the `· ` clause shed at that width; the clause itself at a wider room is still unchecked. The rest owed and named rather than implied, on §9.53's rule. darwin/arm64 cross-builds, and `go test -race ./internal/council` once. New goldens: `route-headroom`, `route-auto`, `replay`, `replay-gate`, `replay-ascii`. No existing golden moved: a room with no near-full reading and no `@auto` draws the routing cell it drew before, and a live room's header, badge rows and footer are untouched by the replay flag. **Not verified here: a live recording.** No vendor was run by the session that wrote this, so a recording of a real room EXISTS as of 2026-09-03 — the owner ran `--record demo.jsonl` over a five-seat gated write room at turn 3 with one `@all` brief, and `replay-check` read it clean: 75 records over 24s, five seats, one dispatch, 59 text events, no tool calls, no gate cards (the file carries verbatim prose and session ids, so it stays local). `--replay` of that file has NOT yet been driven; the replay fixture in the suite is still synthesized (fake ids, fake paths, three seats, one card, one finished seat). What only a live run can show is a `--record` of a five-seat room with real timing, that its replay reads as the room did, and what `replay-check` lists off a real capture. The hero decision stays the owner's. **Measured 2026-09-03: `--replay` of a real recording, and it found a defect.** The owner recorded a four-seat gated write room (`--vendor claude,codex,agy,grok`) over seven dispatches: 1863 records over 39m36s, `replay-check` clean. `--replay` of that file drew every persistent seat as `failed` right after each dispatch, with an elapsed figure of an hour and the note "the vendor process ended mid-turn". The live room had shown those turns done. The mechanism, read off the file: each dispatch to a persistent seat was followed 0.1s to 11s later by a `done` from that seat, before its first text of the turn. That is the seat's PREVIOUS process ending after the respawn that replaced it (`stopProc`). The live room attributes that exit to the old process by one liveness test on the current one, and discards it. The recording carried only the vendor name, so the replay handed it to `applyEvents` as the new turn's terminal event. Two fixes, and the reason for two. The recorder now writes from inside `applyEvents`, after the stale-exit guard (`Model.staleExit`, the one liveness test the `KindDone` and `KindError` branches used to make inline), so a file carries what the room applied and not what the channel delivered. That is the fix at the source, and it is exact. The reader (`markStaleExits`) repairs a file recorded before the guard, which the first real recording is: a process's exit is the last event it emits, so a persistent seat's exit that another event from the same seat follows before its next dispatch was not this turn's end. A replay consumes such a line and does not apply it, and `replay-check` prints how many it will skip. The same rule reads the first-turn retreat to a batch adapter correctly, where the room also discards the exit and the seat goes on. Three tests pin it over a fixture with one stale exit and one real death (`stale-exit.jsonl`). Two smaller findings from the same file. The room line listed five seats and the room-open rebuild reported `5/4 seats rebuilt`: a seat `--vendor` left out was still marked `Restored` on reattach, so the rebuild launched it and the recorder listed it. `reattach` now marks a seat restored only if `State.seats` says it takes turns, and the room line lists the columns `VisibleColumns` draws. And `replay-check` printed one bare `grok` line per tool RESULT (an act carrying an outcome and no text, 112 of them for one seat); those fold into one count per seat. The replay half of the owed measurement is done; the hero decision still stays the owner's, and the rendering of that recording belongs to the density pass. **2026-09-16: the replay draws its provenance on the room line.** Before this date, the recording's stamp lived only in `replay-check`'s stdout, and the scrubbed claim lived on the entry notice and the closing notice. The entry notice is gone at the first dispatch, so a reader who looked at a frame in the middle of a replay saw `REPLAY` and nothing that said which run it was or whether the prose was real. The room line (`roomline.go`, `replayFact`) now prints the provenance on every replayed frame, first among the room facts: `recorded 2026-09-03 21:14 -0400` for a capture, and `recorded 2026-01-01 09:00 UTC · scrubbed: the shape is real; the date and every word are synthesized` for a scrubbed file. Each clause follows §4a.1. The date is the file's own stamp in the zone the recorder wrote. A file with no readable stamp draws no date; it never draws the epoch or this machine's clock. A scrubbed file's stamp is `scrub.go`'s constant, so the clause says the date is synthesized with the words. The vendor CLI versions are NOT drawn, because the format carries no version field (`recordLine`) and a version read off this machine would describe a room that ran somewhere else. A version field on the room line, written when the room learns one, is owed and is not started here. The two notices keep their words. The cost is one room-line row on every replay frame, so the `replay`, `replay-gate`, `replay-ascii`, `demo-gate`, `demo-final` and `demo-ack` goldens moved by that row, and `demo-compare` and `demo-compare-route` are new (the compare is `docs/room-identity.md`'s 2026-09-16 section; `panes-keys` moved by the `c compare` cell and `panes-compare` is new). `TestTheReplayFactIsHonestAboutWhatTheFileCarries` pins each clause. ### 9.57 three seats stay up between briefs, on a reading rather than a run (2026-09-02) Until this change the room kept exactly two processes alive across turns: the Claude seat (`claude -p --input-format stream-json`, §9.8, measured) and the Cursor seat (`cursor-agent acp`, §9.36, measured). Codex, Antigravity and Grok each paid a whole process per brief — `codex exec --json`, `agy -p`, `grok --single=` — and two of those cold starts were measured on the reference box at **5.6 s** (Cursor, before §9.36) and **6.4 s** (Antigravity). A crew tool that dispatches a brief every few minutes spends most of a seat's wall clock on startup that way, and a seat that starts fresh every turn has no channel on which to be asked anything. This section records the move of all three to long-lived processes, what each was built from, and — the part that matters more — the fact that **none of the three shapes has been driven from this repository.** Every badge on those columns says `unmeasured` and names the version it was read at, and the checklist at the end is the price of removing the word. **This is a departure from §9.50's ruling and says so.** §9.50 left the codex app-server protocol unseated because a read-posture turn on it was measured failing to inspect in two of three arms, and named the read posture's liveness as the measurement owed before any flip. That measurement has not been made. The seat moved anyway, on the crew ledger above, with three things holding the honesty line: the measured batch invocation stays one step away as the fallback, the seat owns its own kill, and the badge says what nobody has watched. The owed measurement is now the FIRST item of the checklist rather than a precondition, and that ordering is a decision, recorded here so nobody reads the registry and assumes the debt was paid. #### What was built, and from which pages | seat | live shape | fallback (measured) | built from | |---|---|---|---| | Codex | `codex app-server`, one process, `thread/start{cwd, sandbox, approvalPolicy}`, `turn/start` per turn, `item/*/requestApproval` answered through the room's card | `codex exec --json` (codex.go, 0.149.1) | the protocol capture of §9.50 at **0.149.1**; the app-server README on `openai/codex` main and `app-server-protocol/src/protocol/v2/{shared,item}.rs`, read 2026-09-02, for the `approvalPolicy` enum (`untrusted`, `on-request`, `never`), the v2 decision enum (`accept`, `acceptForSession`, `decline`, `cancel`), the approval params (`itemId`, `command`, `cwd`, `reason`, `grantRoot`), and `turn.status: "interrupted"`. Installed build **0.152.1**, undriven | | Grok | `grok agent stdio`, the ACP client of §9.36 under a second dialect, `session/new{cwd}`, `session/prompt` per turn, `session/request_permission` answered by kind through the room's card | `grok --single=` (grok.go, 1.0.4) | docs.x.ai/build/cli/headless-scripting and zed.dev/acp/agent/grok-build, read 2026-09-02, for the subcommand and the `--cwd` / `--resume` flags this seat deliberately does not pass; agentclientprotocol.com/protocol/schema for `loadSession`, the option `kind` values and the `cancelled` outcome. Grok Build **1.0.13**, undriven at the time; driven 2026-09-04 (the checklist below) | | Antigravity | `agy --input-format stream-json --output-format stream-json`, one process, `{"event":"user","message":{"content":…}}` per turn, `result` ends the turn | `agy -p` (agy.go, 1.1.13) | antigravity.google/docs/cli/headless and the changelog entry for **1.1.15** (2026-08-19), read 2026-09-02, for the flag, the envelope, "one turn per message in a single conversation", and "close stdin … the process exits after the input pipe is closed and the current turn completes". Installed build **1.1.24**, undriven | Three shapes in the package carry the move: - **`vendors.LiveFallback`** names the measured batch adapter behind each live seat. `FallbackRegistry()` is the whole room after a retreat, and it exists so that state can be constructed rather than scripted — the give-up and re-send suites build "an ordinary one-shot seat" from it, because the default registry no longer drives one. `Registry` became a var for that reason, on the spawn vars' precedent. - **`vendors.GracefulStop`** is the seat owning its kill. §9.50 measured a closed stdin NOT reliably ending `codex app-server` (four exits in 1.5–3.3 s, one alive at 15 s), so the app-server protocol's `Closing()` cancels any held approval and interrupts an open turn, `Grace()` bounds the wait at 4 s, and the runner's job-object kill is unchanged behind both. The ACP client and the Antigravity seat implement the same pair. - **`acpDialect`** is everything the shared ACP client may vary per vendor, and it is three fields: the read-posture mode id (cursor: `plan`, measured; grok: none), whether permission answers use the measured option spelling or pick by kind from the request (cursor: fixed; grok: by kind, `cancelled` when the kind is not offered), and whether `session/load` waits for the server to advertise it (grok only). `cursoracp.go` is `acp.go` now, with every measured sentence intact. **Codex's approval policy is a choice, and the alternative is named.** Read posture asks for `never`; write postures ask for `on-request`, the vendor's own interactive default, under the `workspace-write` sandbox — so the vendor asks when it wants more than the sandbox allows and the room cards that. `untrusted` would ask about every command off the vendor's trusted list, and the pages read for this do not say whether a command approved under it then runs outside the sandbox. A policy that might trade the sandbox for a keystroke is not one to adopt unmeasured. The gated posture never reaches the seat as itself: `spawnPosture` collapses it to write for every Conversational seat, and whether a request becomes a card or an automatic yes is the room's decision when one arrives. **Two things the codex badge did, in opposite directions.** On Windows the read level HOLDS at `ro:enforced`, because the same 0.149.1 session measured the sandbox on the app-server path — a write through cmd.exe denied with no file on disk — and the detail now carries that path's sharper liveness residual. Off Windows the level DROPS to `ro:requested`: the macOS enforcement was `codex exec`'s (2026-08-05, 0.146.0), every app-server arm ran on Windows, and a seat move re-measures rather than inherits. That is the one badge this change lowered. **What is unchanged, stated so it is not inferred.** `canGate` still names one seat. Two more can now be ASKED — a request the vendor raises reaches a person — and neither has a coverage measurement, so both stay `WRITES` with `asks · unmeasured` where the argument lives. Grok's cost figure comes only from the `--single` fallback: the ACP prompt response is `{stopReason, _meta?}` and nothing read says what grok puts in `_meta`, so on the ACP seat cost renders **absent**, never zero. Antigravity's `--conversation` on the stream argv is unmeasured for composition, and the §9.43 fork tell is what makes sending it safe: the seat implements `SilentResumeFork`, so the room compares the id it asked for against the id `init` reports. **What the core does with the two interfaces (landed 2026-09-02, in the crew integration).** This paragraph said "nothing yet" while the six crew lanes ran apart; the core was patched once they merged, and it now reads both. `LiveFallback`: a seat whose protocol reports `Dead()` — on the brief that finds it dead in `handTurnToSeat`, or on the failed-turn event the handshake refusal produces — and a persistent seat whose process dies on its FIRST turn before it names a session, both retreat to the batch adapter for the rest of the room (`fallback.go`). The retreat is on the SAME dispatch: the brief the operator typed goes down the batch branch with the operating brief applied, the column stays on its turn with a note naming the invocation, the badge reads `seatShape(v, true)` through `postureClaimFor`, and the seats in flight beside it are untouched. `GracefulStop`: teardown, a respawn, and a cancel that has to become a kill all run `stopProc` — `Closing()` down the pipe, `runner.Session.CloseInput`, `Grace()`, then the kill that §9.50 measured necessary — with teardown waiting on every seat's grace at once. Both are pinned over stubbed sessions in `fallback_test.go`; the live run each one owes is on the checklist below, and STATE.md carries the consolidated list. #### The live measurements owed, as a checklist Each item names the command to run on the reference box and the sentence it would let the badge change. Capture the wire under `vendors/testdata/wire/--.jsonl` per the README there; a claim that is not captured is not made. - [x] **Codex, read liveness (the §9.50 debt).** PAID. MEASURED 2026-09-04 at codex-cli 0.151.0, Windows 11: a hosted read room (`telltale council --host --read --cd --vendor codex`, driven through the plain client with stdin piped), the brief "list this directory and read README.md" three times; every turn ran `cmd.exe /c 'dir /a'` and `cmd.exe /c 'type README.md'` and settled `done (exit 0)`. Three of three inspected. The transcript is filed privately (desk/research, 2026-09-04). The Windows detail's "two of three read turns" is retired by this run; it was true at 0.149.1 and is not true at 0.151.0 on this path. What this run does NOT re-measure: the sandbox's refusal of a write under `-s read-only` (no write was attempted), so PARITY's `ro:enforced` still rests on the 2026-08-29 measurement. - [ ] **Codex, the approval flow, both branches.** A write-posture thread (`workspace-write`, `on-request`) asked to write OUTSIDE the workspace and to reach the network: does `item/commandExecution/requestApproval` or `item/fileChange/requestApproval` arrive, does the vendor BLOCK until answered, does `accept` run it and `decline` stop it, and does the file land or not. This is what lets `asks · unmeasured` become `asks · measured`. - [x] **Codex, `turn/interrupt` and teardown order.** MEASURED 2026-09-05 at codex-cli 0.151.0, Windows 11, five runs of `codex app-server` through the same `codex.cmd` shim `doctor` names as the seat's entry point, with the seat's own `initialize`, `thread/start{sandbox:"read-only", approvalPolicy:"never"}` and `turn/start` frames on this repo as cwd (transcript filed privately, desk/research). **The interrupt half passes clean:** `turn/interrupt` on a running read turn produced `turn/completed` with `status: "interrupted"` on every run, 0.10–0.13 s after the frame. **The teardown half fails the grace:** with no turn live `Closing()` sends nothing, and after stdin close the process exited 0 after 6.79, 1.76, 6.69, 7.77 and 6.42 s — four of five over the 4 s grace, so the room's reaper ends this seat before it ends itself most of the time. No stderr on any run. **DIAGNOSED the same day, same build, forty runs, stdin closed with a thread open:** the shim is innocent (`cmd.exe /c codex.cmd`, `node.exe codex.js` and the native `codex.exe` each took 6-8 s, five runs each); no shutdown frame exists in the schema, and `thread/unsubscribe` and `thread/archive` before the close changed nothing; a process with no thread exits in 0.03 s; with `--disable hooks` the same process exited in 0.06-0.08 s with its MCP servers still configured, and with `mcp_servers={}` and the hooks kept it took 4.5 s. The cost is the operator's own `SessionEnd` hooks in `~/.codex/hooks.json`, one of which took 4.5-6.6 s run by hand, and the server writes nothing to stdout while they run. The room may not disable them (the fleet's guard hooks ride the same flag), so the fix is the last resort the chip named, with the diagnosis attached: `Grace()` is 20 s, a bound on the operator's hooks rather than on the vendor. Five runs after, through the shim with an interrupted turn: 4.46, 4.33, 14.73, 7.55 and 2.45 s, all inside it. The 14.73 s run is the same hook under load. - [ ] **Codex, macOS.** The read sandbox on the app-server path, on the Mac, before the off-Windows badge may return to `ro:enforced`. Record in PARITY.md. - [ ] **Codex, the exec fallback trigger.** Run the seat against a build without `app-server` (or an unauthenticated one) and watch what the room shows; this is the arm the room's retreat (`fallback.go`) is built for, and the note and the `exec · unasked · fallback` badge are what it should show. - [x] **Grok, the handshake and a turn.** PAID 2026-09-04 at grok 1.0.13 (5e9a58528b76), `grok agent stdio` driven with the seat's own frames from a scratch cwd (the transcript is filed privately, desk/research). `agentCapabilities`: `loadSession: true`; `sessionCapabilities: {list, resume, close}`; `promptCapabilities.embeddedContext: true`, image and audio false; `mcpCapabilities: {http, sse}`; and an `_meta` block advertising `x.ai/hooks` with blocking events `pre_tool_use`, `stop`, `subagent_stop` and decisions `deny`, `block`. `authMethods`: `cached_token`, `grok.com`. **`session/new` at 1.0.13 returns `sessionId`, `models` and `_meta` — no `modes` and no `configOptions`.** The frame this adapter's header quotes (a `modes{currentModeId:"agent"}` block) was 1.0.4's; at 1.0.13 the server advertises no mode at all, so what `session/set_mode` does for the read posture is now an open question (the permissions item below). A fenced brief streamed back as `agent_message_chunk` (853 chunks), beside `agent_thought_chunk`, `tool_call`/`tool_call_update`, `available_commands_update` and `session_info_update`, and resolved `stopReason: end_turn`. **A brief beginning `/` is NOT eaten on this path**: `/help` reached the model as text and was answered with a description of the TUI's slash commands. stderr carried `BatchLogProcessor.ExportError … network error` lines throughout — grok's own OTLP exporter pushing at a collector that was not running (§7.16a), harmless here. - [x] **Grok, a permission request.** PAID 2026-09-04 at grok 1.0.13 (5e9a58528b76), Windows 11, with the seat's own frames (the transcripts are filed privately, desk/research, one per arm; `vendors/acp.go`'s header carries what each showed). Two findings, and a third the item did not ask for. **The server does not ask by default.** A write session given "run `mkdir zzz` and create probe.txt" raised NO `session/request_permission` in two trials; the wire showed `_x.ai/session_notification{pending_interaction, kind: "permission"}` followed at once by `interaction_resolved`, `_x.ai/sessions/changed` reported `"yolo":true`, and both trials put `zzz` and `probe.txt` on disk. After the prompt `/always-approve off` — handled by the agent as a host turn, no model call — the `write` tool DID raise the request: `toolCall.kind: "edit"`, options `allow-edits-session` (allow_always), `allow-once` (allow_once), `reject-once` (reject_once), cursor's spelling to the letter and the tool's `_meta` naming the `opencode` namespace. `reject-once` ended the turn `stopReason: cancelled` with no `probe.txt`; `allow-once` put it on disk; `mkdir zzz` through `run_terminal_command` asked in NEITHER trial and landed in both. The seat sends no toggle, so its badge now reads `acp · unasked · measured at 1.0.13` and the write detail says why. **The read posture's mode.** `session/new` advertises no `modes` at this build, and `session/set_mode` answers `{}` to every id — `plan`, `agent`, `read`, `bogus-mode-xyz` — so the acceptance is no evidence; only `plan` and `agent` echoed a `current_mode_update` in both runs. Under `set_mode plan` the same write brief ran no shell and wrote nothing in the workspace, two of two: the agent wrote its plan into `~/.grok/sessions//…`, raised the server request `_x.ai/exit_plan_mode{planContent}`, the client's empty result for an unknown request was read as "the user wants to revise the plan", and the turn ended. A read brief under `plan` still ran `list_dir` and `read_file` and answered from the file. So `grokDialect.readModeID` is `plan` since this date and the ACP seat's read badge is `ro:requested`, on exactly the evidence class cursor's is: a mode the model obeys, two trials, never `ro:enforced`. The `acpDialect` bullet above that says "grok: none" was true until this date. **`x.ai/hooks` is not a wire protocol**: it is the agent's on-disk hook system (`~/.grok/hooks`, listed by the agent-handled `/hooks-list`; every tool call produced a `hook_execution` notification naming the operator's own hook). Declaring the block under `clientCapabilities._meta` changed nothing — the initialize response was byte-identical and no hook-shaped request reached the client — so a client cannot hold a posture through it. Also measured: closing stdin ends the process in 2.2–4.4 s (exit 0, twelve of twelve), above the 2 s grace, so the kill lands first; and a `/`-brief is eaten when its name is an advertised command and passed to the model when it is not (`/help`, item above). Not measured: macOS, and whether the seat should send `/always-approve off` itself so the room's card can gate this seat's file writes — that is a design choice, recorded as open in STATE.md rather than made here. - [x] **Grok, cost — HALF PAID 2026-09-04 at grok 1.0.13 (5e9a58528b76).** No `total_cost_usd` anywhere. But the prompt response's `_meta` now carries `usage` with `inputTokens`, `outputTokens`, `cachedReadTokens`, `cacheCreationTokens`, `reasoningTokens`, `modelCalls`, `apiDurationMs`, a per-model `modelUsage` map keyed `grok-4.6-build`, and **`costUsdTicks`** (129,988,800 on a two-call turn). That is a cost figure in a unit nobody has measured: a tick could be 1e-9 or 1e-8 USD and the two readings differ tenfold, so under §4a.1 the seat may render the TOKEN counts and must keep the cost cell absent until the tick is measured against grok.com's own billing page for the same turn. The adapter header's "no cost, and no token usage anywhere" was true at 1.0.4 and is not true at 1.0.13; the header says so now. Owed: the tick's unit. - [x] **Grok, `session/load`.** PAID 2026-09-04 at grok 1.0.13 (5e9a58528b76). `loadSession: true` is advertised on every handshake and honoured: `session/load{sessionId, cwd, mcpServers}` in a NEW process, from the SAME cwd, answered with `models` and `_meta` (no session id, as the adapter assumes) after streaming the whole prior conversation back as `session/update` lines — `user_message_chunk`, `agent_message_chunk`, `tool_call`, `hook_execution` — which is the replay `acp.go`'s guard drops; a prompt after it recalled the earlier turn's files by name. The refusal shape is `-32603 "Path not found."` with `data.code: FS_NOT_FOUND`, and it is the SAME for an unknown id and for a known id loaded from a different cwd: the store is keyed by cwd, so a moved room cannot resume. The process survives the refusal and `session/new` answers in it, which is the branch the cursor capture measured at `-32602`. The `costUsdTicks` unit (the cost item above) is still owed. - [x] **Antigravity, the stream handshake.** PAID 2026-09-05 at agy 1.1.26, Windows 11, the seat's own argv and line shape from a scratch workspace (transcripts filed privately, desk/research). Two `{"event":"user",…}` lines down one stdin: one pid, the same `conversation_id` on both `result` events (`num_turns` 1 then 2), and the second turn answered a question only the first could ("copper"). stdin closed → exit 0 after **0.08 s**, well inside the 3 s grace. No stderr. - [x] **Antigravity, `--conversation` under stream input.** PAID 2026-09-05 at agy 1.1.26: the id from the run above, resumed in a NEW process with `--conversation ` after `--input-format` — the SAME id came back, `num_turns` 3, and the word was recalled. Resumed, not forked, on this path. The unknown-id arm (§9.43's fork, measured on `agy -p` at 1.1.11) was NOT re-run under stream input and stays as `agystream.go` states it. - [x] **Antigravity, `--print-timeout` under stream input.** PAID 2026-09-05 at agy 1.1.26, with a stand-in: `--print-timeout 20s`, turn one answered, **35 s idle**, turn two answered on the same pid. The bound is per turn and idle does not count against it; a 30m bound therefore does not end a seat half an hour into a room. The 30m value itself was not waited out. - [x] **Antigravity, the fallback trigger.** MEASURED 2026-09-05 at agy 1.1.26, and it does NOT produce the shape the retreat keys on. With the seat's batch argv, no `-p` and no `--input-format`, a stream line on stdin was read as the PROMPT: `init` with a fresh `conversation_id`, three `step_update`s, then `result` `status: SUCCESS` with the answer the JSON asked for, exit 0, **zero stderr**, 14.3 s. No first-turn death, no missing session — a silent misparse. So on 1.1.26 the retreat is never triggered by this arm; it remains built for a build that lacks the flag, and this reading confirms only that 1.1.26 is not that build. #### Verification No vendor was started. `go test ./internal/council/vendors -count=1` and the same with `-race` pass over synthesized fixtures: the app-server handshake, both approval methods answered each way in both vocabularies, the interrupt cancelling held approvals first, `Closing()` in order and silent when idle, every fallback trigger reported by `Dead()`, the approval policy per posture; the grok dialect's handshake, its `loadSession` gate against the cursor dialect's measured behaviour, a permission answered by kind and cancelled when the kind is absent, the read posture's own refusal, interrupt and closing order, no cost on the prompt response; the Antigravity stream's argv, envelope, `result` ending the turn, the fork fixture replayed, and its refusals. `internal/council/seatshape_test.go` pins every badge word and the fallback registry. `go vet ./...`, `GOOS=windows GOARCH=amd64 go build ./...` and `GOOS=darwin GOARCH=arm64 go build ./...` are clean. The `help-postures` golden moved by exactly the codex detail's new sentences. ### 9.58 the codex seat failed in both rooms and said only `exit status 1` (2026-09-01) Two live drives on 2026-09-01 and 2026-09-02 showed the same seat in the same state. The hosted read room drew `codex — failed (exit 1)`. The gated write room drew `Codex ✗ failed 12s` with `⚠ exit status 1` under it. Neither room showed a sentence, and neither room showed an answer. The failure was in both postures and in both rooms, so it was not a posture defect and it was not a host defect. #### The measurement The seat was last measured at codex-cli 0.149.1 ([§9.2](#s9-2)'s 2026-08-29 amendment). The machine ran codex-cli 0.151.0. A chip re-ran the seat's exact argv outside council, from a scratch directory, with the prompt on stdin, three times: the read seat's first turn, the write seat's first turn, and the resume shape. Every run ended the same way: ``` {"type":"thread.started","thread_id":"..."} {"type":"turn.started"} {"type":"error","message":"You've hit your usage limit. ... try again at 11:45 PM."} {"type":"turn.failed","error":{"message":"You've hit your usage limit. ... try again at 11:45 PM."}} exit 1, stderr empty ``` Two facts follow from the capture, and they are different facts. **The turn failed because the account had no quota.** That is the vendor's condition and council cannot change it. The three argv shapes parsed and each produced `thread.started`, so no flag moved between 0.149.1 and 0.151.0. The sandbox probes could not run on a turn that never reached a tool, so §9.2's claims stay pinned at 0.149.1 and [STATE.md](../STATE.md) carries the re-measurement as owed. **The room lost the sentence because the vendor moved it.** Through codex-cli 0.147.0 a failed turn wrote its reason to stderr and put zero bytes on stdout, and `codex.go` recorded that as the reason it modelled no error frame: the runner's exit event carried the stderr tail, and that was the whole failure signal. At 0.151.0 the reason rides stdout as a `turn.failed` frame and stderr is empty. The adapter dropped the frame as an unknown type, on the rule that an unlisted type is dropped rather than guessed at, and the exit event arrived with `exit status 1` and an empty tail. Both rooms rendered exactly what they were handed. **Which path this is, after [§9.57](#s9-57).** The drives above ran on a build that seated `codex exec --json`. On the same day this section landed, §9.57 seated `codex app-server` first and kept `codex exec --json` as the measured fallback. The app-server path reports a failed turn on its own `turn/completed` with `status: "failed"`, as a `KindError` that ends the turn, so the sentence reaches the room there already. The adapter change below is the fallback path's, and the exit guard below is what the fallback needed: a spawn-per-turn seat is the one whose failure sentence a process exit can overwrite. #### What changed **`turn.failed` is parsed, and its sentence is the card's note.** The event takes agy's shape (`vendors/agy.go`): a `KindError` with exit code 0, no error and no `EndsTurn`, because the process has not exited. That puts it on the branch [§9.33](#s9-33) built and `failedturn_test.go` pins. The column settles at the vendor's sentence and the exit retires it. **The `error` line is not parsed.** In the capture it always paired with a `turn.failed` that carried the same text. Whether a lone `error` line ends the turn is unmeasured. A column failed on a line that may be a recoverable hiccup would be the room inventing the verdict, and `turn.failed` IS the verdict. **The exit no longer overwrites the sentence, in either room.** A spawn-per-turn failure now produces two events: the vendor's sentence, then the process exit. The second arrived last and replaced the note, in `dispatch.go` and in `councilhost/room.go` alike. Both now hold one rule. A process exit that lands on a column already failed with a note keeps that note, because only the runner's exit event carries `Err`, and the only way a column is failed with a note before its exit lands is a failure the vendor reported in its own stream. In the single-process room the exit's own sentence moves to the note's detail line, so the card reads the vendor's reason as its title and `exit status 1` as its body, and nothing the exit said is dropped. In the hosted room the exit code was already on the head line, so the seat keeps its note and gains `(exit 1)` beside the phase. A bare exit on a column nobody had failed is unchanged: it still names itself. **The failure stays `Unclassified`.** `runner.FailureClass` is grounded in strings captured off a run that positively never reached the conversation. A usage-limit refusal says nothing either way about the thread: the resume shape produced `thread.started` with the requested id and then died the same way. Whether the room may keep a restored thread across that is a ruling under [ADR-008](#adr-008)'s sixteenth amendment, and this section does not make it. #### Verification `go vet ./...`, `go build`, `go test ./internal/council -timeout 20m`, `go test ./internal/councilhost`. The capture is pinned as `vendors/testdata/wire/codex-0.151.0-turn-failed.jsonl`, sanitized to its thread id only, and `TestCodexFailedTurnWireIsPinnedAt_0_151_0` replays it. The adapter's three new tests pin the verdict, the empty-payload verdict, and the dropped `error` line. The dispatch test feeds the real adapter's event and then the runner's exit shape, and asserts the note, the detail, the settle and the retirement. The hosted-room test does the same over `Room.Apply`. **Not verified here: a live turn that answers.** Every turn the chip ran died at the usage limit, so the fix was checked against the captured failure and never against a codex turn that succeeded at 0.151.0. The 0.147.0 fixture still replays green, which is the evidence that the success path did not move in the parser. The first live turn after the account has quota is the check, and it is the operator's. ### 9.59 the agy racer wrote its attempt into the vendor's own scratch directory (2026-09-03) A recorded race (`2026-09-03-rebuttal.jsonl`, turn 9) asked four seats to add `haiku.md`. Claude, Codex and Grok wrote the file in their worktrees. The Antigravity column said `no changes against 5664d51` and ranked 3rd of 4. Its trace showed why: `write_to_file` on `~\.gemini\antigravity-cli\scratch\sailboat-telltales\haiku.md`, then on `~\.gemini\antigravity-cli\scratch\haiku.md`. The seat also listed its own scratch directory and read a task log under its own `brain\` folder. The rank was honest. The seat never touched its worktree. #### The measurement The adapter set `Spec.Dir` to the worktree, and the vendor's `init` line reported that cwd. The transcript agy stored for the turn holds the cause, in its own system prompt: *"The user does not have any active workspace. If the user's request involves creating a new project, you should create a reasonable subdirectory inside the default project directory at C:\Users\sanle\.gemini\antigravity-cli\scratch."* Every past headless conversation on this box that had no `--add-dir` carried the same sentence, back to the first arena on 2026-08-08. The conversations that did name an active workspace were the interactive ones, opened in `~`. Eight probes at agy 1.1.25, each one `agy --output-format stream-json --disable-slash-commands --print-timeout 3m … -p ""` from the named directory: | cwd | extra flags | where `probe.md` landed | turn | | --- | --- | --- | --- | | arena worktree (`.git` is a file) | none | `~\.gemini\antigravity-cli\scratch` | SUCCESS | | the plain scratch repository (`.git` is a directory) | none | scratch | SUCCESS | | a directory with no `.git` | none | scratch | SUCCESS | | `code\telltale`, an exact `trustedWorkspaces` entry | none | scratch | SUCCESS | | arena worktree | `--mode accept-edits` | scratch | SUCCESS | | seat worktree | `--add-dir ` | **nowhere**: `write_to_file` reported DONE, no file, stderr `a tool required the "write_file" permission that headless mode cannot prompt for, so it was auto-denied` | CANCELED, empty response | | seat worktree | `--add-dir --mode accept-edits` | **the worktree** | SUCCESS | | arena worktree | `--add-dir --mode accept-edits --input-format stream-json`, one user line on stdin | **the worktree** | SUCCESS | A ninth run resumed the CANCELED conversation with `--conversation --add-dir --mode accept-edits` and asked for a second line. It edited the file in the worktree, echoed the same id, and reported `num_turns: 2`. So two facts, and each one needs its own flag. The cwd is no workspace to this vendor. Neither git shape nor the trust list changes that, and `--add-dir` is the flag that names one. A write inside a named workspace needs a permission that print mode cannot ask for, and `--mode accept-edits` is the vendor's edit-only grant. A write into the vendor's own scratch directory never needed the grant, which is why the defect looked like a wrong directory and not like a denied write. #### The fix `baseArgs` in `vendors/agy.go` takes the workspace and the posture. It passes `--add-dir ` in every posture, and `--mode accept-edits` in the write postures. Both precede `-p`, by the adapter's standing rule. The stream-json session builds from the same function. `PostureWriteGated` takes the write argv: this seat has no gate channel, and the posture contract says a gated seat may do anything write mode allows. The read posture keeps the workspace named and the edits unaccepted. That is the closest thing to a read posture this vendor has offered: the one in-workspace write measured under it was denied. **The badge does not move.** A single denied write is not a sandbox. `run_command` still runs under the operator's own allow rules in `settings.json`, and a write outside the named workspace still lands, so the column stays `unsandboxed` and its detail now says which two flags council passes and why. `--dangerously-skip-permissions` stays refused: the edit grant is not the approve-everything class ADR-008's fifth and seventh amendments refuse. One side finding, recorded in [§9.37](#s9-37)'s AGENTS.md table: with a workspace named, the seat ran `list_dir` and then `view_file AGENTS.md` before it wrote, on a brief that never named the file, and then wrote the file the brief in AGENTS.md asked for. That is the Claude row's fact. The codename probe was not run. #### Verification `go vet ./...`, `go build ./cmd/telltale`, `go test ./...`. `TestAgyNamesItsWorkspaceOnArgv` pins `--add-dir ` ahead of `-p` on the first turn and the resume, and pins that an empty workspace sends nothing. `TestAgyAcceptsEditsOnlyWhenTheRoomWrites` pins the grant on both write postures, on both paths, and its absence on read. The two posture tests now pin that the grant is the WHOLE difference between the postures. `TestAgyStreamSessionKeepsEveryFlagAndNoPrompt` pins both flags on the stream session. **Verified live by the operator, 2026-09-04, race t13.** The section shipped saying the council TUI takes no scripted input, so the racer's exact argv had been run by hand instead (rows seven and eight above) and the live race was left as the operator's check. That check has now run: four seats, the same `haiku.md` brief, in `~\Desktop\telltale-rooms\scratch`. The Antigravity column reported `4th of 4 · done · 8m19s`, `committed 3c9df8c`, and a stat of `haiku.md | 3 +++`. Its trace showed `write_to_file` and `view_file` on `scratch-arena-t13-agy\haiku.md`, then four `run_command` steps inspecting the tree with `git status` and `git diff`. On disk afterwards, `3c9df8c` on `arena/t13/agy` carries `haiku.md` and three insertions, and the file holds the haiku. Agy's own scratch directory received nothing: the `haiku.md` sitting there is still the turn-9 one, dated 2026-09-03. **The rank is the honest part.** This seat placed 3rd of 4 before the fix, on a worktree it had never touched, and 4th of 4 after it, on work it actually did — the slowest of the four at 8m19s, where the other three finished in 29s, 36s and 1m41s. The fix did not make the column win. It made the column true.