# telltale — design doc
Status: v1 is cut — **v0.2.0, released 2026-08-14**, the snapshot gates held (§1). The honest-gauge
rule requires every segment's data source to be named here before that segment ships; the
tables below are the authority the eval harness tests against.
## Citing this document
Code comments, PR bodies and the other docs cite this document by section number. Every
numbered section therefore carries a stable HTML anchor, so a citation can be a link.
**The anchor comes from the section number alone.** Write `s`, then the number, and
replace each `.` with `-`. So §9 is `#s9`, §4a.1 is `#s4a-1`, and §7.16a is `#s7-16a`.
```
[§7.16a](docs/design.md#s7-16a) from another file
[§7.16a](#s7-16a) from inside this document
```
The anchor is keyed on the number and never on the heading text. GitHub derives its own
anchor from the heading text, and the headings here carry dated prose that gets amended,
so a text-derived anchor breaks on every rewording. `#s7-16a` survives it. Add the
matching anchor line when you add a numbered section.
## ADR index
The decision records themselves are archived outside this repository, because the
2026-08-05 ruling stopped adding them here. The ADR numbers stay in the prose, so this
table says what each one decided and points at the section that carries the substance.
Each row is also an anchor target: ADR-002 is `#adr-002`.
| ADR | What it decided | Where the substance is |
|---|---|---|
| ADR-001 | The honest-gauge rule. A displayed value must come from measured vendor output, and an inferred value is omitted or marked as an estimate. | [§4a.1](#s4a-1), [§3.4](#s3-4) |
| ADR-002 | Go with Bubble Tea and Lipgloss, one binary, two modes. Windows is the primary target, and the statusline path initializes no TUI framework. | [§6](#s6), [§5](#s5) |
| ADR-003 | Gemini CLI gets a built-in adapter. Its verification hold is released, and an addendum records the consumer-tier withdrawal. | [§3.7](#s3-7) |
| ADR-004 | Antigravity CLI gets a statusline, routed on the payload's documented `product` field. | [§2.1](#s2-1) |
| ADR-005 | External adoption is an explicit product goal, so the roadmap carries an adoption track. | [§8](#s8) |
| ADR-006 | Antigravity also gets a HUD adapter. A re-survey found the transcript that the first verdict recorded as absent. | [§3.8](#s3-8) |
| ADR-007 | Cursor gets a built-in HUD adapter, because its seam is on disk rather than behind a CLI. | [§3.9](#s3-9) |
| ADR-008 | `telltale council` is the dispatch room. It spawns vendor CLIs, so it sits off the gauge data layer. | [§9](#s9) |
ADR-006 is the one number above that the current text never prints; §3.8 carries its
account under the re-survey heading instead.
**ADR-010 and ADR-012 in this repository are not telltale's.** They belong to the
`agent-ops` decision series, and the prose always names that series beside them. The two
series number independently, so telltale's ADR-007 is the Cursor HUD adapter while
`agent-ops` ADR-007 is a different decision.
## 1. Product shape
**`telltale council` is the product. The gauges — `telltale statusline` and `telltale
hud` — are the infrastructure under it.** That ranking had been stated out loud and
recorded nowhere, which meant it bound nothing and every argument about what to build
next started over from memory. It is written here so it stops depending on who was in the
room. It does not demote the gauges: they are where each vendor's on-disk seam was
surveyed and written down (§3), they are what the honest-gauge rule was built and tested
against (§5), and council inherits both — it renders through the same `internal/model`
vocabulary and `internal/theme` palette the two gauge paths share. They are finished, they
are load-bearing, and they are not the thing this is for.
**v1 is a snapshot, not a freeze — it cuts when three gates hold, not when the room goes
quiet (re-cut 2026-08-08, owner's ruling).** The original hold said "until council
settles," and the first attempt to operationalize that — five consecutive days with no
merged PR touching council's visible surface — was falsified within a day: this project is
driven daily by its owner, so a quietness clock measures abandonment, not stability. What
the hold was actually protecting is narrower and checkable: a stranger who reads the
launch post and installs must find the room the post described. So v1 cuts when, on the
day of the tag:
1. **Nothing on the surface is half-finished or landed-but-never-driven** — every
recently re-founded piece (a seat's protocol, a new command) has been used by the
owner for real work for a few days;
2. **The README is verified against tip** — every claim, keybinding and badge checked
against the code, with the frame-freshness tests holding the renders;
3. **No breaking change to the routing grammar, room commands or keymap is planned** —
churn after the tag is welcome; a *known upcoming* contract break is not.
The owner's dogfood bar (two weeks of daily use, clock from 2026-08-01) still applies and
closes no earlier than 2026-08-15. Development never pauses for any of this: work merged
after the tag becomes the next minor version, and the tag itself is one command (§8).
The standing alternative — cut v1 as gauges only, statusline and HUD with declared vendor
version pins — remains rejected, because a v1 that named the gauges would name the wrong
product.
Two gauge surfaces over one data layer:
```
vendor adapters ──► normalized session model ──► renderers
(claude, codex) (one schema, documented) (statusline / HUD)
```
One Go module, one binary (`telltale.exe`). ADR-002 specified two modes; council (ADR-008)
is the third, and it does not sit on the pipeline above:
- **`telltale statusline`** (Claude Code and, since ADR-004, Antigravity CLI — routed
on the payload's documented `product` field, §2.1): reads the vendor's JSON on
stdin, prints one line, exits. **Bubble Tea is never initialized on this path** —
a convention until 2026-08-16, gated since by `TestFastPathNeverReachesTUIFramework`
(§5). Latency budget: single-digit milliseconds of telltale's OWN work, and that
much is measured — parse+render is 14 µs (`BenchmarkRender`). The end-to-end cost
the operator actually pays is larger and is process start, not this code: ~25 ms
median per invocation on the reference workstation. Budget-conscious output (every
character renders on every prompt).
- **`telltale hud`** (cross-vendor): a Bubble Tea/Lipgloss watch-mode TUI listing live
sessions across vendors with per-session gauges. **First-class UI surface** — a UI
design section (layout grid, color/threshold system, motion rules, empty/degraded
state designs) is written here BEFORE the HUD is built, and degraded-state renders
are eval fixtures. Windows Terminal is the reference rendering environment.
- **`telltale council`** (ADR-008, §9): the dispatch room — one brief typed once,
answered by the seated vendor CLIs side by side, each column claiming only what was
measured about that vendor. It spawns vendor CLIs instead of reading their session
files, which is why it is off the data layer above and specified separately in §9.
- **`telltale hook `** (§7.16): the vendor-hook relay — a per-turn payload on
stdin, token counts to `~/.telltale/usage/`, and **nothing on stdout**, because a
hook's stdout is parsed by the vendor as a hook result. Not a gauge and not a room; it
renders nothing and is never run by a human.
**The gauges never write, with three bounded exceptions — all under `~/.telltale/`, all
numbers and keys only, never content.** `telltale council` keeps `council/room.json` (the
session ids reattaching needs); the statusline relays `quota/.json` after its line
is on stdout (§7.15); and `usage/.json` accumulates per-turn token counts from two
writers, `telltale hook` (§7.16) and the `telltale otel` collector (§7.16a). Each store
is atomic (temp+rename), best-effort, self-expiring on read, and pinned by a test that
walks the serialized form field by field. No transcript, prompt, reply, path or address
reaches any of the three. Anything else that wants to write from `internal/hud` or
`internal/statusline` is in the wrong package.
**Amended 2026-08-11: a FOURTH store exists, it carries content, and the paragraph above
was never corrected for it.** The event sink (`telltale events`, §7.21) writes each hook
payload VERBATIM under `~/.telltale/events/`. It does not widen the rule above, because
what contains it is scope rather than redaction: it is its own foreground mode the
operator starts, its server binds loopback only and refuses any other host, and nothing
in the gauges reads or renders those files. So the three counted above stay
numbers-and-keys, and the fourth is named as an exception instead of being folded into
them. §7.21 carries the record and CLAUDE.md's boundary section carries the same
exception.
**Amended 2026-09-04: the numbers-and-keys list is FOUR, not three.** `telltale probe`
(LEDGER, 2026-09-04) writes `probe/.json` per seat: the vendor id, the version
string that binary printed, the day, the telltale build that probed, and one result plus
a millisecond count for each of its three checks. It is the strictest of the four,
because its writer DRIVES an agent: the brief, the reply, the session id and the
workspace path are all within reach of it and none is written, and neither is the failure
reason, because a vendor's own error line carries paths. It is not a gauge, so the
paragraph above is unchanged in what it says about `internal/hud` and
`internal/statusline`. SECURITY.md and CLAUDE.md carry the same four; a count in prose
goes stale with nothing to catch it, so `internal/history/boundary_test.go` enumerates
the writer packages by name and states none.
The two gauge paths share exactly two packages: `internal/model` (the schema) and
`internal/theme` (thresholds and value formatters). `internal/theme` is stdlib-only and
holds no style type, which is what lets both surfaces share the numbers while the
statusline links no TUI framework.
## 2. Statusline segments (v1)
| Segment | Source (exact field) | Empty/degraded state | Status |
|---|---|---|---|
| Model | stdin `model.display_name` (falls back to `model.id`) | hide if both empty | **built** |
| Context % | stdin `context_window.used_percentage` (input-token based per docs) | hide segment | **built** |
| Cache hit % | stdin `prompt_cache.hit_ratio` (vendor-COMPUTED over the session's main-conversation requests; ×100 is a unit conversion), gated on `prompt_cache.caching_observed` | absent block, null ratio, or `caching_observed` false all hide; 0 with caching observed is a reading and renders `cache 0%` | **built** (§7.16c) |
| Session cost | stdin `cost.total_cost_usd` | hide segment | **built** |
| Quota pacing (5h) | stdin `rate_limits.five_hour.used_percentage` + `resets_at` (unix s) | rate_limits absent on API-key logins; each window independently absent → hide, never zero; countdown hides without `resets_at` | **built** |
| Quota pacing (7d) | stdin `rate_limits.seven_day.*` | same rule | **built** |
| Worktree | stdin `worktree.name` (present only in `--worktree` sessions) | hide segment | **built** |
| Folder | stdin `workspace.current_dir` (fallback `cwd`), basename only — no filesystem/git calls | hide segment | **built** |
Deliberately not shown **on this path**: git branch (would require an exec; the
statusline path reads nothing beyond stdin — revisit only with a measured budget),
permission mode (not in the stdin payload; same call the predecessor script made).
The path's one write is the quota relay (§7.15): after the line is on stdout, the
payload's rate-limit windows go to `~/.telltale/quota/` for the HUD — numbers only,
best-effort, never ahead of the render.
> Both exclusions are properties of the **stdin seam**, not of Claude Code. The
> transcript carries `gitBranch` and `permissionMode` directly, so the HUD's disk path
> gets both for free (§3.1). The statusline's exclusions are unaffected.
Threshold colors (applies to any percentage segment): green < 60, yellow ≥ 60, red ≥ 85,
from `theme.WarnPct` / `theme.CritPct`. `NO_COLOR` env strips styling. Derived displays
(reset countdown `↻2h13m`) are arithmetic on `resets_at` only.
Schema verification record: full stdin JSON schema captured from
code.claude.com/docs/en/statusline on 2026-08-01, including per-field absence semantics
(`rate_limits` Pro/Max-only and only after first API response; each window independently
absent). Statusline updates are debounced at 300ms and in-flight scripts are cancelled —
which is the empirical backing for the fast-exit budget. Parsing ignores unknown fields
by design (vendor adds fields between versions).
**Re-measured 2026-08-16 at CLI 2.1.233 by source read (§7.16b), and the payload had grown.**
`context_window` now also carries `total_input_tokens`, `total_output_tokens` and a
`current_usage` object, and the payload carries an undocumented `prompt_id`. The first three
are modelled and **render nothing and relay nothing** — they describe one API call, and
`total_input_tokens` is the window's occupancy rather than a spend, so no total may be built
from them. That is the whole of §7.16b, and the tolerance of unknown fields above is why
every release between 2.1.90 and 2.1.233 was unaffected by the growth. The segment table
above is unchanged: no new segment was added.
**Re-measured 2026-09-04 at CLI 2.1.260 by source read (§7.16c), and this time the growth
DID earn a segment.** The payload gained `prompt_cache` at 2.1.251 — the vendor's own
prompt-cache statistics for the main conversation — and one of its fields is a number the
vendor computes rather than a count telltale would have to combine. The segment table above
carries the new row; §7.16c carries the measurement and the reason the transcript adapter
was left alone.
**Known divergence from `theme`, deliberate and unresolved:** the statusline's local
`pct` uses `%.1f`, which rounds, and its local `shortDur` has no days branch. The shared
helpers `theme.Percent` and `theme.Countdown` floor and carry days respectively, and the
HUD uses them. Unifying the statusline onto them changes its rendered output (99.96 would
stop reading as `100.0%`, a 7-day window would stop reading as `↻120h00m`), which is a
behaviour change to a shipped surface and is therefore a separate change with its own
fixture updates. The thresholds are already unified; the formatters are not.
### 2.1 Antigravity CLI statusline (added 2026-08-02, ADR-004)
`telltale statusline` serves a second vendor: Antigravity CLI (`agy`) hands statusline
commands a JSON payload on stdin, same seam shape as Claude Code. **Routing is the
documented `product` field** — agy stamps `"product": "antigravity"` on every payload
(observed on all six live captures) and Claude's payload has no product field; one
binary, one subcommand, no flag. Stdin is read once and handed to whichever parser the
marker selects (`internal/antigravity`).
| Segment | Source (exact field) | Empty/degraded state | Status |
|---|---|---|---|
| Model | stdin `model.display_name` (falls back to `model.id`) | hide if both empty | **built** |
| Context % | stdin `context_window.used_percentage` (vendor-reported; the payload also carries `context_window_size`) | hide segment; 0 is a reading and renders `ctx 0%` | **built** |
| Quota buckets | stdin `quota..remaining_fraction` + `reset_in_seconds` (fallback `reset_time`) — one segment per NAMED bucket, ids rendered verbatim, sorted for stability; used% = (1−remaining)×100, a unit conversion | bucket without `remaining_fraction` hides; absent map hides all | **built** |
| Agent state | stdin `agent_state` — the first vendor-REPORTED liveness signal on any seam; `tool_confirmation_pending: true` outranks it and renders `confirm?` | hide if empty; unknown vocabulary renders verbatim in dim | **built** |
| Branch | stdin `vcs.branch` (+`*` when `vcs.dirty`) — in the payload, so no exec; the no-I/O-beyond-stdin rule holds | hide segment | **built** (documented; not yet observed live — §3.8) |
| Folder | stdin `workspace.current_dir` (fallback `cwd`), basename only | hide segment | **built** |
Not rendered, deliberately: `cost` does not exist anywhere in the payload (nothing is
priced); `email` and `plan_tier` are identity, not gauges; `transcript_path` is not
displayed because it is unverified, not because it is absent — the transcript IS
written on disk (§3.8's 1.1.13 re-read found it in 81 of 81 conversations; the claim
that agy "never writes that file" was false and is corrected there), but whether the
PAYLOAD's value points at that real file needs a live capture nobody has run, and
displaying an unverified path would be narrating.
**Amended 2026-08-17 — `transcript_path` is no longer unverified. It is verified WRONG.**
The live capture that paragraph asks for was taken (§3.8's re-capture block). The payload's
path drops the `-cli` segment: it names `~\.gemini\antigravity\brain\\…`, while the real
transcript for that same session sits only under `~\.gemini\antigravity-cli\brain\\…`.
The advertised directory exists and is EMPTY, so a reader that trusted the value would open
nothing and could not tell a missing file from a missing session. The refusal above stands
unchanged and its GROUND changes: it was caution about an unverified value, and it is now a
measurement of a false one. No code ever followed the path, which was re-checked across the
repository rather than assumed.
Schema verification record: documented contract (antigravity.google/docs/cli/statusline)
cross-checked against a six-payload live capture from a real interactive session on agy
1.1.9, 2026-08-02 (§3.8). **Re-captured 2026-08-17 at agy 1.1.13 — fifteen payloads, one
live interactive turn** (§3.8's re-capture block): quota confirmed at FOUR named buckets
(`3p-5h`, `3p-weekly`, `gemini-5h`, `gemini-weekly`), each carrying `remaining_fraction`,
`reset_time` and `reset_in_seconds`; `agent_state` observed live, including one value the
documented vocabulary omits; `context_window` grown to the §7.16b shape. `vcs` is still the
one documented segment nobody has observed — both capture sessions ran outside a git repo.
Fixtures are synthesized to the observed shapes.
### 2.2 Cursor CLI statusline (added 2026-08-16)
`telltale statusline` serves a third vendor. `cursor-agent` reads a top-level `statusLine`
object from `~/.cursor/cli-config.json` and hands the command a JSON payload on stdin —
the same seam shape as the other two. Measured at **cursor-agent 2026.08.04-aaa8809** on
Windows 11, live capture first and source read second; §7.16's dated amendment carries the
per-surface measurement and `internal/cursorstatus`'s package doc carries the shapes.
```json
{"statusLine":{"type":"command","command":"telltale statusline --vendor cursor",
"padding":0,"updateIntervalMs":300,"timeoutMs":2000}}
```
**Routing is an explicit flag, and that is a finding rather than a shortcut.** This
payload carries no `product` field, no `hook_event_name`, and no other vendor name —
`version` holds the CLI's own build string, which is a value and not a marker. It is
Claude-shaped on purpose: the vendor's bundled `statusline` skill says the spec "is
aligned with Claude Code's status line", and the two overlap on `session_id`,
`transcript_path`, `cwd`, `model.*`, `workspace.*`, `version` and `output_style.name`.
So §2.1's affirmative-marker scheme cannot be extended to it, and guessing from structure
(`render_width_chars` present, `cost` absent) would be a heuristic over a payload the
vendor may grow at any release — with a silent failure mode, since a misrouted Claude
payload renders a plausible line with its quota missing. `--vendor cursor` is written once
into the config above and wins over the marker probe.
| Segment | Source (exact field) | Empty/degraded state | Status |
|---|---|---|---|
| Model | stdin `model.display_name` (falls back to `model.id`) | hide if both empty | **built** |
| Context % | stdin `context_window.used_percentage` — the ONE context number this vendor sources rather than computes | hide segment; a session before its first API call sends every `context_window` key as `null`, which is an unread field and not a zero | **built** |
| Autorun | stdin `autorun` — no counterpart in either other vendor's payload; renders only when `true` | `false` and absent both hide (see below) | **built** |
| Worktree | stdin `worktree.name`, same `⌥` mark as the Claude path | hide segment | **built** (documented; not observed live) |
| Folder | stdin `workspace.current_dir` (fallback `cwd`), basename only | hide segment | **built** |
**Two fields are REFUSED, and they are the reason this section is careful.**
`context_window` arrives with six keys; two of them are computed by the CLI from a third
and named as though they were read. From the bundle's own payload builder in
`./src/ui.tsx` at the pinned build, where `Ve` is the usage reading and `Ke` the window
size:
```
m = null!=Ve?Ve:null, // used_percentage
f = null!=m ? Math.max(0,Math.round(10*(100-m))/10) : null, // remaining_percentage
v = null!=Ke&&null!=m ? Math.round(m/100*Ke) : null // total_input_tokens
```
The vendor documents it too — its `statusline` skill calls `total_input_tokens`
"Estimated input tokens (derived from used_percentage)". A token count that is really a
rounded percentage, under a name that reads like a meter, is the ADR-001 violation print
mode's `inputTokens` already cost this repo once (§7.16). **They are absent from
`internal/cursorstatus`'s structs rather than parsed-and-ignored**, on the
`internal/cursorhook` rule: the struct is the allowlist, `encoding/json` drops every field
with no destination, and a field that does not exist cannot be reached by a later change
that did not read the comment. `TestCursorDerivedFieldsNeverRender` feeds a fixture
carrying both, populated and plausible, and asserts neither number reaches the line.
Also not rendered: `context_window_size` (genuinely vendor-reported, but nothing draws it
— add it back with the segment that wants it) and `current_usage` (observed only as
`null`, so its populated shape is unmeasured and declaring one would be inventing a
schema). There is **no quota, no rate limit and no cost anywhere in this payload**, so
this path writes no quota relay at all (§7.15) — the absence is the honest answer, not a
gap.
**The autorun asymmetry is deliberate and is not the zero-vs-absent rule being bent.**
That rule governs gauge READINGS, where 0% and "no source" are two facts a user must tell
apart. `autorun` is a posture flag with a default: `false` is the ordinary state of every
session, and a segment that says "nothing unusual" on every line teaches the reader to
stop looking. Off and absent both render nothing; on renders a word, in yellow, because
it is the state where the agent may run a command without asking.
**Budget.** The vendor's `timeoutMs` defaults to 2000 with a floor of 50, and
`updateIntervalMs` is clamped to >= 300. The 2000ms is its kill deadline, not an
allowance: the binary is respawned on every debounced update, so ADR-002's
single-digit-millisecond target for telltale's own work is unchanged, and
`BenchmarkRenderCursor` measures it at 14 µs against that 300 ms floor. Two honest
limits on that sentence, added 2026-08-16: a benchmark is not a gate — CI runs no
`-bench`, so `BenchmarkRenderCursor` fails nothing — and parse+render is not what the
respawn costs. The respawn's end-to-end price is ~25 ms median, essentially all of it
process start. §5's amendment says which half CI now holds.
Schema verification record: two live payloads captured 2026-08-16 from a real interactive
session (the shape with every `context_window` key null — that session had made no API
call), cross-checked against the vendor's bundled `statusline` skill and a source read of
`./src/hooks/use-status-line.ts` and `./src/ui.tsx` at 2026.08.04-aaa8809. Fixtures are
synthesized to the observed shapes; the populated-context fixture is synthesized to the
documented one, because no captured payload ever carried a number there. **That gap is
CLOSED, 2026-08-17**: a live interactive session at cursor-agent `2026.08.11-e8db854`
rendered `ctx 12.7%` after its first reply, so `used_percentage` is observed populated at
one-decimal precision. §7.16's amendment carries the capture and the build caveat. The
fixture's assumed shape was correct and did not move.
## 3. HUD (v1)
One row per live session, both vendors; per-row: vendor, session identity, model,
context/quota gauges **where the vendor provides them**, last-activity age. The rendered
grid, its responsive tiers and every degraded state are specified in §7.
### 3.1 Claude Code adapter sources — VERIFIED LIVE 2026-08-01, Claude Code 2.1.219; RE-MEASURED 2026-08-16, Claude Code 2.1.233
Read-only survey of `%USERPROFILE%\.claude\` on the dev PC: 33 project dirs, 837
sessions, 13,211 records walked. Nothing from that survey is reproduced here or in the
fixtures; the fixtures are synthesized to shape only.
**Discovery glob — `~/.claude/projects/*/*.jsonl`, non-recursive, UUID basename.**
Measured: non-recursive = 837 files, recursive = 2021. The extra 1,184 are subagent
transcripts under `/subagents/**`, plus `tool-results/` and `workflows/`
sidecars; a `**/*.jsonl` glob inflates the session list 2.4x and double-counts every
token. Non-`.jsonl` neighbours share the directory (`.memory-sync-manifest.json`), so
the basename is checked as a UUID, not just the extension.
The project-directory slug is **lossy and must never be decoded to a path**: `\` and a
literal `-` both encode as `-` (`C--Users-dev-code-my-app` could be `code\my-app` or
`code-my\app`), and the drive-letter case is not stable — the same tree can hold
`C--Users-…-app` and `c--Users-…-app` as siblings. `cwd` is read from the record; the slug is an
opaque grouping key. Because those sibling directories can hold the same session id,
`Discover` also de-duplicates by id, newest mtime winning: a duplicate id would break
the HUD's row matching. The tree also mutates during a sweep (a project dir vanished
between enumeration and open during the survey), so `Discover` swallows ENOENT on dirs
and files and continues rather than aborting.
**Record fields the adapter reads** (all on `assistant` records unless noted):
| Normalized field | Exact source | Absent when |
|---|---|---|
| Session id | `sessionId` (present on every record type) | never |
| Working dir | `cwd` | metadata-only record types |
| Git branch | `gitBranch` (carried as a display-only extra) | outside a repo |
| Model | `message.model` | non-`assistant` records; `""` is rejected |
| Tokens in context | `message.usage.input_tokens + .cache_read_input_tokens + .cache_creation_input_tokens` | no `usage` |
| Last activity | file mtime | never (but see clock skew below) |
| CLI version | `version` (display-only extra) | never |
| Title | `custom-title.customTitle`, else an `ai-title` record | untitled sessions |
| **Sub-agent count** (v1.1) | **stat only:** entries matching `.jsonl` under `/subagents/`, mtime within 15 min | never — an absent directory is a measured **zero**; only an unreadable one is absent |
**The sub-agent count is `CapDerived`, and the reason is worth stating precisely.** The
files are counted *exactly* — one `ReadDir`, one `Info` per entry, no file opened and no
byte parsed, which is what makes it affordable on the 1 s poll. What is inferred is the
15-minute recency boundary that turns "written lately" into "a fan-out is running now".
That inference is the thing the estimate marker exists to expose, so the chip renders
`⑂~2` rather than `⑂2` (§7.13). The boundary is `model.DefaultLivenessThresholds.Idle`
rather than a second constant: the chip sits on a row whose state dot already classifies
"recent" at that boundary, and two definitions of recent on one line is how a display
starts contradicting itself.
Two absences are distinguished, per §4a.1: the directory **not existing** means the
session never fanned out, which is a countable zero; the directory existing and the OS
**refusing** is nil plus a diagnostic, because we do not know. A sub-agent transcript
whose mtime is ahead of the local clock is not counted at all — the same rule the
session's own mtime gets, for the same reason: a timestamp ahead of the clock is not a
readable time, so it cannot be evidence of recency.
**Amended 2026-08-17 — `subagentStatusLine` was evaluated as a replacement for that inference,
and refused.** The estimate marker exists because of the 15-minute boundary, so a vendor surface
that *reports* which sub-agents run would retire `CapDerived` honestly. Claude Code 2.1.233 ships
one. It is a settings key with the same `{type:"command", command}` shape as `statusLine`, and the
bundle's own schema describes it as a "Custom per-subagent status line shown in the agent panel;
receives row context as JSON on stdin".
Measured by a source read of the shipped bundle at **Claude Code 2.1.233** — the instrument
§7.16b used, labelled here for the same reason. The payload is the shared session-basics block
(the `py` helper §7.16b identified) plus `columns` and a `tasks` array. Each `tasks` entry carries
`id`, `name`, `type`, `status`, `description`, `label`, `startTime`, `model` and `effort`. The
command runs with a 5 s timeout, and its stdout is parsed as JSON lines of `{id, content}`. A
React effect drives it: a 300 ms debounce on change, then a repeating 5 s tick while any row is
live, with an overlap guard.
**The count it carries is not the count this adapter defines, and that is the refusal.** Three
measured reasons, in order of weight:
- **The population is wider.** `tasks` holds panel rows, and the observed `type` values include
`local_agent`, `remote_agent`, `in_process_teammate` and `local_workflow`. §3.1 counts
transcripts in a session's `subagents/` sidecar. `len(tasks)` answers a different question
under this field's name.
- **The list keeps rows that finished.** The tick filters on `evictAfter !== 0`, so a row
survives its own completion until eviction. A length taken from it overstates a fan-out in
progress. The 15-minute boundary makes the same error, but it declares itself an estimate.
- **It is the interactive UI's path, and the HUD reads disk.** The effect is a React hook over
the agent panel and the terminal's column count, and print mode mounts no panel. To source it,
telltale needs a new relay mode plus operator wiring in `settings.json`. It would then cover
only the sessions where both are true, while `countSubagents` covers every session on disk
today.
So `FieldSubagents` stays `CapDerived` and `countSubagents` is unchanged. The payload's `status`
field is the one genuinely stronger signal here, because a reported status beats an inferred
recency window. It is reachable only through that relay, so this ruling is re-openable against
that build rather than closed.
**One arm is owed, and it is named rather than dropped.** A live payload capture was attempted
and blocked. To exercise `subagentStatusLine`, the key must sit in a `settings.json`, and this
machine's credential guard default-denies writes to that path — including the throwaway,
project-local copy the probe wanted. The block was accepted rather than worked around, so the
shape above rests on a source read with no live capture behind it. That is weaker than §7.16b
ended up: it closed its own source read with a capture on the same day.
Only `assistant` and `user` records carry `message`. `custom-title`, `last-prompt`,
`mode` and `ai-title` carry `{type, sessionId, }` and have **no `timestamp`
and no `cwd`** — a parser that assumes those fields exist will nil-deref. Full observed
`type` set: `assistant, user, attachment, last-prompt, queue-operation, custom-title,
pr-link, mode, system, ai-title, permission-mode, file-history-snapshot`.
The survey verified the **shape** of an `ai-title` record but not the name of its payload
key. The adapter therefore matches on the verified structure — exactly one key beyond
`type` and `sessionId`, holding a string — rather than guessing a field name from memory.
A record that does not match yields no title and the row falls back to its workspace
name, which is absence rather than a wrong label.
**Claude capability gaps (grepped, zero matches across the corpus):**
- **No cost.** `cost.total_cost_usd` exists only on the statusline stdin payload.
- **No quota.** `rate_limits.*` likewise stdin-only.
- **No context window size**, so **context % is not derivable** — the denominator varies
by model and by the `[1m]` variant. Token counts are sourced and carried as a
display-only extra; the CONTEXT cell for a Claude row is absent (§6 Q7).
**Honest-gauge traps pinned by fixtures:**
- `input_tokens` alone is not context usage. Measured live: `input_tokens=2,
cache_read=213388, cache_creation=2464`. Reading `input_tokens` renders **2 tokens**
for a ~216k-token context.
- `message.model` can be `""` (locally generated notices, zeroed usage). It
must never reach the model cell and its zeros must not overwrite a real reading.
- An mtime ahead of the local clock has **no readable age**. It is left nil and marked
degraded rather than clamped, because "0s" claims the session was active this instant.
The HUD renders `—`.
**Liveness — and why the PID registry is not read in v1.** The honest primitive is
**mtime = last activity**, rendered as an age. The survey found an undocumented registry
at `~/.claude/sessions/.json`
(`{pid, sessionId, cwd, startedAt (unix ms), version, kind, entrypoint, name,
nameSource}`); at survey time all 7 entries mapped to live PIDs and to top-level
transcripts. **The adapter does not read it.** Every use of it reduces to "a process
with this id exists", which §4a.4 names explicitly as evidence a process exists rather
than evidence the session is doing anything — and a liveness hint is the one value
`Validate` cannot check, so the bar for emitting one is a signal that actually separates
working-now from process-exists. `liveness` is therefore `CapNone` for Claude and the
HUD classifies every vendor identically from `last_activity`. The registry stays recorded
here as a verified observation on 2.1.219 (not a vendor contract) in case a later version
adds a turn-start/turn-end signal worth reading.
**Read strategy.** Transcripts routinely reach 7.7 MB. The adapter reads a bounded
**head** (64 KiB — session id / cwd / git branch / title, verified present on the first
record of 60/60 files sampled) plus a bounded **tail** (256 KiB, first fragment discarded
as partial), scanning for the newest `assistant` record with `message.usage` and a
non-synthetic model. A file smaller than the tail window is read once, not twice.
Records with `isSidechain == true` are skipped defensively — 0 of 837 top-level
transcripts contain one on 2.1.219, but the filter is free.
**RE-MEASURED 2026-08-16 — Claude Code 2.1.233.** `telltale doctor`'s drift notice (§9.42)
reported this survey stale on its first live run, and this is the re-survey it asked for. Same
machine, same read-only method: **119 project directories, 1,045 top-level transcripts, 179,614
records walked**, up from 33 / 837 / 13,211. The pin moves to `Claude Code 2.1.233`; §3.10's cell
inherits it from the adapter's own constant.
*The corpus is mixed-version, and that is a method change, not a footnote.* **16 CLI builds wrote
these records**, and 2.1.233 wrote only 1,704 of them. So "surveyed at 2.1.233" cannot mean "this
is what 2.1.233 writes" — every claim below about *when* a field arrived is attributed by the
record's own `version` field, not by the version of the binary installed. Two records types carry
no `version` at all and cannot be attributed either way; the block says so where that bites.
*A raw grep is no longer a safe method here, and the first survey's headline finding was wrong
because of it.* This corpus now contains telltale's own development sessions, which **discuss
these field names in their own text** — so a grep for `context_window_size` matches prose about
the absence of `context_window_size`. The re-measure therefore parsed every record and looked for
each token as a **JSON key at any nesting depth**. Raw-token hits: `rate_limits` 420 records,
`total_cost_usd` 288, `context_window_size` 87. Hits as an actual key: **one**.
**What held.**
| claim | 2026-08-01, 2.1.219 | 2026-08-16, 2.1.233 |
|---|---|---|
| first record carries `sessionId` | 60 of 60 sampled | **1,045 of 1,045** |
| `isSidechain` in top-level transcripts | 0 of 837 | **0 of 1,045** |
| recursive glob inflates the session list | 2021 vs 837 (2.4x) | **2,001 vs 1,045 (1.9x)** |
| `message.model` can be `""` | observed | **34 records** — the trap holds |
| `custom-title` payload key | `customTitle` | **`customTitle`, 6,327 records** |
| sessions with a `subagents/` sidecar | present | **107 of 1,045** |
The `ai-title` payload key is no longer unverified. §3.1 above says the survey established the
record's *shape* but not its key name, and that the adapter therefore matches on structure. The
re-measure names it: **`aiTitle`**, exactly one key beyond `type` and `sessionId`, on 2,331
records. **The structural matcher stays as it is** — it was correct, it is now confirmed correct,
and hard-coding the name buys nothing a measured structure does not already give.
**What changed, and what it costs.** Every capability gap stays `CapNone`. One of them keeps the
ruling and loses its stated reason:
- **`quota` — the reason was wrong, the ruling was right.** The 2026-08-01 pass grepped the
snake_case `rate_limits` and recorded zero matches. **The on-disk key is camelCase
`rateLimits`**, it hangs off `error` on API-error records, and **2.1.219 itself wrote it** — the
original grep missed a key that was already there, and this survey's own spelling hid it for two
weeks. It is still not a quota source, for a better reason than absence: it was **`null` in 32
of 32 records**, and it only appears where a request FAILED, never on a normal turn. A key that
is present and null is not a reading (§4a.1).
- **`context_pct` — unchanged and re-confirmed.** No `context_window_size`, `context_window` or
`contextWindow` key occurs at any depth. `message.context_management` exists (36 records) and is
`null` in 34 of them; the other two carry an empty `applied_edits` array. No denominator.
- **`cost` — unchanged.** No `cost` or `total_cost_usd` key at any depth. Still stdin-only.
- **`liveness` — unchanged, and the registry was re-opened on purpose.** §3.1 recorded
`~/.claude/sessions/.json` explicitly so a later build could be checked for a turn-start or
turn-end signal. 2.1.233 adds two keys, **`peerProtocol` and `procStart`**. Neither is that
signal. `procStart` hardens process *identity* — a pid plus its start time survives PID reuse,
which a bare pid does not — but it still answers only that a process exists, which §4a.4 rules
out. `CapNone` stands.
**Fields that appeared, and are deliberately modelled by nothing.** Recording an arrival is not
the same as reading it; per §7.16b, model-and-render-nothing needs a reason, and absence of need
is itself a finding.
- **`message.usage.output_tokens_details.thinking_tokens`** — the one genuinely new field.
Written only by 2.1.228, 2.1.229 and 2.1.233 (3,462 records), zero at 2.1.219. It breaks down
**output** tokens, and this adapter's token figure counts what entered **context**, so it feeds
no cell. Not modelled.
- **`message.usage.speed`, `.inference_geo`, `.server_tool_use`, `.iterations`** — all present at
2.1.219 as well, so not drift at all. None carries a window size.
- **a top-level snake_case `session_id`** on some `assistant`, `user` and `attachment` records
(680), beside the camelCase `sessionId`. Those records carry both. The adapter reads `sessionId`.
- **a `file-history-delta` record type** (9 records), which the observed type set above does not
list.
**One claim above is narrowed, and it is the canary's.** §3.10 called `sessionId` the field on
every JSONL record. At 2.1.233 it is not: **`file-history-snapshot` (37) and `file-history-delta`
(9) carry no `sessionId` at all** — 46 of 179,614 records. Neither type carries a `version` field
either, so *when* this changed cannot be attributed, and this block does not guess. The canary is
unaffected and the reason is worth writing down so nobody re-widens the claim: those two types
carry no `message`, no `cwd` and no title, so they feed nothing the adapter reads; `Saw()` fires
on the first record carrying the field rather than requiring all of them to; and the first record
of 1,045 of 1,045 transcripts still carries it, which is what the head read actually depends on.
§3.10's cell is reworded to *"on every JSONL record that feeds a field"*.
### 3.2 Codex CLI adapter sources — RESEARCHED FROM SOURCE, **NOT LIVE-VERIFIED**
Codex is now installed on the dev PC (2026-08-01): **Codex Desktop** (VS Code app
26.727.51351, bundling `codex-cli 0.146.0-alpha.9.2` under `%LOCALAPPDATA%\OpenAI\Codex`)
plus the **npm CLI** (`codex-cli 0.146.0`). The claims below were first read from
`github.com/openai/codex` at commit `1e85ca09` (2026-08-01): `codex-rs/utils/home-dir/src/lib.rs`,
`codex-rs/rollout/src/{lib,recorder,compression,policy,metadata}.rs`,
`codex-rs/protocol/src/{protocol,models}.rs`,
`codex-rs/login/src/auth/default_client.rs`, `codex-rs/thread-store/README.md` — and then
checked against the live corpus. **§3.4 carries the verified results and the itemized
remainder**; the adapter is not "done" until the remainder is discharged.
**Layout.** `$CODEX_HOME` (default `~/.codex`) `/sessions////rollout--.jsonl`,
fixed depth, no recursion. The date directory is **local** time, not UTC — deriving
today's directory from a UTC clock silently loses sessions across midnight and DST, so
the adapter walks the tree instead of computing a path. Files older than 7 days are
compressed in place to `.jsonl.zst` (zstd level 3); the adapter reads `.jsonl` only — a
`.zst` file is by construction ≥7 days cold and cannot be a live row, so skipping it
avoids a zstd dependency. `rollout-compression.lock` and `*.tmp` in the same tree are not
sessions. `archived_sessions/` is deliberately ignored.
**Envelope.** `RolloutLine { timestamp, ordinal?, #[serde(flatten)] item }` with
`RolloutItem` tagged `#[serde(tag="type", content="payload")]`:
```json
{"timestamp":"…","ordinal":42,"type":"session_meta|turn_context|response_item|event_msg|compacted|world_state|…","payload":{…}}
```
`EventMsg` is **internally** tagged, so its discriminator sits *inside* `payload`
alongside its fields (`payload.type == "token_count"`, with `info` / `rate_limits` as
siblings). This differs from the outer envelope and is the easiest thing to get wrong.
**Field mapping:**
| Normalized field | Exact source |
|---|---|
| Session id | filename uuid, cross-checked against `session_meta.payload.id` / `.session_id` |
| Working dir | `session_meta.payload.cwd`, then the last `turn_context.payload.cwd` |
| Git branch | `session_meta.payload.git.branch` (display-only extra) |
| Model | **last** `turn_context.payload.model` (not on `session_meta`) |
| **Context %** | **derived**: last `token_count` → `info.last_token_usage.total_tokens ÷ info.model_context_window` |
| **Quota** | `payload.rate_limits.primary` / `.secondary` → `{used_percent (0–100), window_minutes, resets_at (unix s)}` |
| Plan / CLI version / history mode | `rate_limits.plan_type`, `session_meta.payload.cli_version`, `.history_mode` (display-only extras) |
| Sub-agent thread | `session_meta.payload.agent_nickname` / `agent_role` non-null → not a session |
| Last activity | file mtime (matches Codex's own `updated_at`/`recency_at` derivation) |
`session_meta.payload.history_mode` is `legacy` (default) or `paginated` and changes
which *message* records exist. The adapter does not branch on it: `policy.rs` persists
`SessionMeta`, `TurnContext` and `EventMsg::TokenCount` under both modes, and those are
the only records it reads. The mode is carried as a display-only extra so a live
verification pass can see which one produced a given fixture.
`TokenCountEvent.info` and `.rate_limits` are both `Option`. **Judgement call, UNVERIFIED
(§3.4):** a `token_count` whose `info` or `rate_limits` is null is treated as *clearing*
that datum rather than leaving the previous value standing. `protocol.rs` annotates the
neighbouring field with *"`None` is unavailable, not a sparse-update recovery"*, which
reads as "we do not have it" rather than "unchanged"; it is also the conservative side of
the honest-gauge rule, since it never shows a number the vendor's most recent statement
did not contain. The 2026-08-01 live pass could not settle it — no session in the corpus
emitted a mid-stream null after a populated event — so the conservative reading stands
unfalsified rather than confirmed (§3.4 "still owed").
**Codex capability gaps:** no cost in USD anywhere; **no process-liveness registry** (mtime
is the only signal); no session title, so rows fall back to the workspace basename; cold
`.zst` sessions unreadable under minimal deps. Reading the SQLite state DB
(`codex-rs/rollout/src/state_db.rs`) is a **rejected** path — it would add a sqlite
dependency for metadata the JSONL already carries, and `thread-store/README.md` confirms
JSONL stays canonical and readable without SQLite.
#### Re-measure 2026-08-16 — `codex-cli 0.147.0`; the map holds, and two new fields are traps
`telltale doctor` reported this pin drifted (0.146.0 surveyed, 0.147.0 installed), so the
survey above was re-run against the live corpus rather than re-read from source. **The pin
now reads `codex-cli 0.147.0` and nothing in the field map changed.**
**Corpus scanned.** `~/.codex/sessions/`, walked exactly as the adapter walks it
(`archived_sessions/` not visited): **330 native rollouts**, 35 imported transcripts
filtered on the `external-import-turn` marker, 8,408 `event_msg` records of which **1,313
are `token_count`**. By writer version: 169 rollouts at `0.147.0` and 2 at
`0.147.0-alpha.6.5` — so **171 rollouts written by the installed build**, beside 144 at
`0.146.0` and 9 at `0.146.0-alpha.9.2`. The re-measure rests on rollouts 0.147.0 wrote,
not on old files re-read.
**What held.** Every path in the field-map table above still resolves on 0.147.0-written
rollouts, at the same rate or better than on 0.146.0 ones:
| path | 0.147.0 | 0.146.0 |
|---|---|---|
| `session_meta.id` / `.session_id` / `.cwd` | 169/169 | 144/144 |
| `session_meta.git.branch` | 149/169 | 120/144 |
| `session_meta.history_mode` | 169/169 | 144/144 |
| `turn_context.model` / `.cwd` | 163/169 | 31/144 |
| `info.model_context_window` + `.last_token_usage` | 57/57 rollouts that carry a `token_count` | 25/29 |
The sub-100% cells are **session shape, not drift**: a rollout that never reached a user
turn has no `turn_context`, and one that never reached a model call has no `token_count`
(112 of the 169 at 0.147.0 are `codex_exec` runs of that kind). `git.branch` is absent
exactly when the `cwd` is not a repository — `session_meta.git` itself is present, carrying
`commit_hash` and `repository_url`. Also holding: both canaries (`envelope type`,
`session_meta record`) on all 324 rollouts that carry any parseable record; `secondary`
still **null in all 1,278** populated `rate_limits`; `used_percent` / `window_minutes` /
`resets_at` unchanged as the only window keys; and `plan_type: "plus"` on 1,257 of them.
**What changed — additions only, no rename.** 0.147.0 adds fields and moves none, which is
the case §3.10 says costs this program nothing because every reader here addresses keys by
name. New on `session_meta`: `model_provider` (`"openai"`, 324/324), `base_instructions`
(now an object `{text}`), `context_window`, `dynamic_tools`, and `git.commit_hash` /
`git.repository_url`. New on `turn_context`: `turn_id`, `workspace_roots`, `current_date`,
`timezone`, `approvals_reviewer`, `permission_profile`, `comp_hash`, `personality`,
`collaboration_mode`, `multi_agent_version`, `realtime_active`, `file_system_sandbox_policy`.
`effort` gained `ultra` and `xhigh` beside the previously-observed `low`/`medium`/`high`.
**Two of the additions are traps, and are deliberately NOT read:**
- **`session_meta.context_window` is `{window_id: }` — an IDENTIFIER, not a size.**
The name invites reading it as the context denominator, and 324 of 324 carry it while
only 93 rollouts carry an `info.model_context_window`, so wiring it up would look like it
*widened* coverage. It cannot: a window id is not a token count, and dividing by one
would be an invented number of exactly the kind §4a.1 forbids. The denominator stays
`info.model_context_window`.
- **`turn_context.multi_agent_version` is the literal `"v2"` on all 288 turn contexts**, so
it is a format version, not a sub-agent marker. Treating it as one would reject every
session as a sub-agent thread. The sub-agent filter still keys on `agent_nickname` /
`agent_role` alone.
A third addition was checked and left alone: `collaboration_mode.settings.model` carries a
model id, and it **equals `turn_context.model` on all 288** turn contexts. The model source
is unchanged rather than merely still-working.
**What stays `CapNone`, now cited at 0.147.0.** Cost in USD, session title, sub-agent count
and process liveness are all still absent — a `rate`/`limit`/`cost`/`title`/`pid` sweep over
the 0.147.0 rollouts matches nothing beyond the `rate_limits` block already modelled. The
§3.3 matrix row is unchanged.
**What this pass could NOT exercise, stated rather than glossed.** No sub-agent thread and
no imported transcript written by 0.147.0 appeared in the corpus — `agent_nickname`,
`agent_role` and `external-import-turn` are all unobserved at this version. `ErrSubAgentThread`
and `ErrImportedTranscript` therefore still rest on the 2026-08-01 observation, and this
block does not claim otherwise: those markers are **unobserved here, not measured gone**.
Two further observations that are new and cost nothing: `history_mode: "paginated"` finally
appeared (one rollout, 0.147.0) and it carries `turn_context.model` and a populated
`token_count` exactly as `legacy` does — so §3.2's "the adapter does not branch on
`history_mode`" is now confirmed against a real paginated rollout instead of a source read.
And 6 rollouts on disk contain **zero parseable records**, which exercises for real the
`sampled <= 0` branch `internal/adapter/drift` calls its load-bearing case: they produce no
drift report, correctly.
### 3.3 Cross-vendor capability matrix — the asymmetry is a design fact, not a bug
| Field | Claude (disk) | Codex (disk) | Gemini (disk, §3.7) | Antigravity (disk, §3.8) | Cursor (disk, §3.9) | Grok (disk, §3.9a) |
|---|---|---|---|---|---|---|
| session id, cwd, git branch | yes | yes | id yes; cwd via `projects.json`; branch no | id yes; cwd via the trajectory blob's `file:///` URI; branch no | id yes; cwd via `workspaceStorage//workspace.json`; branch no | id yes; cwd verbatim in `summary.json`; branch **yes, unused** (`head_branch`, only when the cwd is a repo) |
| model | yes | yes | yes (per message) | yes (per generation, id + display name) | yes (`modelConfig.modelName`, one string; sometimes the literal `default`) | yes (`current_model_id`, one string) |
| token counts | yes | yes | yes (per message) | yes (per generation, self-checking) | context totals yes; per-message counts present and **always 0** | yes (per turn in `updates.jsonl`, plus a context total in `signals.json`) |
| context window size | **no** | yes | **no** (static table in CLI source only) | **no** (statusline payload only) | yes (`contextTokenLimit`) | yes (`contextWindowTokens`) |
| context % | **not derivable** | **derived** | **not derivable** | **not derivable** | **reported** (the vendor persists its own; derived from raw counts only if it is missing) | **reported** (`contextWindowUsage`, an integer the vendor truncates) |
| quota / rate limits | **no** (statusline stdin only) | yes | **no** (runtime 429 handling only) | **no** (statusline stdin only; never persisted) | **no** — plan *entitlements* on disk, no consumption record | **no** — nothing account-level anywhere in the store |
| cost USD | no (stdin only) | no | no | no | **no** — `usageData` `{}`, token counts unpopulated zeros | **per turn yes, session total no** — `costUsdTicks`, unit measured; no cumulative figure exists |
| process liveness | registry exists, deliberately unread (§3.1) | none | none | `steps.status` exists, structural only (never observed in-flight) | `status`/`generatingBubbleIds` exist, structural only (never observed in-flight); Hooks is the real seam | `active_sessions.json` exists and was **measured empty during a live turn**; `events.jsonl` phases outlive the process |
| session title | yes | no | yes (`summary` metadata) | **no** — the only free text on disk is prompt content | yes (`value.name`, vendor-generated) | yes (`generated_title`, vendor-generated; absent on headless runs) |
| sub-agent count | **derived** (`subagents/` sidecar, §3.1) | **no** | **derived** (`chats//` nest) | **no** — `parent_references` observed empty | **no** — `isSubagent`/`numSubComposers` observed zero throughout | **no** — a `spawn_subagent` tool exists, nothing about it reaches disk |
Codex is `CapNone` for the sub-agent count and not merely empty. Sub-agent *threads* do
exist in the Codex format — `session_meta.payload.agent_nickname` marks one, and the
adapter rejects those rollouts with `ErrSubAgentThread` — but they are whole top-level
rollout files carrying no link back to a parent session, so there is nothing to attribute
a chip to. Declaring the field and always emitting zero would assert "this Codex session
is running no sub-agents", which is not something the format lets us check.
Claude's quota lives on the statusline seam; Codex's lives on the disk seam. So the HUD's
quota block is Codex-sourced today, and the CONTEXT column carries a Codex number beside
a Claude em dash. **Nothing sources cost**, so the COST column auto-hides in every real
v1 frame — see the `v1-capabilities` render in §7.3.
Cursor is the first vendor to put a context percentage on disk as a number *it* computed,
which makes it the only unmarked bar in that frame. Everything else about it is the
asymmetry again from the other side: it is also the first vendor whose store holds live
credentials, so the adapter's most load-bearing property is the list of things it does
not read (§3.9, decisions/007). Grok is the second to report a percentage, and the first
to write a **dollar figure** to disk at all — and the COST column still auto-hides on its
rows, because what it writes is one turn's cost and never the session's (§3.9a). "Nothing
sources cost" became "nothing sources a session cost", which is a narrower sentence and
the same column.
**Percentage comparability.** Codex's own
`TokenUsage::percent_of_context_window_remaining` subtracts `BASELINE_TOKENS = 12000`
from both numerator and denominator; Claude's `context_window.used_percentage` is raw
input-token-based. **They are not the same statistic.** The adapter therefore does *not*
reproduce Codex's baseline-normalized figure: it computes a plain
`last_token_usage.total_tokens ÷ model_context_window`, declares it `CapDerived`, and the
HUD marks it with an estimate marker. See §6 Q7 for the resolution and its alternatives.
### 3.4 Live verification (ADR-001) — first pass run 2026-08-01; remainder itemized
**Environment:** Codex Desktop app 26.727.51351 bundling `codex-cli 0.146.0-alpha.9.2`
(every live rollout in the corpus was written by it, `originator: "Codex Desktop"`,
`source: "vscode"`), with npm `codex-cli 0.146.0` installed alongside. The source read
above was taken at CLI `2.1.219`; no contradiction between the two surfaced except where
noted below.
**Confirmed:**
- `sessions////` is the **local** date: events stamped `2026-08-02T00:12Z`
(UTC) sit under `08/01`. Walking the tree instead of computing today's path was right.
- `session_meta` writes **both** `id` and `session_id`, identical values.
- `history_mode` is `"legacy"` on every fresh thread.
- `model_context_window` is populated (`258400` for `gpt-5.6-terra`), so the derived
context percentage works as designed.
- `ordinal` is **not emitted** — the envelope is `{timestamp, type, payload}` only.
Fixtures 0002/0003 keep their `ordinal` deliberately (the field must stay tolerated);
fixtures 0006/0007 pin the observed no-`ordinal` shape.
- `rate_limits`, live values: **free plan** = `primary` only with `window_minutes: 43200`
(a 30-day window), `secondary: null`; **plus plan** = `primary` with
`window_minutes: 10080` (7 days), `secondary: null` so far, `plan_type: "plus"`. The
"record real values instead of hard-coding 5h/7d" instinct was right — neither plan
matches the guessed pair, and labels derive from `window_minutes` alone. Newer fields
(`limit_id`, `credits{}`, `plan_type`, `rate_limit_reached_type`) are parsed loosely;
`credits.balance` has been observed as both `null` and the string `"0"`, so nothing in
it is typed strictly.
- Go can `os.Open`, head-read, and tail-read a rollout **while a live codex process holds
it** (verified against an active session; Windows sharing mode is permissive).
**Learned, not on the checklist:**
1. **Imported external-agent transcripts.** Desktop onboarding imported 35 Claude
sessions into `sessions//` as rollout files. Markers: `session_meta` lacks
`thread_source` (native threads carry `"user"`); every turn's `task_started.turn_id`
is `external-import-turn-` (inside the head window in all 35 observed files); the
single `token_count` is synthetic (zero components, non-zero `total_tokens`, null
window, null `rate_limits`). The adapter rejects these with `ErrImportedTranscript`
on the affirmative `turn_id` marker only — absence of `thread_source` is not used, so
pre-`thread_source` CLI rollouts are unaffected. Rendering an imported Claude
transcript as a Codex row is a cross-vendor double count; the filter is not optional.
2. **`archived_sessions/` semantics confirmed the hard way.** The first inspection pass
found every real session in flat `archived_sessions/` and only imports in
`sessions//`, which read as "Desktop sessions are invisible to the adapter."
A later live session disproved that: Desktop writes live rollouts under
`sessions//` and threads move to `archived_sessions/` when archived — the
Desktop auto-archives its onboarding threads, which is what emptied the first hour.
Ignoring `archived_sessions/` remains correct.
3. **Windows mtime does not reliably advance mid-session.** On an active session the
newest records were stamped ~100 s *after* the file's mtime: NTFS defers the mtime
update while the writer holds the handle. `LastActivity` from mtime therefore
under-reports on live sessions (never over-reports). ~~The ruling is §6 Q8~~ —
**ruled and implemented 2026-08-01**: `LastActivity = max(mtime, newest record
timestamp)`, both adapters; see §6 Q8 for the rules.
4. Desktop threads run in per-thread scratch workspaces
(`Documents\Codex\\`), so the workspace-basename fallback shows the
thread slug, not a repo name. Cosmetic, vendor-truthful, unchanged.
**Discharged 2026-08-01, same evening:**
- ~~A rollout written by the **standalone CLI**~~ — observed live: two sessions from
the npm CLI at 0.146.0 write `session_meta` with `originator:"codex-tui"`,
`source:"cli"`, `thread_source:"user"` — same record shape and tree layout as the
Desktop writer, so no adapter change (nothing keys on originator). The observation
also reproduced §6 Q8 a second time: the first CLI rollout's mtime settled roughly
twenty minutes after the session ended, when the writer released the file.
**Still owed** (re-scannable on demand: `tools/scan-passive-tail.py`):
- Null `info`/`rate_limits` mid-stream, "cleared" vs "unchanged" (§3.2): the corpus
contained **no mid-stream nulls**, so the conservative "clearing" reading stands
unfalsified rather than confirmed.
- An **API-key login** capture (rate_limits expected absent), and whether a paid plan
ever populates `secondary`. Capture path when wanted: `codex login --with-api-key`
(reads the key from stdin), run one short session, re-scan, then plain
`codex login` to return to the ChatGPT plan.
- The 7-day `.zst` compression pass — unobservable until the corpus is a week old.
*Re-scan 2026-08-02:* 5 native rollouts (35 imports filtered), including one new
Desktop session — still zero mid-stream nulls; plus-plan `secondary` still null
across all 32 populated `rate_limits`; no API-key-signature session; no `.zst`
anywhere under `sessions/`. Oldest native rollout is 2026-08-01, so the `.zst`
pass stays unobservable before ~2026-08-08.
*Re-scan 2026-08-11:* 316 native rollouts, 1,251 populated `rate_limits`. All three
owed items stay negative. Two of them now rest on 316 rollouts rather than the 5 of
2026-08-02, and a fourth observation changes what the API-key capture must look for.
- **Mid-stream nulls: still zero**, now across 316 rollouts rather than 5. The
conservative "clearing" reading stays unfalsified rather than confirmed.
- **`secondary`: still null** in all 1,251 populated `rate_limits`. A plus plan does
not populate it.
- **`.zst`: still zero files** under `sessions/` — and this item is no longer blocked
on time. The oldest native rollout is 2026-08-01, so it was 10 days old at the scan,
and Codex wrote new rollouts on 9 later days. The pass is **measured absent on this
box, not unobservable**. The prediction above expired on ~2026-08-08.
- **A `rate_limits` object can report "no windows" without being absent.** 21
`token_count` records across 4 native `codex_exec` sessions (`cli_version`
0.146.0, 2026-08-07 and 2026-08-08) carry a `rate_limits` OBJECT whose `primary`,
`secondary`, `plan_type`, `limit_name`, `individual_limit` and
`spend_control_reached` are all null, beside `limit_id:"premium"` and a `credits`
block. Each of those sessions is null from its FIRST `token_count`, so this is not
a mid-stream clear and it does not settle §3.2.
**What it costs the owed capture.** `tools/scan-passive-tail.py` detects the
API-key signature as `rate_limits is None` alone, so a session of this shape
passes it unseen and reports as negative. Whoever runs the capture must check
both signatures: an absent `rate_limits`, and a present one whose windows are all
null. **The cause is deliberately not stated here** — nobody recorded which auth
mode those four sessions ran under, so a claim that they are API-key sessions
would be an inference, and §4a.1 forbids one dressed as a reading.
*Re-scan 2026-08-16* (during the §3.2 re-measure to `codex-cli 0.147.0`; that block carries
the field-map results, this line carries only the owed items). 330 native rollouts, 1,313
`token_count` records, 1,278 populated `rate_limits`. **All three owed items stay negative**,
and the newer corpus does not move any of them:
- **Mid-stream nulls: still zero**, now across 330 rollouts and 1,313 `token_count` records.
The conservative "clearing" reading stays unfalsified rather than confirmed.
- **`secondary`: still null** in all 1,278 populated `rate_limits`. A plus plan still does
not populate it.
- **`.zst`: still zero files** under `sessions/`, with the oldest native rollout now 15 days
old. Measured absent on this box, as the 2026-08-11 entry already ruled.
- **API-key capture: still not taken**, and this pass checked **both** signatures the entry
above demands. Signature A (`rate_limits` absent on every `token_count`): **zero sessions**.
Signature B (a `rate_limits` object whose windows are all null): **4 sessions** — the same
four from 2026-08-07/08 at `cli_version` 0.146.0, and **no new ones at 0.147.0**. So the
shape has not spread, and the auth mode behind it stays unrecorded and unclaimed.
### 3.5 Framing rule — now measured, not assumed (see §4)
The §4 hazards were quantified against the live Claude corpus:
- **64 KiB cap: firing.** 107 of 13,211 records exceed 64 KiB; the longest single line
is **1,004,230 characters**. `bufio.Scanner` at its default cap returns
`bufio.ErrTooLong` on ~0.8% of records, and an unchecked `Err()` silently truncates the
file — reading as "no more sessions". Adapters use `bufio.Reader.ReadBytes('\n')` via
`internal/jsonl`, which is the one tested implementation of this rule.
- **U+2028/U+2029: not observed** (0 raw `E2 80 A8`/`E2 80 A9` bytes across the 40 newest
transcripts). That is absence of evidence, not absence of hazard — the records carry
model-authored text and both characters are legal unescaped inside a JSON string. The
§4 byte-level rule is unchanged and both fixtures embed the character to pin it, with
a `.gitattributes` entry plus a byte assertion in each adapter's tests so a checkout
rewrite fails the build instead of silently disarming the test.
- **Trailing partial line:** all 7 live transcripts ended on `0x0A` at survey time, so
writes look line-atomic — a sampled observation, not a guarantee. The hold-until-`\n`
rule stands, and both fixtures end in a deliberate truncated record with no trailing
newline.
- **Windows concurrent read:** all 7 live transcripts opened with share-read/write while
their processes were running. Go's `os.Open` already requests
`FILE_SHARE_READ|WRITE|DELETE`, so no special handling is needed — recorded here so
nobody "fixes" it later.
### 3.6 Degradation rule
A vendor field the adapter cannot read renders as `—` (absent), never as a zero or a
stale value presented as fresh. A record that parses only partially degrades the fields
it could not source to `—` and keeps the rest. A truncated trailing line is not a record.
The exact renders are §7.7.
### 3.7 Gemini CLI seam — source-verified 2026-08-02; first live pass itemized
**Environment:** gemini-cli 0.53.1 installed via npm 2026-08-02; the persistence layer
read at tag v0.53.1 (`packages/core/src/services/chatRecordingService.ts` +
`chatRecordingTypes.ts` for the writer and record shapes, `config/storage.ts` for the
tree, `config/projectRegistry.ts` for the slug registry). This is the writer's own
source, not its docs — the same standard as the Codex `rollout` read (§3.2).
**Layout (from source):**
- Sessions: `~/.gemini/tmp//chats/session--.jsonl`.
The filename embeds only the session id's first 8 characters; the full id is on the
first record. `~/.gemini/projects.json` maps absolute project paths (lowercased on
Windows) to slugs; the slug scheme replaced sha256-hash directory names in 0.5x, and
the registry self-heals from `.project_root` markers.
- Sub-agent transcripts nest at `chats//.jsonl` — a structural
parent link, which is why Gemini declares `subagents` (derived) where Codex cannot
(§3.3): Codex's sub-agent threads are top-level files with no path back to a parent.
- Legacy pre-JSONL sessions are single-document `*.json`; the adapter skips them.
**Record shapes (from source):** the first line is metadata (`sessionId`,
`projectHash`, `startTime`, `lastUpdated`, optional `kind`/`directories`); message
records carry a string `id`, `timestamp`, `type`, and on `type:"gemini"` a `model` and
a per-message `tokens` summary (`input` = promptTokenCount, `cached` a subset of it,
`output`, `total`); `{"$set":{...}}` records patch metadata (including `summary`, the
session title, and whole-array `messages` checkpoints that can put megabytes on one
line — the §4 framing rule is earning its keep here); `{"$rewindTo":id}` truncates.
**Messages are upserts**: the writer re-appends the full record under the same id when
tokens or tool calls settle, so a linear last-wins pass needs no dedup map.
**Traps encoded in the adapter:**
- The writer **deletes** a session file on exit when it holds no resumable content, so
a file vanishing between Discover and Read is normal operation (`ErrSessionGone`,
row dropped silently).
- Nothing quota-shaped is persisted — rate limiting exists only as runtime 429
handling (`googleQuotaErrors.ts`, `retry.ts`). `quota` is CapNone, not empty.
- No context-window size reaches disk; the CLI's own percentage divides by a static
per-model table compiled into its source. An assumed denominator is an invented
gauge, so `context_pct` is CapNone — the §4a.7 sketch guessed
"derived" here, and the source read falsified the guess.
- `workspace` is read verbatim from the vendor's registry entry (REPORTED, a lookup
not a computation), with a fidelity caveat: the vendor lowercases the recorded path
on Windows.
- The adapter replays the writer's grammar, not just its records: `$rewindTo`
truncates the ordered message log (a rewind to an id outside the read windows
conservatively clears it), and a `$set` messages checkpoint clears and rebuilds it —
both mirroring the vendor's own loader. Independent review (2026-08-02) caught the
first cut ignoring both; a rewound-away 215k-token reading would have kept rendering.
- **Bounded-read limitation, stated:** the head/tail windows share the seam behaviour
of every adapter — a record crossing the boundary is read by neither window, and a
single line larger than the tail budget (256 KiB) is outside the read entirely. On
Gemini that line can be a whole-conversation checkpoint, so a giant checkpoint's
values are invisible until the next ordinary record re-establishes them. Accepted as
the same tradeoff the other adapters carry; the live pass below sizes real
checkpoints to check whether the budget needs raising.
**Market note (2026-08-02, post-merge):** Gemini CLI stopped serving consumer tiers
(free/Pro/Ultra) on 2026-06-18; it remains live for Gemini Code Assist
Standard/Enterprise licenses and paid API keys, with Antigravity CLI (`agy`) as the
consumer successor (ADR-003 addendum). This adapter therefore covers the
enterprise/API-key flavour. (A live session was nonetheless produced on this machine
2026-08-03 — the auth flavour behind it was not investigated; recorded as an observed
fact only.)
**First live pass — RUN 2026-08-03 and PASSED** (gemini-cli 0.53.1, the same version
the source read pinned; one real session, ~1.6 MB, 50 records, written live during
the check). Adapter output against it: discovered 1; model `gemini-3.5-flash`;
workspace `c:\users\sanle` via `projects.json` (lowercased-path registry confirmed);
LastActivity rode the mtime side of the Q8 fold (the file was being touched after its
newest record timestamp — the live-write pattern the fold exists for); subagents
derived 0; zero diagnostics; Validate green. Name is absent and honestly so — the
header carries no title field; the HUD label falls back to the workspace basename.
The itemized checks resolve as follows (original list kept below for the record):
- Metadata line is the first line: **confirmed.** Main sessions **carry
`kind:"main"`** — the fixture's omitted-field assumption was falsified; the adapter
is unaffected (only `"subagent"` branches) and the healthy fixture now carries
`kind:"main"` to match reality.
- Filename shape and prompt registry entry: **confirmed** (entry present, lowercased).
- Upserts against real traffic: **confirmed** — 27 message records, 20 distinct ids
(7 in-place updates). Per-message `tokens` **confirmed live** with shape
`{cached, input, output, thoughts, tool, total}` — unused by the adapter (context
stays CapNone per §4a.7's falsification), recorded for fidelity. Delete-on-exit for
non-resumable sessions: not observable (this session persisted).
- Checkpoint sizes vs the read budgets: the two live `$set` messages checkpoints were
**7.3 KiB and 14.5 KiB — neither approached the 64 KiB scanner cap**; the "expected:
yes, on any long session" guess did not materialize (checkpoints snapshot compactly).
The budget pressure came from elsewhere: **single message lines up to 746 KiB** were
observed, dwarfing both budgets. The newest checkpoint sat wholly inside the 256 KiB
tail and the bounded read produced correct output — a tail that starts mid-line
resyncs at the next newline, exactly the framing design.
- `$rewindTo`: **not exercised by this session** — remains fixture-verified only.
**ADR-003's verification hold is RELEASED**: the launch post may claim the Gemini
adapter live-verified.
Original itemized list (as written before the pass):
- Confirm the metadata line is the FIRST line in every live file, and whether main
sessions carry `kind:"main"` or omit the field (the fixture assumes omitted).
- Confirm filename shape and that `projects.json` gains the entry promptly (the
registry can lag a fresh project; the adapter treats a missing entry as absence).
- Confirm the upsert pattern and per-message `tokens` against real traffic, and the
non-resumable delete-on-exit behaviour.
- Observe whether a live `$set` messages checkpoint exceeds the 64 KiB scanner cap in
practice (expected: yes, on any long session), and whether any exceeds the 256 KiB
tail budget (which would put whole checkpoints outside the bounded read — see the
limitation above).
- Observe a real `$rewindTo` and confirm the truncate-or-clear replay against the
loader's behaviour on the same file.
### 3.8 Antigravity CLI seam — surveyed live 2026-08-02; statusline is the seam
**Environment:** agy 1.1.9, installed 2026-08-02. Closed source (the GitHub repo,
google-antigravity/antigravity-cli, is docs and examples only), so this survey is the
Claude Code method: documented contracts cross-checked against live observation, no
source read possible.
**Disk verdict (first survey — superseded the same day by the re-survey below).**
Interactive and headless sessions both write
`~/.gemini/antigravity-cli/conversations/.db`: SQLite, protobuf blobs
in every payload column (`steps.step_payload`, `gen_metadata.data` — inspected
read-only). The first survey judged the blobs unparseable without guessing (and the
repo had already rejected a SQLite dependency once, §3.2), and found no transcript: the
docs and the live payload both advertise
`~/.gemini/antigravity/brain//.system_generated/logs/transcript.jsonl`,
that path did not exist, and `antigravity-cli/brain/` held only empty `scratch/` dirs
at survey time. Verdict then: no honest HUD adapter; agy shipped as a statusline-only
vendor (ADR-004), with a standing watch item on the transcript.
**Re-survey (2026-08-02, later the same day; ADR-005 decision 5, prompted by ccusage
issue #1402): the disk seam is OPEN, and the watch item is RESOLVED — the transcript is
real.** Corpus: four conversations, all agy **1.1.9** — the same version the first
survey ruled on; this is a correction, not a version change.
- `brain//.system_generated/logs/transcript.jsonl` **exists for all four
conversations**, non-empty, plus an undocumented untruncated sibling
`transcript_full.jsonl` (`transcript.jsonl` marks its cuts with `truncated_fields`).
Written default-on, with `enableTelemetry: false`, under `antigravity-cli/` — the
docs' advertised `antigravity/` tree still does not exist. Why the first survey saw
only empty dirs is unresolved (a flush-timing artifact vs. a missed `.system_generated`
subdir); both observations stand as recorded. The transcript is plain JSONL, one line
per `steps` row (verified exact, 38/38 across the corpus): step_index, source, type,
status, created_at (RFC3339 UTC, second resolution), content, thinking, tool_calls,
exit_code.
- `gen_metadata.data` **decodes with a stdlib protobuf wire walk** (no schema, no
deps). Per-generation token counts live at `#1.#4`: uncached input (`#2`), total
output (`#3`), cache-read (`#5` — inferred from position and magnitude, lower
confidence), thinking (`#9`), answer (`#10`), and the per-generation response id
(`#11`, the dedup key — the *top-level* `#4` UUID is constant per conversation and
must not be used). Model id at `#1.#19` (`gemini-3.6-flash`) and display name at
`#1.#21`, matching the live statusline string verbatim. **Self-check: `thinking +
answer == output` held in 15/15 decoded generations** — the arithmetic identity that
promotes this from field-guessing to a schema. An adapter must assert it at read time
and degrade to absent rather than render a number that fails it.
- What an adapter could report as **measured**: Name, Model, Workspace (trajectory-blob
URI), LastActivity (max transcript `created_at`; the trajectory-blob timestamp is
session *start* — do not use it). **Partial**: context used-tokens (numerator
measured; the 1,048,576 window denominator appears nowhere on disk — it must come
from the statusline payload or a constant with stated provenance). **Absent**: cost
(consumer auth; no pricing on disk), quota (server-refreshed in memory, never
persisted). **Structural only, never observed live**: liveness (`steps.status` — all
38 observed rows `DONE`; no in-flight session sampled yet) and subagents
(`parent_references` empty, `has_subtrajectory` all zero, `define_subagent`/
`invoke_subagent` present in the tool registry).
- Build cautions, recorded before any adapter work: **WAL sidecars are load-bearing**
(a `-wal` larger than the `.db` was observed — copy `.db`+`-wal`+`-shm` together, or
open the live file strictly read-only and let SQLite replay; never open it for
write). `conversation_summaries.db` is a stale index (1 row for 4 conversations) —
enumerate `conversations/*.db`, never trust it. `cache/last_conversations.json` is
written at session start and points at the *previous* session for a reused
workspace. Transcript `content`/`thinking` and the request blobs are PII (full
prompt text, file contents; the account email appears in `cli.log`) — an adapter
reads structural fields only and never surfaces those into the HUD or any log. All
protobuf field numbers are reverse-engineered and unversioned; the corpus is one
day, one model — field stability across models is untested.
**Statusline seam — verified live.** Capture method: a temporary statusline command
(`telltale-capture.cmd`) appending each stdin payload to a file; six payloads captured
across one interactive session. Confirmed against the documented schema: `product:
"antigravity"` on every payload (the routing marker); model with display_name/effort;
full context accounting (`context_window_size` 1,048,576 observed, `used_percentage`,
per-request `current_usage` incl. cache reads); **named weekly quota buckets**
(`gemini-weekly`, `3p-weekly` on the Starter tier) each carrying `remaining_fraction`,
`reset_time`, `reset_in_seconds`; `agent_state` observed transitioning `tool_use` →
`idle`; `tool_confirmation_pending` observed `true` while a permission prompt was on
screen. Documented but not yet observed live (the capture session ran outside a repo):
`vcs`, `artifact_count`, `task_count`, `execution_mode` — the parser carries them as
optional; the branch segment's live confirmation is the itemized remainder here.
**Adapter built, 2026-08-02 (`internal/adapter/antigravity`).** What the
adapter took from this survey and what it left:
- **Took:** Model (`#1.#21` display name, `#1.#19` id fallback), Workspace (the
trajectory blob's URI, converted to a native path), LastActivity (the Q8 fold over the
transcript's newest `created_at` and the mtimes of the transcript, the database and its
sidecar — the sidecar because on a live conversation that is the file being written).
All three REPORTED; nothing is derived. Name was sourced from the conversation id at
first shipping (shortened for the grid, the only label on disk that is not somebody's
prompt) — **revised 2026-08-12 (ruled):** the HUD row's Name showed the id prefix
instead of the workspace basename every other vendor with no on-disk title falls back
to, so Name moved from REPORTED to `CapNone` and the adapter no longer writes it; the
HUD sources the row's label itself, the same fallback a Gemini row takes with no
summary of its own.
- **Left:** context % (the numerator is measured and the denominator is not — the token
totals are display-only extras instead), cost, quota, liveness, subagents and (as of
the 2026-08-12 revision) name, all `CapNone` for the reasons itemized above. The
liveness and subagent deferrals are pending live observation, not pending effort.
- **The identity is asserted, not assumed:** every generation must satisfy `thinking +
answer == output` before its numbers count. A generation that fails contributes
nothing and the row says a self-check failed. Across the live corpus on the day the
adapter landed the identity held **16/16** (one more generation than this survey's 15,
a conversation having advanced in between).
- **Zero dependencies added.** The `.db` and `-wal` bytes are read by `internal/sqlite`,
a read-only reader written for this seam: header, `sqlite_master`, table b-trees,
overflow chains, and a WAL overlay applying SQLite's own recovery semantics. The
rejected alternative and the reasoning are decisions/006.
- **Live verification, same day:** all five local conversations discovered and read;
model `Gemini 3.6 Flash (High)` on every one; three workspaces resolved to
`C:\Users\sanle` and two absent (those conversations genuinely carry no URI — absence,
not degradation); a real 284 KiB `-wal` parsed and accepted with no diagnostic; every
session passed `Validate` with an empty degraded set.
**Fidelity notes:** the docs' storage path (`antigravity/`) and the real one
(`antigravity-cli/`) disagree — the re-survey confirmed the real tree (the docs path
has never existed on this machine); `model.id` equals the display
string ("Gemini 3.6 Flash (High)"), not a machine id; the payload carries the signed-in
email and plan tier, so real captures are PII and never enter `testdata/`.
**Re-verification, 2026-08-15 — the pin moves to `agy 1.1.13` and the field map does
not.** agy self-updates, so the version below was read from `agy --version` at read
time, never from a release note. The corpus is 20× the one §3.8 ruled on: **81
conversations, 4,284 transcript records, 1,926 decoded generations.** The adapter was
run over the live tree unmodified and every field it declares still sources.
| field | §3.8 (agy 1.1.9) | re-read (agy 1.1.13) | verdict |
|---|---|---|---|
| model | 5/5, one id string | **81/81** | sources; id string DRIFTED, see below |
| workspace | 3 of 5, 2 absent | **23 of 81**, 58 absent | sources; absence confirmed genuine |
| last_activity | measured | **81/81** | sources, no drift |
| tokens (extras + §7.17 sum) | identity 15/15 | **identity 1,926/1,926** | sources, no drift |
| liveness | `CapNone` — all 38 rows DONE | `CapNone` — **164 RUNNING now seen** | stays `CapNone`, for a NEW reason |
| subagents | `CapNone` — never observed | `CapNone` — **3 `INVOKE_SUBAGENT` steps** | stays `CapNone` |
**0 rows degraded, 0 diagnostics, 0 unparseable transcript records** across the whole
corpus — including conversations whose `-wal` sidecar is 25× the `.db` it belongs to,
which is the sidecar contract still holding at scale.
- **The two vendor release-note claims are NOT corroborated at this seam, and stay
vendor claims.** 1.1.13 claimed a transcript-corruption fix during compaction: 0 of
4,284 records failed to parse, so there is no corruption here to have been fixed and
nothing measurable changed. 1.1.12 claimed a Windows drive-letter fix in a transcript
path converter: all 23 workspace URIs present read `file:///C:/Users…` — the same form
§3.8 recorded — and `pathFromFileURI` converted 23 of 23 to `C:\Users\sanle`. Whatever
the vendor fixed, it was not the shape this adapter reads.
- **Workspace absence is absence, not a broken converter.** This is the check that
separates the two, and it was run rather than assumed: of 81 trajectory blobs, **23
contain a `file:` substring anywhere in their bytes and 58 contain none.** The 58
carry no URI to convert, so the field is correctly absent (§4a.1's zero-vs-absent
distinction), and the converter has no silent failure hiding behind them.
- **`status` gained `RUNNING`, and it is still not liveness — the new evidence makes the
refusal stronger.** §3.8 declined the field because every one of 38 rows read `DONE`.
The re-read found the missing state: **164 `RUNNING` rows in 27 of 81 transcripts**,
158 of them on `RUN_COMMAND`. But the oldest is dated **2026-08-03, thirteen days
before the read**, in a conversation nothing has touched since. `RUNNING` therefore
means "no terminal status was ever written for this step", which is equally true of a
live command and of one whose process was killed. Wiring it would have reported dozens
of long-dead sessions as working — the exact dishonest-gauge failure ADR-001 exists to
prevent. The field stays `CapNone` and the HUD keeps classifying age from
`last_activity`.
- **`INVOKE_SUBAGENT` is now observed (3 steps) and still buys no count.** It is the
first evidence the feature is used at all. A step recording that a subagent was invoked
at some past moment is not a number of subagents running now, which is what the field
means, so the `CapNone` stands.
- **The model id gained a second spelling.** `gemini-flash-3.6-high-control` (37 rows)
now appears beside `gemini-3.6-flash` (41 rows) for the one display string "Gemini 3.6
Flash (High)", and **3 rows carry the display name at `#1.#21` with nothing at
`#1.#19`** — the first corpus to exercise the adapter's either-half-will-do fallback.
Both strings are read verbatim; normalizing them would be inventing a vocabulary.
- **New transcript keys, none of them read:** `error` (23), `error_code` (9), alongside
the `truncated_fields` (345) §3.8 already described. The `step` struct is an allowlist,
so added keys cost nothing — and no key the adapter *does* read has gone missing.
- **A `logs/chunks/` tree appeared (first seen 2026-08-14, on the 4 newest
conversations)**: `chunks/transcript/00000000.jsonl` plus a `chunks/transcript_full/`.
This is the likeliest mechanism behind the compaction claim. It changes nothing today —
the flat `transcript.jsonl` is still written and was **byte-identical (md5) to its
single chunk on all 4**. All four hold exactly one chunk, so the case that would matter
— whether the flat file stays complete once a second chunk exists — is **unobserved**,
and is the standing watch item this block leaves behind.
- **The docs' `antigravity/` tree now EXISTS and is still not the data.** §3.8 recorded
that it had never existed. At 1.1.13 `~/.gemini/antigravity/` holds a `bin/`, a
`builtin/skills/` tree and crash logs, while its `brain/` and `conversations/` are
**empty** and all 81 conversations remain under `antigravity-cli/`. An adapter that
picked its root by probing which path exists would now pick the empty one; this one
roots by name.
**`/quota` cross-check: MEASURED, and deliberately NOT built.** The probe was gated on
proving it is free before anything could be wired to it, because a probe that quietly
started a turn would plant phantom rows in telltale's own gauge. Two `agy -p "/quota"`
invocations (PowerShell — Git Bash mangles the leading slash) against a quiescent store,
each bracketed by a full snapshot:
- **No residue in the store the HUD scans.** 81 conversations / 81 brain dirs / 81 HUD
rows / 1,926 generations / identical token totals, before and after both runs. The
probe writes only `cli.log`, `log/`, `last_check.timestamp`, `updater/update_status.json`
and a **0-byte** `crashes/crash__.log` sink — none of them in the adapter's
scan path.
- **No token cost.** Both runs returned identical remaining figures (87% / 98% / 100% /
100%) and neither added a generation. A spent turn writes a conversation and a
`gen_metadata` row, as every other agy turn in this corpus did.
**So the gate passes and the feature is still wrong to build**, on evidence the
measurement itself produced: the probe takes **2.9–3.8 s** and its numbers come from the
server. Wiring it into a gauge would break two rules that are not about cost at all — the
statusline path "reads nothing beyond stdin" (§2, revisitable "only with a measured
budget", and 3 seconds is not one), and the gauges make no network calls
(`CLAUDE.md`'s read/write boundary). It would also spend that round trip cross-checking a
number the vendor already hands us free on stdin, which is where the quota buckets come
from today. Nothing is wired; this paragraph is the record of why.
**`/config` in `doctor`: NOT built, and the measurement that would license it was not
made.** `doctor`'s charter is a cheap local `--version` parse that "never starts a turn,
spends quota, reads a credential, writes under `~/.telltale`, or calls the network"
(`cmd/telltale/main.go`). The `/quota` timings above show an agy slash command is a
different animal: seconds, not milliseconds, and server-backed. Its behavior with the
network or auth absent was **not measured** — doing so means disabling networking or
de-authenticating the operator's account — and an unmeasured network-backed probe in
`doctor` would red-cross exactly when the operator is offline, which is the dishonest
cell that mode exists to avoid. `doctor` is left alone.
**Amended 2026-08-16 — the multi-chunk case is PINNED, and it is still not MEASURED.**
The re-verification block above left a standing watch item: whether the flat
`transcript.jsonl` stays complete once a second chunk exists. That is still unobserved.
No live multi-chunk conversation has been captured on any machine, and this amendment
does not claim one. What changed is that the adapter's side of the question now has
tests instead of a comment — `internal/adapter/antigravity/multichunk_test.go`, over a
**synthesized** two-chunk fixture built by `testdata/gen_fixtures.py`.
**The distinction is the whole point of writing this down.** Every other fixture in that
directory reproduces a shape this section MEASURED. This one reproduces a shape nobody
has seen, extended from the measured single-chunk shape along the vendor's own naming
(`logs/chunks/transcript/0000000N.jsonl`, `logs/chunks/transcript_full/`), with the one
relationship the corpus did measure — the flat file is byte-identical to its chunk —
carried forward to two chunks. So the fixture pins **the adapter's contract with
itself**, never a vendor claim, and nothing here is admissible as evidence about what
agy writes.
**What the pin found: the adapter is correct under that contract, and no code changed.**
Six properties now have tests, and each was mutation-checked — the guard it names was
deliberately removed and the test failed with the intended message, because a green test
that cannot fail pins nothing.
- **The flat file is what gets read.** The chunk tree's `transcript_full` files carry a
poison step dated 23:59 against the flat file's newest at 09:11:39, dated in the past
so the future-skew guard cannot silently swallow it. An adapter that followed the
chunk tree moves `last_activity` by fourteen hours and the test says so by name. Note
the limit honestly: `chunks/transcript/` is byte-identical to the flat file, so a
switch to *that* is invisible to any test — which is precisely why the switch must not
be made on the strength of a directory name.
- **A flat file frozen at the first chunk costs precision, not the row.** This is the
worst plausible form of the unobserved case: the vendor stops appending to
`transcript.jsonl` and writes only into chunk 1. The §6 Q8 mtime fold carries it — the
database is the file still being written — so `last_activity` stays correct, nothing
degrades, and nothing is invented. That is an honest outcome rather than a lucky one:
the adapter never claimed the transcript was its only clock.
- **Damage inside the second chunk degrades the reading, not the row** (ADR-001's
partial-read rule). Both branches are exercised: JSON that does not parse, and JSON
that parses carrying a timestamp that does not. Three torn records, counted once,
stated once in `Diagnostics`, and the newest surviving step still dates the row.
- **The read budget's blind middle is absence, not corruption.** A transcript big enough
to chunk is the first one this adapter splits into a head and a tail read at all, and
the ~120 records it skips must never be reported as unparseable. An adapter accusing
the vendor of damage that is not there is the same class of dishonesty as rendering a
number it did not read.
- **The head/tail overlap guard now has a test, and it had none before.** Every other
fixture transcript on this vendor is under 1 KiB — swallowed whole by the tail read,
never reaching `jsonl.Head`. The narrowing branch executes only between the tail budget
and the full budget, so the test truncates a copy into that band and damages one record
in the contested bytes. With the guard the count is 1; with it removed the measured
result is **2 unparseable transcript records skipped** for one torn record. Duplication
is invisible to every other signal here, because the newest-timestamp fold is a maximum
and survives being fed a record twice; the unparseable counter is a sum, and it is the
one place a doubled read surfaces.
- **The PII boundary holds over a transcript 400× the size of the others**, chunk tree
included.
**What is still owed is the capture, and it is unchanged by any of this.** A synthetic
fixture can prove the adapter is self-consistent across the behaviors the vendor might
have. It cannot say which one agy has, and the four conversations that grew a chunk tree
still hold one chunk each. The instrument this seam wants is a live multi-chunk
conversation, and nothing above substitutes for it.
**Re-capture, 2026-08-17 — the statusline pin moves to agy 1.1.13, and one field is
FALSIFIED.** The statusline seam block above is pinned at 1.1.9 and six payloads. That
record is superseded here and kept as the older reading. Capture method: the same
temporary statusline command as 2026-08-02, appending each stdin payload to its own file;
**fifteen payloads across one live interactive turn**, in Windows Terminal. The session
ran outside a git repo again, so `vcs` is still unobserved.
**The four-bucket prediction is CONFIRMED, with ids rather than labels.** §3.8's `/quota`
cross-check reported four windows under human labels and refused to call that an
observation of bucket ids, because the labels are not the ids. The payload now names them:
**`3p-5h`, `3p-weekly`, `gemini-5h`, `gemini-weekly`** — a five-hour and a weekly window
for each of two model families. **All four carry `remaining_fraction`, `reset_time` AND
`reset_in_seconds`**; the 1.1.9 record claimed those three keys for two buckets, and they
hold for four. Nothing in the renderer changed, which is what §2.1's verbatim-id rule was
built to buy. The `week` page's fixture already carried all four names, so the fixture was
right before the capture was taken.
Two quota readings are recorded because both are honesty cases rather than trivia:
- **Three of the four buckets reported `remaining_fraction` exactly 1**, serialized as the
bare literal `1` rather than `1.0`. That renders `0%` used — a MEASURED zero, drawing a
full empty track, not an absence. §4a.1's zero-versus-absent distinction, arriving off
the wire this time instead of out of the renderer. `week.txt`'s `3p-weekly ─── 0%` row is
the shape this produces, and it was synthesized correctly.
- **`remaining_fraction` never moved across the fifteen fires.** Only `reset_in_seconds`
counted down. The quota a line draws is therefore not the turn it is drawing. This is a
vendor property, it is not hidden, and nothing here compensates for it.
**`agent_state` is observed live, and the documented vocabulary is incomplete.** The turn
moved `authenticating` → `idle` → `working` → `idle`. `idle` and `working` are on the
documented list; **`authenticating` is not**. The renderer draws an unknown state verbatim
in dim, so an unlisted value costs nothing — which is the point of having built it that
way. The list is documented, not closed, and this is the measurement that says so.
`thinking`, `tool_use` and `initializing` did not appear on this turn.
**The payload GREW since 1.1.9, into the same shape Claude's block has.** `context_window`
now carries `total_input_tokens`, `total_output_tokens`, `context_window_size`, float
`used_percentage` and `remaining_percentage`, and a populated `current_usage` object
(`input_tokens`, `output_tokens`, `cache_creation_input_tokens`, `cache_read_input_tokens`)
— the §7.16b block's shape, on a second vendor. Also present and not previously recorded
here: `conversation_id` (**equal to `session_id` byte for byte on all fifteen fires**),
`exceeds_200k_tokens`, `sandbox.enabled`, `terminal_width` and `model.effort`. The parser
was checked field by field against this capture: it models fourteen of the sixteen
top-level keys. `email` is excluded deliberately and that exclusion is tested. The one
field genuinely unmodelled by silence was `exceeds_200k_tokens`, and
`internal/antigravity/stdin.go` now carries the agy-side reason instead of inheriting a
ruling made about Claude's payload. **Nothing new was wired to a renderer or a relay** —
§7.16's display hold binds this seam.
**The numbers populate on the LAST fire only, and that is a zero-versus-absent collapse in
the vendor's own serializer.** Through fourteen of fifteen fires the vendor sent
`used_percentage: 0` with the token counters at `0`; every real number arrived at once on
the final fire, after the turn. On the very first fire it sent `used_percentage: 0`
alongside `context_window_size: 0`, which is a placeholder that no pointer type can tell
from a measured zero — the vendor writes a literal `0` instead of omitting the key. So the
statusline honestly draws `ctx 0%` for the length of a turn and then jumps. This is
recorded, not compensated for: inferring "not yet known" from a zero the vendor stated
would be exactly the invention §4a.1 forbids, and the collapse happens upstream of
anything this repo controls.
**`transcript_path` is FALSIFIED, and §2.1's refusal now has a measured reason.** §2.1
declined to display the field because the payload's value was *unverified* — the transcript
is real on disk, but nobody had captured a payload to check whether its path pointed at it.
This capture checked. **It does not.** The payload roots the path at
`~\.gemini\antigravity\brain\\.system_generated\logs\transcript.jsonl`, and the real
transcript for that same session exists only under `~\.gemini\antigravity-cli\brain\\`.
The payload drops the `-cli` segment. Both roots exist at 1.1.13 and that is what makes the
trap sharp: `antigravity/brain/` is present and EMPTY while `antigravity-cli/brain/` holds
every conversation, so a reader that trusted the payload would open nothing and could not
distinguish a missing file from a missing session. This is the docs-versus-data tree
disagreement §3.8 recorded in 2026-08-02, still unfixed at 1.1.13 and now measured on the
statusline seam as well as on disk.
**No code trusted it, and that was checked rather than assumed.** Every use of
`transcript_path` and `TranscriptPath` in the repository was read. No non-test code
dereferences the field — it is parsed and held, never opened, stated or rendered — and
`TestTranscriptPathIsHeldButNeverAPath` already pins that. The HUD adapter reaches the
transcript by its own root by name (`internal/adapter/antigravity`), never by trusting a
path handed to a gauge, which is the decision that makes this vendor bug a non-event here.
So nothing is fixed, because nothing is broken; §2.1's refusal is upgraded from caution to
measurement and the field stays parsed-and-unused, because deleting it would delete the
evidence. One further shape is recorded for whoever reads a raw capture next: before a
session id exists, the vendor joins an EMPTY id segment rather than omitting the key, so
the path collapses to `…\brain\.system_generated\logs\transcript.jsonl`.
**One inventory cell went stale with the re-verify, and is now fixed.** §3.10's canary
table read `agy 1.1.9` for the Antigravity row after the adapter moved to `agy 1.1.13`;
the row was corrected in the same-day ledger pass. The weakness this exposed stands
recorded: `TestTheCanaryInventoryMatchesThisAdapter` substring-matches the pin anywhere
in `design.md`, so a dated block quoting the new pin turns the guard green while the
table it guards stays wrong. Scoping that assertion to the table is unowned.
### 3.9 Cursor (Composer) seam — surveyed live 2026-08-02; the store is open, and it holds credentials
**Environment:** Cursor 3.14.7, Windows. Closed source, and the store format is
**undocumented and unversioned** — there is no changelog for the 3.12–3.14 line and no
schema anywhere. So this is the Claude Code / Antigravity method again: a read-only live
survey, every field cross-checked, nothing claimed that was not observed.
**Store inventory.** One SQLite database backs every Composer session:
```
%APPDATA%\Cursor\User\
globalStorage\state.vscdb 9.3 MB at survey time
globalStorage\state.vscdb-wal 4.6 MB — LIVE, Cursor was running
workspaceStorage\\workspace.json
workspaceStorage\\state.vscdb 4096 bytes
workspaceStorage\\state.vscdb-wal 300–540 KB
```
Three tables. `composerHeaders(composerId, workspaceId, createdAt, lastUpdatedAt,
isArchived, isSubagent, recency, checkpointAt, value)` is one row per session, `value`
being a JSON blob. `cursorDiskKV(key, value)` is key/value: `composerData:` holds
per-session state, `bubbleId:*` and `ofsContent:*` hold the message payloads.
**`ItemTable(key, value)` holds the credentials** — see below.
**Field map.**
| Field | Verdict | Source, and what was measured |
|---|---|---|
| name | **MEASURED** | `composerHeaders.value.name` — a vendor-GENERATED session title ("Multi-vendor orchestration"), same class as the Claude summaries the HUD already shows. Absent on a session the vendor has not titled yet; the composerId's first eight characters then. |
| model | **MEASURED** | `composerData:` → `modelConfig.modelName`. Observed `composer-2.5`, `grok-4.5`, `gpt-5.6-sol`, and the literal `default`. One string, no display name beside it. |
| workspace | **MEASURED** | `composerHeaders.workspaceId` → `workspaceStorage//workspace.json` → `.folder`, a `file:///c%3A/...` URI (lower-case drive letter, percent-encoded colon). `value.workspaceIdentifier.uri.fsPath` carries the same path and confirms it. |
| context % | **MEASURED** | `composerData.contextUsagePercent`, a float the vendor persists: 37.05, 29.008, 12.38, 10.99 observed, alongside raw `contextTokensUsed`/`contextTokenLimit` (94854/256000, 24763/200000, 44k/1M). The header row's `value.contextUsagePercent` mirrors it and agreed **exactly** on all four rows carrying both. |
| cost | **ABSENT** | `usageData` was `{}` in all 8 blobs, and `tokenCount.inputTokens`/`outputTokens` read **0 in all 310 message rows**. The schema is present and never populated. |
| quota | **ABSENT** | No consumption record on disk anywhere. What IS there is plan ENTITLEMENT (`credit_dollars: 25`, `included_usage_dollars: 40`) — what the plan grants, not what has been spent. |
| last_activity | **MEASURED** | `lastUpdatedAt` / `recency` / `checkpointAt`, epoch **milliseconds**. Not every row has all three (`lastUpdatedAt` was NULL on 4 of 9). |
| liveness | **PARTIAL, never observed in flight** | `composerData.status` (`completed`, `aborted`, `none`), `generatingBubbleIds`, `hasBlockingPendingActions`. All read terminal or empty across the corpus; no session was ever sampled mid-generation. |
| subagents | **ABSENT (structural only)** | `isSubagent` 0, `numSubComposers` 0, `subComposerIds` `[]` on every row. The fields exist; the observation does not. |
**Amended 2026-08-11: a larger survey HARDENED the `cost` row and CORRECTED the reason
behind the `quota` row. The corrected reading lived only in §7.16 until now.** The two
rows above rest on the 2026-08-02 corpus — 8 `composerData` blobs and 310 message rows.
A re-verification on 2026-08-08 (§7.16) read a bigger one and found the same thing
harder: `usageData` was `{}` in **19 of 19** blobs, `tokenCount` was zero in **1,622 of
1,622** message rows, and 78 `turn_ended` records and 51 transcripts carried status and
no numbers. `cost` **ABSENT** therefore gets stronger, not weaker, and the original
counts above stay as the smaller measurement they were.
The `quota` row's VERDICT stands and its REASON does not. The row reads `credit_dollars:
25` and `included_usage_dollars: 40` as plan ENTITLEMENT — what the plan grants. The
2026-08-08 survey measured those constants as **Statsig experiment values stamped
`is_user_in_experiment:false`**, and they were the only account figures anywhere on that
disk. So they describe an experiment this account is not in, not what this account was
granted. §7.17's absence table already renders Cursor as `no quota anywhere · its store
holds experiment values, not usage` on exactly that measurement.
What follows from both rows is the same: nothing about consumption reaches the store as
a byproduct of a turn, so the number is FETCHED rather than found. **Cursor Hooks is the
real token seam** — the vendor's own documented `afterAgentResponse` step hands a command
hook the turn's token counts on stdin, and §7.16 is the record of it, including why print
mode's derived `inputTokens` was refused.
**Most rows are not sessions.** 9 header rows, of which **5** were: the empty-state
draft (`composerId` literally `empty-state-draft`, `isDraft` true), two pre-created
composers a new window makes before anyone types (no title, no `lastUpdatedAt`), and two
archived threads. Filtering on `isDraft`, `isArchived`, `isSubagent`,
`workspaceId == "empty-window"` and the draft sentinel left 4 real sessions, which is
what the HUD shows.
**Build cautions, recorded before any adapter work:**
- **The WAL is where the data is.** This is stronger than the usual "read the sidecar
too" (§3.8): every workspace-level `state.vscdb` was **4096 bytes — one empty page —
with 300–540 KB in its `-wal`**. A reader that opens only the `.db` there does not get
stale data, it gets an empty database. The global store was 9.3 MB with a 4.6 MB live
sidecar. Read both as bytes; never open or lock the file Cursor owns.
- **THE STORE HOLDS LIVE CREDENTIALS.** `ItemTable` carries `cursorAuth/accessToken`,
`cursorAuth/refreshToken`, `mcpOAuth.secret.*` and git-IPC auth tokens;
`composerData` blobs carry `blobEncryptionKey` and
`speculativeSummarizationEncryptionKey`. This is the first vendor where "read the
store" and "read the user's tokens" are the same sentence unless an adapter is
explicit about what it will not touch. The allowlist is decisions/007 and it is
asserted by a test, not promised in a comment.
- **`ItemTable['composer.composerHeaders']` is a legacy JSON mirror and it is STALE.**
At survey time it named **3** composers to the table's **9**, and all three were rows
the filter drops — so an adapter reading the mirror would report *zero* sessions on a
machine running four. Read the table. (The mirror is in `ItemTable` anyway, so the
credential rule forbids it independently.)
- **Timestamps are mixed.** The header columns are epoch milliseconds; ISO-8601 UTC
strings live in the same store at `composerData.fullConversationHeadersOnly[].createdAt`
(per-message, structural). That path is a finer-grained activity signal than
`lastUpdatedAt` and the adapter deliberately does **not** read it: it is outside the
allowlist, and widening the allowlist for precision is exactly the trade this seam
should not make. The timestamp reader accepts both encodings anyway, because an
unversioned INTEGER column is not promised to stay one.
- **One store, many sessions**, so the store's file mtime dates the STORE. Folding it
into `last_activity` — which is what §6 Q8 prescribes for every other vendor — would
mark every Cursor row live whenever Cursor wrote anything, forever. The Q8 fold runs
over the per-row timestamps only, and degrades when none is readable. This is the one
deliberate departure from the Q8 shape, and the reason is that the shape assumes one
file per session.
- **Version fragility.** No changelog, no schema version in the file, no documentation.
The adapter addresses columns by NAME (read out of the CREATE statement `sqlite_master`
stores) rather than by position, and a store missing `composerHeaders` or one of the
columns the field map needs is reported **unreadable with the reason** rather than as
zero sessions — a wrong "your agents are idle" is worse than a visible "I cannot read
this".
**The needs-input seam, for later.** Cursor's *documented* surface is Hooks
(cursor.com/docs/hooks): the base payload carries `conversation_id`, `model`,
`workspace_roots` and `transcript_path`, and `preCompact` carries context numbers. That
is a supported, versioned contract and it is where a liveness/needs-input signal should
come from — not from reverse-engineering `status` out of the store. Recorded as the
watch item; §8 carries it. Separately, the `cursor-agent` CLI keeps its own store, which
is **not installed on this machine** and therefore an unverified surface, out of scope.
(That last sentence held until 2026-08-29. `cursor-agent` is installed here now and its
session manifest IS read; see this section's 2026-08-29 addendum. The Cursor Hooks watch
item above is untouched.)
**Adapter built, 2026-08-02 (`internal/adapter/cursor`).**
Name, model, workspace and last_activity REPORTED; context % declared DERIVED and marked
per read only when the adapter computed it; cost, quota, liveness and subagents
`CapNone`. `internal/sqlite` gained two additions rather than being worked around:
`Columns` (split the column list out of the stored CREATE statement, so columns are
addressed by name) and `Rows` (stream a table to a callback that can stop, so filtering a
key/value table by prefix retains nothing).
**Live verification, same day:** 5 sessions discovered and read against the real store
**with Cursor running** and a 4.6 MB live sidecar; workspaces resolved to `agent-ops` and
`faithfulness-judge`; models `grok-4.5`, `composer-2.5`, `gpt-5.6-sol`; context
percentages 4.42 / 12.93 / 37.05 all vendor-REPORTED, none marked derived; one session
with no `composerData` row rendering an absent model rather than an empty one; every
session passed `Validate` with an empty degraded set; no cost, no quota, no sub-agent
count anywhere. First `Discover` 78 ms, second 0 ms (the store had not moved). What this
does **not** cover, itemized: no in-flight session was sampled, no fan-out was observed,
the derived-percentage path did not fire on live data (no real session was missing
`contextUsagePercent` while carrying raw counts), and the corpus is one machine, one day,
one Cursor version.
**Observed 2026-08-17: a live `cursor-agent` CLI session drew no HUD row — and that is the
DESIGN, not a defect.** While the operator drove the interactive session that paid §7.16's
context capture, no HUD row appeared for it. The reason is structural and already on the
record: this adapter reads exactly one store,
`%APPDATA%\Cursor\User\globalStorage\state.vscdb`, which is the **IDE's** Composer store,
and the `cursor-agent` **CLI keeps its own** (§7.16's three-excluded-fields block says so in
as many words, and it is why a `conversation_id` is not stored — it would dangle a join that
does not exist). A CLI session was therefore never a row this adapter could draw. The
observation confirms an existing claim rather than opening a question, and it is written
down because "the gauge showed nothing" is the kind of report that gets re-investigated
every few months unless the expected answer is recorded beside it. **This paragraph
describes the build it was written at and no longer describes the HUD: a CLI session draws
a row as of 2026-08-29. See this section's addendum below.**
**What the same observation does NOT settle, and it is the measurement worth taking.** The
operator's Cursor configuration carries `ghostMode: true` and `privacyMode: 2`. Neither
setting can explain the missing row above, because that row was never expected. They bear on
a different and live question: **whether those settings suppress session writes to the IDE
Composer store this adapter DOES read.** If they do, then every Cursor row in the HUD is
conditional on a vendor privacy setting nobody has measured, and a reader would see an empty
Cursor section with no way to tell "no sessions" from "sessions not written". That is a
zero-versus-absent failure one layer below the renderer, where §4a.1 cannot reach it.
Stated as an open item rather than a finding, because nothing here was measured: the hypothesis
is untested, and it must not be repeated as though it were observed. The measurement is
cheap and needs the IDE rather than the CLI — drive a Composer session in the IDE with those
settings on, then off, and compare what `composerHeaders` gains in each arm. Until somebody
runs it, this paragraph is a hypothesis with its reason attached and nothing in the adapter
changes.
**Amended 2026-08-29 — the CLI store is readable on this machine, and the 2026-08-17 gap is
CLOSED. `cursor-agent` CLI sessions draw HUD rows now.** The paragraph above stays correct
about the build it described. Its premise moved: `cursor-agent` is installed here, and a
2026-08-29 read-only survey of its trees measured a per-session manifest the 2026-08-02
survey could not have seen. `internal/adapter/cursor` reads it, in `chats.go`, beside the
Composer store it already read.
**First, the claim that did NOT survive the survey.** A JSONL transcript tree exists at
`~/.cursor/projects//agent-transcripts//.jsonl`, and a vendor-seam refresh
candidate described it as Claude-compatible JSONL carrying per-turn token counts. Both halves
are false, measured structurally over 71 files and 1,951 records with 0 unparseable. The
envelope carries exactly three key sets — `message`+`role` (1,852), `status`+`type` (87),
`error`+`status`+`type` (12) — and none of `sessionId`, `timestamp`, `cwd` or `version`, so
`internal/adapter/claudecode` cannot be reused: its canary is `sessionId`. A key sweep for
`usage|tokens|input_tokens|output_tokens|totalTokens|cacheRead|cost` over all 71 files
returns zero matches. That CONFIRMS §7.16's 2026-08-08 measurement on a corpus 1.4x larger at
a CLI build one week newer, so §7.16's held display is untouched and cost stays `CapNone` for
this vendor. The transcript tree is not read and no adapter opens it.
**What IS built, and what it measures.** The store is one plain-JSON manifest per session:
```
~/.cursor/chats///meta.json
```
Observed key sets over 43 manifests, at `cursor-agent 2026.08.11-e8db854`: 40 carry
`schemaVersion`, `createdAtMs`, `hasConversation`, `updatedAtMs`, `cwd`; 3 add `title`.
`schemaVersion` is `1` on 43 of 43. **This is the first Cursor surface anywhere that declares
its own format version**, against the "Version fragility" caution above, and the value is
PINNED in `chats.go` and in a fixture.
| Field | Verdict | Source, and what was measured |
|---|---|---|
| name | **MEASURED** | `title`, on 3 of 43 manifests. An absent title is therefore genuine absence, not a failed read. The row then takes the `cwd` basename, which is the fallback `internal/hud`'s own `sessionLabel` applies and `internal/adapter/pi` applies at the adapter. |
| workspace | **MEASURED** | `cwd`, a native path (`C:\...` on 43 of 43 here), taken verbatim. It is NOT the `file:///c%3A/...` URI the Composer store's `workspace.json` carries, so no conversion runs. |
| last_activity | **MEASURED** | `updatedAtMs`, epoch milliseconds, folded with the manifest's own mtime per §6 Q8. |
| model | **ABSENT** | No key of any kind. The Composer store's `composerData` names one; this store does not. |
| context %, cost, quota | **ABSENT** | See the token sweep above. Zero matches in 1,951 records. |
| liveness | **ABSENT** | No in-flight session was sampled, same as the Composer half. The HUD classifies age. |
**The Q8 fold runs here, and that is not a contradiction of the build caution above.** The
Composer half deliberately excludes the store's file mtime because ONE file backs every
session there. The CLI gives every session its own manifest, so that file's mtime dates that
session, and the fold is the ordinary shape every other adapter uses. The two agreed within
0.2 s on 43 of 43 manifests, so the fold is a guard rather than a source: it is what keeps
the age honest if the vendor ever rewrites the file without restamping the key.
**`hasConversation:false` is an empty shell and draws no row. The ruling is measured.** All
three such manifests held nothing but themselves — no `store.db`, no `prompt_history.json` —
and their `updatedAtMs` stood 263–387 ms past their own `createdAtMs`. They are the directory
the vendor stamps when a session is created and nobody types. That is the same class as the
Composer store's `empty-state-draft` and its pre-created composers, so the same filter applies.
An ABSENT `hasConversation` key is NOT read as a declared false and keeps its row: a vendor
that stopped writing the key must not empty the HUD.
**Two more shapes are skipped in silence, and the counts are recorded here so the row count is
not read as a defect.** 22 of the 65 session directories hold `store.db` with no `meta.json` at
all, because the manifest is newer than the tree; a row for one would date a session from a
directory mtime whose meaning was never measured. Together with the 3 shells, 40 of 65
directories drew rows on the survey machine. A manifest whose `schemaVersion` this adapter does
not read is the opposite case and is REPORTED every time: it draws no row (dropfile's rule —
the keys may no longer mean what the reader thinks) and the skip is counted into a diagnostic
that rides on every Cursor row, Composer rows included. A silent skip there would turn a vendor
format bump into "you have no CLI sessions", which is a wrong answer rather than a missing one.
**One adapter, two stores, and every row says which one it came from.** Both stores describe
Cursor sessions, so both feed `model.VendorCursor`; the Adapter composes a sibling reader
rather than shipping a second adapter, because a vendor id is what the HUD's identity column
and the `--vendor` flag address, and two adapters sharing one id would give the registry two
answers. The distinction IS measurable, so it is displayed: every session carries an Extra
labelled `source` naming its store, and a CLI row's id is prefixed `cli:`. The labels are
symmetric — a Composer row carries one too — so that the ABSENCE of a label never becomes the
thing that identifies a store. The alternative was a second vendor id (`cursor-cli`); it was
rejected because it buys one distinction and pays a second vendor line, a second `--vendor`
value and a second doctor pin for what the operator experiences as one tool. The composition
has one honest cost and `chats.go` states it on the row rather than hiding it:
`model.Capabilities` is static per adapter, so a CLI row's empty model and context cells read
as "absent now" when the truthful reading is "this store has no such field". Each CLI row
carries a second Extra, `not in this manifest`, naming those two fields.
**The credential rule is unchanged and narrower here.** The tree sits beside `store.db`
(4096 bytes with a 300 KB–1.7 MB `-wal`, the WAL trap in its strongest form),
`prompt_history.json`, and the config and cache files in `~/.cursor`. The reader opens ONE
file name, `meta.json`, at a fixed depth of two directories below `chats/`. It never opens
`store.db`, never recurses, and never looks at a sibling for any purpose. `chats_test.go`
plants credential-shaped and prompt-shaped markers in those neighbours and asserts none
reaches a displayed field, which is the same standing test the Composer reader carries.
`conversation-search.db` — a new FTS5 index over conversation `body` text that appeared beside
`state.vscdb`, with a `source` column admitting `'cloud-cache'` — stays unopened under the rule
that keeps grok's `session_search.sqlite` closed.
**What this amendment did NOT measure.** `~/.cursor/acp-sessions//meta.json` carries the
same shape (49 manifests, same three key sets, 3 of them shells) and is deliberately NOT read:
those are `cursor-agent acp` sessions, which is the server council's own Cursor seat runs
(§9.36), and whether telltale should draw rows for its own council seats is a separate
question this change does not answer. The `~/.cursor` root is measured on Windows only; see
`PARITY.md`. No transcript, no `store.db` blob and no `ItemTable` row was read.
### 3.9a Grok CLI seam — surveyed live 2026-08-09; the first vendor that writes money down
**Environment:** grok 1.0.0 (3cd0d0cbce), Windows 11, signed in against grok.com, model
`grok-4.5`. Closed source and undocumented — `~/.grok/README.md` is user-facing product
prose, not a format spec — so this is the Claude Code / Cursor method again: a read-only
live survey over the real store, every claim measured, nothing carried over from
`--help`. **Numbered against the sections above rather than after §3.10** so the survey
sits with the surveys and the inventory keeps summarizing everything before it.
The council seat (§9.39) already drives this vendor, and the HUD could not see it at all.
This closes that; the seat's parser and this adapter now read the same wire from two
sides, which is what let the cost unit below be pinned rather than assumed.
**Store inventory.** Sessions are DIRECTORIES, not files:
```
~/.grok/
sessions/
session_search.sqlite a FILE at the root — full-text index
/
prompt_history.jsonl a FILE at workspace level
/ ONE SESSION
summary.json signals.json chat_history.jsonl events.jsonl
updates.jsonl prompt_context.json system_prompt.txt resources_state.json
rewind_points.jsonl announcement_state.json terminal/ *.lock
active_sessions.json auth.json worktrees.db models_cache.json logs/ memtrace/
```
**The variance is a finding, not noise.** Across 30 session directories in 8 workspaces,
only five files were present in all 30: `summary.json`, `chat_history.jsonl`,
`events.jsonl`, `prompt_context.json`, `system_prompt.txt`. `updates.jsonl` was in 29,
`signals.json` in 23, `resources_state.json` in 13, a `terminal/` subdirectory in 5. The
adapter therefore sources every required field from `summary.json` — the invariant — and
treats a field that lives anywhere else as *absent now* when its file is missing, which is
a first-class state here and not a failure.
**Field map.**
| Field | Verdict | Source, and what was measured |
|---|---|---|
| name | **MEASURED** | `summary.json` → `generated_title` (17 of 30), falling back to `session_summary`, which held the identical string on every session carrying both. A headless `--single` run has `session_summary: ""` and **no `generated_title` key at all** — verified by running one — so an unnamed grok row is absence, not a failed read. |
| model | **MEASURED** | `summary.json` → `current_model_id`. `grok-4.5` on 30 of 30. Note the per-model usage blocks key off `grok-4.5-build` instead; the adapter reports the id the vendor puts in the session's own model field, not the billing variant. |
| workspace | **MEASURED** | `summary.json` → `info.cwd`, absolute and native (`C:\Users\sanle\code\telltale`), on 30 of 30. The parent directory name carries the same path percent-encoded and is the **fallback** — see the round-trip note below. |
| context % | **MEASURED** | `signals.json` → `contextWindowUsage`, an integer percentage the vendor computes, beside the raw `contextTokensUsed` / `contextWindowTokens` (500000 on every session). It TRUNCATES: 39656/500000 = 7.93 is written `7`, 22675/500000 = 4.535 is written `4`. The adapter reports the vendor's integer and does not recompute — grok is the second vendor after Cursor with no assumed denominator, and the more precise float would be a number the vendor never said. These three keys are still the source, and `signals.json` holds 61 more beside them — see the 2026-08-29 drift census below. |
| cost | **PER-TURN ONLY, so the field stays `CapNone`** | `updates.jsonl` → each `turn_completed` record's `usage.costUsdTicks`. The unit is measured twice over (below). It is **not cumulative**: one session's three turns read 455412000, 820464000, 747416000 ticks, the third smaller than the second. `"[a-z_]*cost[a-z_]*"` over every `.json`/`.jsonl` in the store matched `costUsdTicks` and **nothing else** — no session total exists anywhere. Summing needs every turn record and `updates.jsonl` reached **818 KB** in one session, past any bounded tail; a tail-window sum is a lower bound, and a lower bound in a column headed COST is a derived number wearing a read one's clothes. The **last turn's** cost is carried as a labeled Extra instead. **This row named one key and the record carries eleven** — see the 2026-08-29 re-measure below, which is a correction to this row's completeness and not to its verdict: the sweep spelled here could not match `inputTokens` and its siblings. |
| quota | **ABSENT** | `"[a-z_]*(rate\|limit\|quota)[a-z_]*"` over the whole store returned only tool-configuration keys (`output_byte_limit`, `head_limit`). No window, no ordinal, no reset time reaches disk. That is a statement about the *disk*; the network half of the same question is measured and closed separately below. |
| last_activity | **MEASURED** | `summary.json` → `last_active_at` (29 of 30) then `updated_at` (30 of 30) then `created_at`, folded with the file's mtime per §6 Q8. `summary.json` is rewritten every turn, which is also why it — not the session DIRECTORY, whose mtime moves only when an entry is added — is the freshness hint `Discover` returns. |
| liveness | **ABSENT, and this one was probed rather than reasoned about** | See below. |
| sub-agent count | **ABSENT** | grok ships a `spawn_subagent` tool (it is in the tool list every headless run prints), and `"subagent[A-Za-z_]*"` matched nothing on disk outside the system prompt's own description of it. No nest, no count, no parent link. Same ruling as Codex (§3.3): declaring the field and emitting zero would assert something the format cannot check. **The sweep spelled here read file contents and a `subagents/` DIRECTORY exists — see the 2026-08-29 drift census below.** The verdict is unchanged and its stated reason is not. |
**The cost unit, pinned twice.** `costUsdTicks` is fixed-point USD at **1e10 ticks to the
dollar**, and neither half of that was inferred from the name. First, grok's headless wire
prints both forms of the same number on its `end` event, and three live runs on this box
on 2026-08-09 gave `0.0306488 / 306488000`, `0.0315248 / 315248000` and
`0.0382104 / 382104000` — exactly 1e10, three times. Second, the disk field is spelled
differently (`costUsdTicks` vs the wire's `total_cost_usd_ticks`), so it had to be shown to
be the same quantity and not merely a similarly named one: those three runs' on-disk values
were read back and matched the wire's tick counts **value for value**. Without the second
step this would be a plausible unit rather than a measured one.
**`active_sessions.json` claims liveness and does not deliver it — measured.** With 30
sessions in the store the file held the two bytes `[]`. It still held `[]` **while a
headless turn was mid-flight**, sampled with `grok.exe` confirmed running by PID and with
the file's own mtime freshly stamped by that run — so the vendor had written the file and
written nothing in it. A registry that is empty during a live session cannot tell "nothing
is running" from "the thing running is not the kind it tracks", and §4a.4 already names
process-existence as the one case where an adapter can lie to the HUD undetectably. Not
read.
`events.jsonl` carries the other tempting signal — `phase_changed` (1765 of them in one
session, spelling `waiting_for_model`, `streaming_reasoning`, `streaming_text`,
`tool_execution`, `permission_prompt`), `turn_started` / `turn_ended` with an `outcome`,
and a `permission_requested` that is a genuine needs-input state. It is left unread in v1
for the reason the corpus itself demonstrates: the newest session ends on an **unresolved
`permission_requested`** written minutes before grok exited, so "the last event is a
prompt" and "a dead session was killed at a prompt" are the same bytes. A hint that stays
true forever after the process is gone is worse than no hint. This is the strongest
needs-input seam any vendor has offered so far and it is recorded here as the watch item,
not spent.
**The percent-encoding round-trips, and that is why this adapter may decode it.**
`C%3A%5CUsers%5Csanle%5Ccode%5Ctelltale` is `C:\Users\sanle\code\telltale`: `:` as `%3A`,
`\` as `%5C`, with letters, digits and a literal `-` passing through unescaped. Every one
of the 8 workspace directory names decoded to exactly the `info.cwd` its sessions recorded,
drive-letter case included. That is the opposite of §3.1's ruling for Claude Code, whose
project slug maps both `\` and a literal `-` onto `-` and is therefore lossy — decoding
*that* would invent a path. Grok's encoding is injective, so decoding it invents nothing,
and the adapter uses it as the workspace fallback. It stays a fallback: the vendor's own
record of its cwd outranks a key we reconstructed.
**What is deliberately not opened, and why it is stated rather than assumed:**
- **`~/.grok/auth.json`** holds the OAuth token. The adapter resolves no path outside the
sessions tree at all.
- **`sessions/session_search.sqlite`** is an FTS5 index whose `session_docs` table carries
`(session_id, cwd, updated_at, title, content, content_hash)` — **`content` is transcript
text**. It would answer "what is this session about" in one query, and that is exactly the
trade this repo does not make: `summary.json` already has the vendor's own label.
- **`prompt_context.json`** inlines the user's `CLAUDE.md`/`AGENTS.md` verbatim (39 KB on
one session). Same rule. A test plants a marker in both files and asserts nothing carrying
it reaches any displayable field, the way §3.9's credential allowlist is enforced.
- **The `.lock` sidecars** beside `summary.json`, `chat_history.jsonl`, `updates.jsonl` and
`rewind_points.jsonl` are never opened and never created. The gauges read.
**The quota question has a network half, and it is closed three times over — probed
2026-08-09**, the same local day as the survey above (the HTTP `Date` headers quoted below
read 2026-08-10 UTC; this box runs UTC−4). The table's `quota` row says nothing reaches
*disk*, which invites the obvious follow-up: the CLI talks to a server, so ask the server. `~/.grok/README.md` ("Using
auth.json for API Access") even documents the call. This block exists so the next person to
ask stops here instead of re-deriving it. Three findings, then three rules.
*The documented recipe does not run on this build.* Its `jq` path
`."https://accounts.x.ai/sign-in".key` matches nothing in this box's `auth.json`, whose one
entry is keyed `https://auth.x.ai::` (`auth_mode: "oidc"`, a six-hour `expires_at`);
sent as written it returns **401**, `www-authenticate: … reason=no auth context`. And
`POST /v1/chat/completions` gates on a header the recipe never mentions — omit it and the
proxy answers **426 Upgrade Required**, body `"Your Grok CLI version (none) is outdated"`.
The header is `x-grok-client-version`, read out of `grok.exe`'s string table beside
`x-grok-client-identifier`, `x-grok-client-mode`, `x-grok-session-id` and the documented
`x-grok-model-override`. The README is product prose here too, exactly as this section's
header warns.
*The free half of the question is answered, and the answer is no.* `GET /v1/models` — the
same request whose result the CLI already caches to `models_cache.json`, so it bills
nothing — returns **200** carrying `etag`, `strict-transport-security`, `cf-cache-status`,
`CF-RAY`, `alt-svc` and `Server: cloudflare`, and **not one `x-ratelimit-*` header of any
kind**. The one proxy response this vendor already persists has no quota in it.
*The billed half is `not checked`, which per §9.42 carries a reason and never a value.*
Whether `POST /v1/chat/completions` returns those headers was **not measured**: it is the
only probe here that spends a turn, and it stopped at this machine's own tool-permission
boundary rather than being run. Two attempts to get it from the vendor's own logging failed
for a reason worth recording, because it looks like evidence and is not: `~/.grok/logs/
unified.jsonl` is **not an HTTP log** — every record's `src` is `shell` or `grok-pager` — so
its silence about rate-limit headers was never evidence about the wire, and a live
`grok -p` run driven with `--debug-file` produced no file at all.
*The disk sweep was re-run with a case gap closed.* The original survey's
`"[a-z_]*(rate|limit|quota)[a-z_]*"` is **lowercase-only** and would have missed a camelCase
`rateLimit` — which matters precisely here, because this is the vendor that writes
`contextWindowUsage` and `contextTokensUsed`, so camelCase is its house style and the
original regex had a live blind spot. Re-swept as
`"[A-Za-z_-]*([Rr]ate[Ll]imit|[Qq]uota|[Rr]emaining)[A-Za-z_-]*"` over the whole sessions
store, including a session directory created by a fresh live turn: the sole hit is
`agents_remaining` (17 occurrences), a **sub-agent budget**, not an account one. The
`quota: ABSENT` verdict survives the stronger sweep.
*And the vendor's own monitoring surface settles it.* `docs/user-guide/24-monitoring-usage.md`
documents an external OpenTelemetry stream — the one place grok is *designed* to report what
an account is doing — and its attribute keys are **a closed enum**, with an export-time
validator that drops any record carrying a key outside it. What it carries is
`grok_code.token.usage` (`input` / `output` / `reasoning` / `cache_read`, by model) and a
`grok_code.api_request` event with the same four counts. What it carries **nowhere** is a
window, a reset, a remaining percentage or a limit of any kind; `subscription tier` sits on
its explicit never-exported list, and the doc states outright that there is no cost metric
("join `grok_code.token.usage` with your own price sheet"). So the absence is not an
oversight in a session file. The vendor built a schema for exactly this question, put
**spend** in it, and put **no quota in it at all** — which is §7.15 versus §7.16 in the
vendor's own hand: a count with none, never a reading against a limit.
That stream is also the honest answer to "then what *could* be measured". If a grok **spend**
row is ever wanted it is the seam — exported to a local collector whose output telltale would
read as a *file*, leaving §4a.5's no-network-calls contract intact, since the push is grok's
and not ours. It is the cursor relay's shape (§7.16), and it is emphatically not a quota
window. Cost: a double opt-in, an endpoint, and a collector to run; the `console` exporter is
no shortcut because the doc says it is suppressed in the `agent` and `headless` entrypoints.
Recorded as the seam, not spent.
**The seam was spent on 2026-08-10 — `telltale otel grok` is the collector (§7.16a).**
Before it was built, the doc's schema table was checked against the wire the way this
section's header demands: a dump collector on 127.0.0.1:4318, and a live headless
`grok -p "hi"` from grok 1.0.0 (3cd0d0cbce) with the double opt-in set. What arrived,
versus what the doc says:
- Transport as documented: OTLP http/protobuf POSTs to `/v1/logs` and `/v1/metrics`,
`Content-Type: application/x-protobuf`, uncompressed, from `OTel-OTLP-Exporter-Rust/0.32.0`.
The batches flushed before the headless process exited, at the default export intervals —
a short `-p` run loses nothing. The fleet-policy startup suppression the doc warns about
was not observed to delay anything on this signed-in box.
- `grok_code.api_request` events carry `input_tokens`, `output_tokens`, `reasoning_tokens`
and `cache_read_tokens` as int attributes — all four on one record — beside `model`,
`duration_ms`, `stop_reason`, `session.id`, a per-session monotonic `event.sequence`,
`user.id` and `team.id`. The event name arrives in the LogRecord's `event_name` field,
not as an attribute.
- The `grok_code.token.usage` metric (delta temporality, by `model` and `type`) carried
**the same four counts value-for-value** as the same turn's api_request event:
20323/56/42/2560 on both sides of one capture. One number, two envelopes.
- `turn_completed` **on the stream** carries outcome and duration and **no token counts**, as
the table says. The qualifier was added 2026-08-29 and it matters: the record of the same
name that grok persists to `updates.jsonl` carries nine of them, which is the re-measure
block below. One event name, two envelopes, opposite contents.
- One departure from the doc's letter: `OTEL_METRICS_INCLUDE_SESSION_ID` defaults on, but
the token.usage data points carried no `session.id` — only `session.count`'s did. Nothing
here reads metrics, so nothing turns on it; recorded because the doc says otherwise.
The quota verdict above is unchanged by any of this: nothing on the stream carries a
window, a reset or a limit. What was added is a **spend count** (§7.16's vocabulary — a
count with no denominator), accumulated per api_request event into
`~/.telltale/usage/grok.json`, display held. §7.16a is the design record.
A measurement of the completions headers would not move the verdict, because three separate
rules already close this and none of them turns on whether the number exists:
1. it needs `auth.json`, and this adapter resolves no path outside the sessions tree (above);
2. it needs a network call, and §4a.5's adapter contract is explicit that implementations
"must not write to vendor state, and must not make network calls or read credentials";
3. §9.42 draws the probing line at **cost and side effect** — `doctor` may run `--version`
precisely because it starts no turn and bills nothing. A quota probe against the
completions endpoint would spend from the very pool it is trying to read.
And the quantity would be the wrong one regardless. `x-ratelimit-remaining-requests`,
`-remaining-tokens` and `x-ratelimit-reset-requests` are documented for **`api.x.ai`**, the
metered developer API. This CLI rides `cli-chat-proxy.grok.com` on an OIDC session token
against a **SuperGrok subscription**, whose binding limit is a shared weekly pool surfaced
only in grok.com's own Settings → Usage. An RPS/TPM header is not the window the usage pane
means by quota: §4a.3's window carries a *length* and a `ResetsAt`, and a per-second request
cap has neither. Relabelling one as the other would be a duration claim with no source —
the exact move §4a.3 already forbids — and would land a real number on screen answering a
question nobody asked.
**Two smaller findings worth writing down.** One session's `summary.json` carries
`"sandbox_profile": "bogus-profile-xyz"` — the invalid profile §9.39 fed the CLI to prove
`--sandbox` validates nothing. grok not only accepted it, it **persisted it**, which is why
that key is not rendered as an Extra: it would put an unvalidated word on screen. And
`summary.json` carries a git block (`git_root_dir`, `git_remotes`, `head_commit`,
`head_branch`) on exactly the one session whose cwd was a git repo — a branch column is
available for later, and is out of scope here.
**The frame, generated by the build** (`internal/hud/testdata/golden/grok-row.txt`, at
120 columns; both rows are synthesized). The COST column is not narrow here, it is
ABSENT — the vendor writes dollars and no row claims a session total. The first row's
bar carries no estimate marker because the percentage was read rather than derived; the
second is a headless run with no title and no `signals.json`, so its label falls back to
the workspace and its CONTEXT is an em dash rather than a zero:
```
telltale │ 2 sessions │ grok 2
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL CONTEXT AGE
● GR │ Adapter Field Map Review C:\src\code grok-4.5 ▊─────────── 7% │ 20s
◐ GR │ example-app C:\src\code grok-4.5 — │ 6m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys
```
**Adapter built, same day (`internal/adapter/grok`).** Name, model, workspace, context %
and last_activity REPORTED; cost, quota, liveness and sub-agent count `CapNone`; the
derived set deliberately EMPTY. `summary.json` and `signals.json` are read whole under a
64 KB cap each (largest observed: 918 and 1538 bytes) and a file past its cap degrades the
fields it feeds rather than being slurped; `updates.jsonl` is tail-read like every other
JSONL here.
**Live verification, same day** (`go test ./internal/adapter/grok -tags=live -run
TestLiveGrokStore`, which drives the HUD's own `Scan` and `Render` over the real store and
is excluded from CI because CI has no store to read): **31 sessions discovered and read**,
**31 of 31** sourcing both a model and a workspace, 18 titled, 24 carrying a context
reading, 24 carrying a turn cost. Every session passed `Validate`; not one produced a cost,
a quota window, a sub-agent count or a liveness hint. The frame showed the zero-vs-absent
distinction doing real work: the 7 sessions with no `signals.json` rendered an em dash in
CONTEXT while their neighbours rendered an unmarked bar, on the same screen. First
`Discover` plus 31 `Read`s completed in 80 ms.
What this does **not** cover, itemized: no session was sampled while a turn was streaming
(the `active_sessions.json` probe ran against a headless turn, not against a HUD scan); no
interactive TUI session was running during a scan, so nothing exercised the NTFS-deferred
mtime the Q8 fold exists for; the summary-past-its-cap path has never fired on real data;
the corpus is one machine, one day, one grok version; and `active_sessions.json` has never
been observed non-empty **at all**, so "it does not track headless sessions" is the most
this survey can say about what it would otherwise contain.
#### `active_sessions.json` re-measured 2026-08-15 — the registry populates now, and liveness stays `CapNone`
**Environment:** grok 1.0.4 (d846eb93d9), the same Windows 11 box, model `grok-4.6`, cwd
`C:\Users\sanle`. The paragraph above is the reason this ran. The 2026-08-09 survey never
saw the file hold anything, so it could not say what a populated registry would mean. The
file holds records now, and a record with a **dead pid** was observed surviving overnight.
That one fact makes the question answerable, so it was measured rather than argued. The
measurement cost **3 grok model turns**; every other step below is a process start, which
calls no model.
**The record shape, from two sides.** The live file holds a JSON array of objects with four
keys, and the binary at this pinned build names the same four: `struct ActiveSession` with
`session_id`, `pid`, `cwd`, `opened_at`, written by `crates/codegen/xai-grok-active-sessions/src/lib.rs`
through `active_sessions.lock`, `active_sessions.json.tmp` and `active_sessions.json`. A
live record read verbatim:
```json
[
{
"session_id": "01a0084f-e3b1-7ff2-be5c-2afaed6a6581",
"pid": 17096,
"cwd": "C:\\Users\\sanle",
"opened_at": "2026-08-16T02:04:08.999641600Z"
}
]
```
`session_id` is the session directory's own UUID, so the record joins to §3.9a's store
directly. `pid` is the grok.exe process, confirmed against `tasklist` at the moment of
capture. `cwd` is the workspace in native absolute form, and it equals the `info.cwd` the
same session writes. `opened_at` is UTC RFC3339 and stamps the OPEN, not the process start:
a resumed session gets a new `opened_at` and the same `session_id`. The vendor also names
the failure mode itself, verbatim from the binary: `active_sessions.json is corrupted,
starting with empty list`.
**The four questions, each measured.**
| Question | Answer | Evidence |
|---|---|---|
| Does a session start gain an entry, and is the pid real? | **Only when a session OPENS, and the pid is real.** | A fresh interactive session registered 0.7 s and 1.3 s after the first prompt was submitted, on two runs, and never before it. `pid 17096 -> ALIVE as grok.exe` at capture. A `--resume` registered 1.6 s after launch with **no prompt at all**, which is how the rest of this block was measured for free. |
| Does a CLEAN exit reap the entry? | **Yes, immediately.** | `/exit` in the TUI returned the file to the two bytes `[]` within 0.3 s of the command, with its mtime freshly stamped. The process removes its own record on the way out. |
| Does a KILLED process leave the entry behind? | **Yes, and the dead pid persists.** | `taskkill /T /F` on the pid the registry itself named left the record byte-identical, mtime unmoved, for the whole sampling window. `pid 17096 -> DEAD` while the record still claimed it. Reproduced twice, and it confirms the overnight observation deliberately. |
| What are the exact fields and their meanings? | The four above. | Live file plus the pinned binary's own struct, which agree. |
**The lifecycle, stated as write points.** Four events touch the file, and only these four
were observed to. A session OPEN appends its record. A clean exit removes that session's
record. Any grok agent start **rewrites the file and drops records whose pid is dead**. A
kill removes nothing, so the record outlives the process until the next start sweeps it.
**The sweep is pid-aware, not a truncation, and that was measured separately.** One stale
record and one live session were put on disk together, then a second grok process was
started. The stale record disappeared and the live record survived unchanged, so the sweep
reads each pid rather than clearing the list. This also explains the overnight persistence
with no contradiction: nothing swept the record because no grok agent started in between.
**What does NOT register, measured one shape at a time.** Each of these ran with the file
sampled every 0.3 s throughout, and each left it at `[]` while the process was alive:
`grok -p ` (a full headless turn, which is the 2026-08-09 result reproduced at
1.0.4), `grok agent stdio`, `grok agent leader`, and an interactive TUI sitting at its
input box with a session directory already created and no prompt yet sent. Every one of
those starts moved the file's mtime, so the vendor wrote the file and put nothing in it,
exactly as the original survey found. `grok --version` does not touch the file at all.
**Verdict: liveness stays `CapNone`, and §4a.4's ruling is not falsified.** A populated
registry is not a liveness source, because both directions fail:
- **Presence proves nothing.** A record can name a dead pid for an unbounded time. The
bound is the next grok agent start, which is an event telltale cannot predict and must
not cause. An adapter that read presence would report a session as live all night.
- **Absence proves nothing.** Four live session shapes register nothing at all, and a
headless turn is one of them. An adapter that read absence would report a running turn as
gone.
§4a.4 rules process-existence out for every vendor, and it named Claude's registry only as
the first instance. A grok registry that populates is a different fact from a grok registry
that is trustworthy, and only the second one would touch that ruling. It does not.
**One usable shape falls out, and it is recorded rather than spent.** The registry supplies
a `session_id` to `pid` binding, and a dead pid is checkable directly. So a stale
`permission_requested` in `events.jsonl` (the needs-input seam §3.9a already recorded and
declined) **can be NEGATED by a dead pid, and can never be ASSERTED by a live one**. The
asymmetry is the whole content of the finding, and it survives its own caveats in one
direction only. A recycled pid reads ALIVE, which withholds a negation rather than
inventing one, so the failure is conservative. A session that never registered cannot be
negated, which costs coverage and claims nothing false. Nothing is built on this. The
measured-silence advisory that would consume it sits on the post-demo shelf by owner
ruling, and this block exists so that work starts from a measurement instead of a memory.
**What this does not cover.** Every registering session ran in one cwd on one box at one
build. No two sessions were ever registered at once, so the multi-record ordering is
unmeasured. `active_sessions.lock` was never opened and `active_sessions.json.tmp` was
never observed mid-write, so the write is atomic by the vendor's own naming rather than by
observation here. The adapter reads none of these bytes, and this measurement did not
change that.
#### The `usage` object re-measured 2026-08-29 — the token counts were always beside the cost, and this section's own sweep could not have seen them
**Environment:** grok 1.0.5 (5115b46bc9) [stable], the same Windows 11 box, model
`grok-4.6`. One billed headless turn (`grok --output-format streaming-json --single=…`)
from a fresh empty workspace, with the session's `updates.jsonl` read back and
cross-checked against the same turn's wire output. **This block corrects the section
above rather than reporting vendor drift, and that is the point of it: the miss was
telltale's instrument, not the vendor's format.**
**What the cost row says, and what the record actually carries.** The row names
`usage.costUsdTicks` and nothing else, because the sweep behind it was
`"[a-z_]*cost[a-z_]*"` — lowercase, and anchored on the substring `cost`. A key spelled
`inputTokens` matches neither half. The regex answered the question it was asked and was
**structurally incapable** of the question a reader takes the row to answer. The
`turn_completed` record's `usage` object carries **eleven keys** — nine scalar counts, a
per-model breakdown of the same nine, and a turn count — read verbatim off the measured
turn:
```json
"usage":{"inputTokens":22772,"outputTokens":27,"totalTokens":22799,
"cachedReadTokens":256,"cacheCreationTokens":0,"reasoningTokens":22,
"modelCalls":1,"apiDurationMs":3481,"costUsdTicks":77047400,
"modelUsage":{"grok-4.6-build":{ the same nine scalars, per model }},
"numTurns":1}
```
**It is not drift, and that was checked rather than asserted.** A session directory
written 2026-08-09 at 09:35 local — the same day, the same build (grok 1.0.0) and the same
model (`grok-4.5`) this section surveyed — carries the identical key set, `numTurns`
included: `"inputTokens":22222, "outputTokens":37, "totalTokens":22259,
"cachedReadTokens":128, "cacheCreationTokens":0, "reasoningTokens":29, "modelCalls":1,
"apiDurationMs":2605, "costUsdTicks":444484000`. The counts were on disk before the
adapter was written. The vendor changed nothing; the survey looked with an instrument that
could match one key and reported one key.
**The wire and the disk spell `input` differently, and the difference is exactly the cache
read.** The same turn's `end` event on `--output-format streaming-json` reports
`input_tokens: 22516`, `cache_read_input_tokens: 256`, `output_tokens: 27`,
`reasoning_tokens: 22`, `total_tokens: 22799`. The disk reports `inputTokens: 22772`,
which is 22516 + 256. Both envelopes agree on 22799 and reach it by different arithmetic:
**the wire's `input_tokens` EXCLUDES the cache read; the disk's `inputTokens` INCLUDES
it.** So `inputTokens + cachedReadTokens` on disk double-counts, and the two seams' "input"
columns are not the same statistic. §7.16a's collector reads the OTLP wire, whose
`input_tokens` follows the wire spelling — anything that ever holds both must convert
rather than add.
**`costUsdTicks` re-pins at 1e10 on the same record.** The wire printed
`total_cost_usd: 0.00770474` beside `total_cost_usd_ticks: 77047400`, and the disk record
carries `costUsdTicks: 77047400` for that turn. The 2026-08-09 unit measurement holds at
1.0.5, value for value, on a fourth run.
**The counts are per-turn and they do not accumulate — measured on a two-turn session.**
Its two `turn_completed` records read `inputTokens 21548 / totalTokens 21588` then
`inputTokens 21958 / totalTokens 21999`. The second is not the first plus anything; each
record describes one API call, and `inputTokens` grows between turns because the context
is resent, not because a counter advanced. `numTurns` is `1` on **every** record, both
turns included, so it counts the turn the record describes and is not a session ordinal.
That is the same shape as `costUsdTicks`, measured again in a second unit.
**What this changes, and what it does not.**
- **The cost row's verdict stands unchanged, and it now covers nine counts instead of
one.** A session total needs every record, `updates.jsonl` reached 818 KB in one observed
session, and a tail-window sum is a lower bound. A lower bound in a TOKENS column is the
same derived number in a different unit.
- **The seam map is what moves.** §7.16a opens with grok's OTLP stream as the vendor's one
designed-for-reporting surface, and that sentence is still true about *designed
reporting*. It is no longer true that the stream is the only place telltale **could** read
grok's per-turn counts. The disk is a second source, and a passive one: no collector
running, no double opt-in, no port held.
- **That second source creates a constraint, recorded here before anything is built on
it.** §7.16a's replay guard exists because one number arriving twice must be counted once,
and the same rule already chose the event stream over the redundant metric. A disk reader
would be a **third** envelope for the same turn, and the guard cannot see it: the
collector's guard lives in collector memory and keys on `(session.id, event.sequence)`,
while a disk reader would key on a file offset. **A disk reader and the OTLP listener must
never both count one turn into `~/.telltale/usage/grok.json`.** One envelope per turn is
the rule. Which envelope is the design question, and this block does not answer it.
**Nothing is built on this, deliberately.** The DISPLAY of grok spend is HELD by owner
ruling (§7.16's amendment, applied in §7.16a), so a reader has no consumer to feed, and a
reader design would first have to settle the one-envelope question above, the read-budget
question the cost row raises, and the `input` spelling. This block exists so that work
starts from a measurement instead of from a row that could not see it.
#### Session-file drift, censused 2026-08-29 at grok 1.0.5 — `signals.json` is 64 keys, and there is a `subagents/` directory
Read-only census over the whole sessions tree: **108 session directories in 27
workspaces**, against the 30 in 8 this section surveyed. Nothing below changes a verdict,
and two entries say the *description* above is narrower than the store.
**The inventory's invariant holds.** `summary.json` is present on 108 of 108, which is the
one file every required field is sourced from. `chat_history.jsonl`, `events.jsonl`,
`prompt_context.json` and `system_prompt.txt` are on 107, `updates.jsonl` on 105,
`signals.json` on 95, `resources_state.json` on 73.
**`signals.json` carries 64 keys, uniform across all 95 files that have one** — measured by
parsing every one and comparing key sets, which returned a single set. The context row
above names three of them (`contextWindowUsage`, `contextTokensUsed`,
`contextWindowTokens`); all three are still present and still the adapter's source, so no
code is wrong. The description is: a fixture synthesized to a three-key file models a file
that does not exist. The other 61 are session statistics — `turnCount`, `toolCallCount`,
`errorCount`, `compactionCount`, `totalTokensBeforeCompaction`, `sessionDurationSeconds`,
`modelsUsed`, `primaryModelId`, plus latency and line-count blocks. **No cost key and no
quota key is among the 64**, so the quota verdict survives one more sweep.
**Fifteen entry names the inventory above does not list**, with the count of session
directories carrying each: `title_refresh_idx` (19), `last_recap_main_turn` (11),
`recap_requests` (11), `hunk_records.jsonl` (9), `subagents` (6), `web_fetch` (5),
`compaction` (2), `compaction_checkpoints` (2), `compaction_requests` (2), `mcp` (2),
`plan.json` (2), `plan_mode.json` (2), `outputs` (1), `plan.md` (1), `workflows` (1). The
adapter reads none of them. They are recorded because a presence-census fixture built from
this section's ten names would call every one of them unexpected.
**One of the fifteen contradicts a verdict's REASON, and it is the cost regex's mistake a
second time.** The field map rules sub-agent count ABSENT because `"subagent[A-Za-z_]*"`
matched nothing on disk outside the system prompt's own tool description. That sweep read
file *contents*, so it could not match a **directory name** — and `subagents/` is a
directory. Six sessions carry one; the largest holds 18 child directories, each named with
a UUID, each holding `meta.json` and `output.json`. `meta.json`'s key set, read as keys and
not values, is `subagent_id`, `subagent_type`, `parent_session_id`, `child_session_id`,
`child_cwd`, `status`, `turns`, `tool_calls`, `duration_ms`, `started_at`, `completed_at`,
`effective_model_id`, `effective_context_source`, `description`, `prompt`. **This is not
drift either**: the largest `subagents/` tree is dated 2026-08-09, the day of this survey.
**The verdict is NOT changed here, and the reason is a measurement nobody has taken.**
`CapNone` on sub-agent count refuses to assert "this session is running no sub-agents", and
a directory that is present on 6 of 108 sessions does not yet establish that its absence
means zero rather than not-yet-written — which is §4a.1's zero-versus-absent question, and
answering it needs a live run that spawns one and watches the directory appear. Two further
facts a future lane must start from: every child UUID observed is **also a top-level
session directory** in the same workspace, so the HUD already lists sub-agents as
independent rows and a count would need to decide whether they stay listed; and
`meta.json`'s `description` and `prompt` are user content, so they fall under the same rule
that keeps `prompt_context.json` closed. Recorded as the seam, not spent.
### 3.9b Pi seam — source-surveyed then live-verified 2026-08-16, `pi 0.84.1`
**Environment and evidence class.** This section started in the §3.2 class (researched
from source, NOT live-verified) and has since been measured against a live corpus on the
Windows reference box. The heading and this paragraph said "not installed here, nothing
live-verified" until 2026-08-16; that is no longer true, and the verdict table below
carries what the bytes actually said. The source read came first, and it is still the
reason the shape was known before any file was opened. That read was the writer's own
code — `packages/coding-agent/src/core/session-manager.ts`,
`src/config.ts`, and `packages/ai/src/types.ts` — in **`earendil-works/pi`**, default
branch on 2026-08-16, newest release `v0.84.2` (2026-08-14). Both former names
(`badlogic/pi-mono`, `earendil-works/pi-mono`) redirect there. The npm package
`@mariozechner/pi-coding-agent` stopped at 0.73.1 in May 2026; current distribution is
the `@earendil-works` scope and compiled Bun binaries. The §3.7 lesson predicted that the
first live pass would falsify something here. It did: `responseModel` (see the verdict
table) does not exist on any assistant message in the live corpus. That is what the
prediction was for, and the falsified claim is corrected in place rather than deleted.
**Why the survey ran when it did.** Pi had no HUD coverage at the time, and §9.1 already
rejected it as a council *seat* under the harness-re-host class. This section is the
adapter half: "we looked", so a future session does not start from "nobody looked". The
coverage gap it names is closed; the section is kept as the evidence the adapter stands
on, not as an open question.
**Store inventory (from source, not from a listing).**
```
~/.pi/agent/
sessions/
----/ one directory per workspace: the cwd with leading
separator dropped and / \ : each replaced by "-"
_.jsonl one session; ts = ISO timestamp with : and . -> "-"
auth.json provider API keys / OAuth tokens — NEVER read
settings.json models.json themes/ tools/ prompts/ bin/
```
Two relocation facts an adapter must pin: `PI_CODING_AGENT_DIR` and
`PI_CODING_AGENT_SESSION_DIR` override the paths, and the package's `piConfig`
(`name`, `configDir`) can rebrand the whole product — a rebranded build stores under a
different dot-directory entirely. The adapter would pin the default `.pi` and say so.
**Format.** JSONL, session `version` 3 (`CURRENT_SESSION_VERSION`). The first record is
a header: `{type:"session", version, id, timestamp, cwd, parentSession?}`. Every other
record carries `{type, id, parentId, timestamp}` — **entries form a tree, not a list**.
Pi branches within a session, so conversation order is a `parentId` walk from the leaf,
not file order. Entry types: `message`, `thinking_level_change`, `model_change`,
`compaction`, `branch_summary`, `custom`, `custom_message`, `label`, `session_info`.
One write-path quirk: a session is not flushed to disk until it holds an assistant
message, so an abandoned prompt may never produce a file.
**Field verdicts.** The prospect column this table shipped with is now a verdict column,
because the live pass ran. The corpus is small and its size is part of every verdict:
**4 sessions, 689 records, 197 assistant messages, `pi 0.84.1`.** A verdict of "present"
means the field was found on every record that should carry it; a verdict of "absent"
means a grep over that corpus returned zero matches, which is the §3.9a standard and not
the same as "the source did not mention it".
| Field | Verdict (live, `pi 0.84.1`) | Source in the format |
|---|---|---|
| name | **absent in this corpus** — no `session_info` entry appears in any of the 4 sessions, so the fallback is the operative path, not the exception | `session_info` entry's optional `name` (user-set). Absent is a state; the workspace-basename fallback (§3.7 precedent) applies. |
| model | **present** — 5 `model_change` entries; `model` on 197/197 assistant messages. `responseModel` is on **0/197**: the source read claimed it and the corpus falsifies it | `model_change` entries carry `provider` + `modelId`; assistant messages also carry `model`. |
| workspace | **present** — `cwd` on 4/4 headers | header `cwd`, verbatim. The directory slug is the fallback, as in §3.7. |
| last_activity | **present** — `timestamp` on 685/685 non-header records | newest entry `timestamp` (ISO), folded with mtime per §6 Q8; assistant and toolResult messages also carry unix-ms timestamps. |
| tokens | **present** — `usage` on 197/197 assistant messages, with all six documented keys | **every assistant message writes `usage`**: `input`, `output`, `cacheRead`, `cacheWrite`, optional `reasoning`, `totalTokens`. |
| cost | **present per message, absent per session**, with the §3.9a ruling attached — `usage.cost` on 197/197, and no session total anywhere | the second vendor that writes money down, and it writes **dollars, per message**: `usage.cost {input, output, cacheRead, cacheWrite, total}`. No session total exists on disk — the TUI sums in memory (`usage-totals.ts`). A summed total is a derived number wearing a read one's clothes (§3.9a), and the tree sharpens it: all-entries and active-path totals differ. The honest carry is the last message's `cost.total` as a labeled Extra, grok-style. |
| context % | **CapNone, now measured** — a grep for `contextWindow`/`maxTokens`/`context_window` over the corpus returns zero, and this box's `~/.pi/agent/models-store.json` is an empty object, so the denominator is in neither place | a numerator prospect exists (last assistant `usage` occupancy) but the denominator is nowhere in the session file — it lives in Pi's shipped model catalog, which is the §3.8 1048576 trap again. |
| quota | **CapNone, measured absent** — the §3.9a grep now has its corpus: `quota`/`rateLimit`/`resetsAt`/`plan`/`subscription` return zero matches | Pi is bring-your-own-key, so structural absence is expected, and the corpus confirms it rather than assuming it. |
| sub-agents | **absent in this corpus** — `parentSession` on 0/4 headers, so the link was never exercised here; the field's existence is still a source claim only | header `parentSession` is a child→parent link (the inverse of agy's structural nesting) — countable only by scanning siblings for parents. |
| liveness | **CapNone, measured absent** — a grep for `pid`/`lock`/`heartbeat`/`isActive` returns zero | nothing in the session files, matching the source read. |
**What this corpus does not cover.** Four sessions from one box, one operator, one day
(2026-08-11), all on the default `.pi` directory. Five of the nine documented entry types
(`compaction`, `branch_summary`, `custom`, `custom_message`, `label`) never appear, so
their shapes remain source claims. No session was observed mid-write, and no two Pi
sessions ran at once, so the tree's branching behavior is documented but not watched. A
`session_info` entry and a `parentSession` header are the two things a wider corpus would
most likely add.
**The credential boundary is cleaner than Cursor's, and the live pass confirms it.**
`auth.json` is a sibling of `sessions/`, not inside the session files — an adapter that
reads only `sessions/**` never opens a credential-bearing file. The directory listing on
the reference box matches: `~/.pi/agent/` holds `auth.json`, `settings.json`,
`models-store.json`, `AGENTS.md`, `extensions/` and `sessions/` side by side, so the
allowlist is a path prefix rather than Cursor's row-by-row filtering of one shared SQLite
file. A grep for `apiKey`/`accessToken`/`refreshToken`/`authorization` over the session
corpus returns zero matches. That lowers the read-allowlist burden; it does not remove
the planted-marker test, because session *content* is still untrusted and a zero today is
a measurement of this corpus, not a guarantee about the format.
**The extension seam is the part no other vendor has.** Pi's product thesis is an
in-process TypeScript extension system. That means a *Pi extension* could write
telltale's relay files (§7.15/§7.16 shapes) directly, with vendor-computed numbers, no
adapter parse at all — which would make Pi the first external writer of the relay
contract rather than the sixth in-tree adapter. Which of the two paths to build is an
owner decision deferred to post-launch demand; this survey only records that both are
open and neither is blocked by the format.
**Where this leaves the two Pi questions.** They are separate questions with separate
answers, and this section is the reason they can be answered separately. The **HUD**
question is settled: `internal/adapter/pi` is the in-tree observer, and the verdict table
above is what it rests on — `Session.Cost` stays CapNone because `usage.cost.total` is
per message, and it is carried as a labeled Extra instead. The **council seat** question
is unchanged and still refused under §9.1's re-host class; a measured format was never an
argument for a seat. The **relay-extension** path — a Pi extension writing §7.15/§7.16
files directly — remains open, is not what the adapter does, and stays gated on demand
rather than on anything this survey found.
### 3.10 The canary set — what each adapter actually watches
Every survey above pins an adapter to a private, unversioned on-disk format. §7 records how
drift *renders*; this is the other half, and without it the next person re-verifying a vendor
cannot know what was being watched. `grep -n canary docs/design.md` used to return nothing,
which was the whole of the gap.
A **canary** is a structural fact the survey established is present on *every* well-formed unit
of that vendor's corpus. A read that examined units and found no canary is reading a corpus that
has moved, and says so. `internal/adapter/drift` holds the mechanism; this is the inventory.
| adapter | verified against | canary | fields it feeds |
|---|---|---|---|
| Claude Code | `Claude Code 2.1.233` | `sessionId` — on every JSONL record that feeds a field | name, model, workspace |
| Codex CLI | `codex-cli 0.147.0` | `envelope type` — on every rollout record | model, workspace, quota, context % |
| | | `session_meta record` — the FIRST record of every rollout | workspace |
| Gemini CLI | `gemini-cli v0.53.1` | `metadata record` | name, subagents |
| Antigravity | `agy 1.1.13` | `gen_metadata table` | model |
| | | `trajectory_metadata_blob table` | workspace |
| Cursor | `Cursor 3.14.7` | `composerHeaders timestamp columns` | last activity |
| | | `meta.json updatedAtMs` — the CLI manifest's clock, on 43 of 43 manifests | last activity |
| Grok CLI | `grok 1.0.4 (d846eb93d9)` | `summary.json info.id` — the identity envelope, on 30 of 30 sessions | name, model, workspace, last activity |
| Pi | `pi 0.84.1` | `session header id` — first JSONL record is `type=session` with a non-empty `id` | name, model, workspace, last activity |
The middle column quotes each canary by the **name the adapter gives it**, not a paraphrase, so
the string in this table is the string in the code — which is what makes the guard tests below
able to check it at all.
**Cursor's row gained a second canary on 2026-08-29, and its `verified against` cell did not
move.** The cell names the Cursor APPLICATION the SQLite store was surveyed inside, and
`internal/adapter/pins` already records that this pin and an installed `cursor-agent` version
cannot be compared. The second canary watches a different store written by a different
program, so it carries its own pin — `cursor-agent 2026.08.11-e8db854`, in `chats.go`'s
`chatsVerifiedAgainst` — rather than borrowing this cell. The `pins` table stays one row per
vendor: `pins.For` answers per vendor id, and a second Cursor row there would make that answer
depend on ordering.
**Grok's row moved to 1.0.4 on 2026-08-14, and the row's two halves were re-checked to different
depths.** The version came from re-measuring the seat after four patch bumps went unnoticed
(§9.39's 2026-08-14 amendment). What was re-read on disk is one session directory that 1.0.4
itself wrote: `info.id` is there, so is every other key `internal/adapter/grok` names, and a
rate/limit/quota sweep over it still matches nothing account-level. What was NOT re-run is the
30-session census the canary's own phrase quotes — that number is still the 2026-08-09 survey's,
and it is left saying so rather than quietly re-attributed to a build it was never counted on.
**One table rather than a paragraph in each of §3.1–3.9**, deliberately, and against the first
instinct that a canary is a survey finding belonging beside its own survey. It is — but the
question this answers is asked *across* adapters ("what is being watched, and where is the gap"),
and five copies of the same claim in five subsections is five places for it to drift out of step
with `internal/adapter`. The survey sections keep the evidence; this keeps the inventory.
**What the columns are not.** `verified against` is CONTEXT, never a trigger — a version
comparison would fire on every vendor release that did not move a byte, and a report nobody
reads is worse than none. `fields it feeds` is what degrades when that canary goes missing, which
is why a canary is the load-bearing subset of a schema fingerprint rather than the fingerprint:
a vendor ADDING a column costs this program nothing, because every reader here addresses columns
by name.
**2026-08-16 — `telltale doctor` now carries the comparison this column would not make, and
that changes nothing above.** The sentence before this one still holds on the read path: no
adapter compares a version, nothing here triggers on one, and a canary is still structural. What
changed is that the `verified against` column acquired a second reader with a different audience.
§9.42's amendment has the full argument; in short, `agy` and `grok` self-update, so a pin in this
table goes stale silently and CI — which installs no vendors — can never notice. The preflight
already asks each seat its version, so it now says on the seat's own line when the installed
build is not the pinned one, and names the § to re-measure. It is a **staleness note about this
repository**, not a check of the operator's machine: no check fails, no tally moves, the exit
code is unchanged, and a seat whose version could not be read gets no verdict in either
direction.
**This table is now machine-read.** `internal/adapter/pins` mirrors it, and every pin in that
table is the adapter's own exported constant rather than a copy. `pins_doc_test.go` parses the
rows above and compares them **cell by cell**, in both directions. That is the fix §3.8 named
and left unowned: the six per-adapter guards match a pin anywhere in this file, so a dated
paragraph quoting a new pin can turn them green over a stale cell — which is precisely how the
Antigravity row survived a release reading `agy 1.1.9`. Editing a pin in this table without
moving the adapter constant now fails the build, and so does adding a row no adapter claims.
## 4. Adapter contract (v1)
One module per vendor implementing:
- `discover()` — find live/recent sessions from vendor-native data on disk
- `read(session)` — return the normalized session model (schema TBD, documented here)
- `capabilities()` — which normalized fields this vendor can actually source
The contract, the normalized schema, and a worked third-party example are documentation
deliverables of v1, not afterthoughts. The Go form is §4a.5 and the worked example is
§4a.7 — whose subject, Gemini CLI, became a real built-in adapter on 2026-08-02; the
example keeps its original sketch precisely because live verification overturned part
of it (see the §4a.7 postscript).
### JSONL framing rule (binding on every adapter)
Both vendors' on-disk sources are JSONL (Claude transcripts, `~/.codex/sessions`), and
both carry model-authored text. **A JSONL record is framed by the `\n` byte (0x0A) and
nothing else.** U+2028 (LINE SEPARATOR) and U+2029 (PARAGRAPH SEPARATOR) are legal
*unescaped* inside a JSON string value, so a reader that splits on "lines" in the
Unicode sense tears one record in two and both halves fail to parse — a HUD row that
silently loses sessions. Node's `readline` has exactly this bug; pi's
`packages/coding-agent/src/modes/rpc/jsonl.ts` hand-rolls a `\n`-only splitter to avoid
it, with the reasoning written down.
Go is structurally safer: `bufio.ScanLines`, `bufio.Reader.ReadBytes('\n')` and
`bytes`/`strings.Split` all match the 0x0A byte exactly, and the UTF-8 encodings of
U+2028 (`E2 80 A8`) and U+2029 (`E2 80 A9`) contain no 0x0A byte. So the rule for
adapters is: **split at the byte level, never adopt a dependency that does Unicode
line-breaking.** The property is pinned by tests in `internal/claude/stdin_test.go`
rather than assumed.
Two adjacent traps to avoid when the adapters land:
- `bufio.Scanner` caps a token at 64 KiB by default and then returns
`bufio.ErrTooLong`. Transcript records routinely exceed that, and an ignored scanner
error truncates the rest of the file — reading as "no more sessions". Use
`bufio.Reader.ReadBytes('\n')`, or `Scanner` with an enlarged buffer, and **check
`Err()`**.
- A trailing partial line (the vendor is still writing) is not a record. Hold it until
its `\n` arrives; a half-record must degrade to `—`, never to a parsed-looking value.
**Audit record (2026-08-01):** the repo was swept for this hazard at the point the pi
harness surfaced it. At that time the only parse site was `claude.Parse`
(`internal/claude/stdin.go`), a streaming decode of a single JSON value from stdin with
no line splitting anywhere — not exposed, nothing to fix. This section exists so the
question is answered before the HUD adapters are written, not re-audited after.
**Implementation (added with the adapters):** the rule now has one tested home,
`internal/jsonl` — `Split`, `Scan`, `Head` and `Tail`. `Tail` additionally discards the
first fragment after a backward seek, because a bounded tail read lands mid-record and
parsing that fragment invents a record the vendor never wrote. Both adapters use it; no
adapter reads bytes on its own.
## 4a. The normalized session model
*(Answers §6 Q2. Implementation: `internal/model/session.go`, package `model`, stdlib
only — the statusline path must never link a TUI framework.)*
Adapters produce `model.Session`; the statusline and the HUD only read it. Nothing
downstream of an adapter knows a vendor's field names, units, or file formats.
### 4a.1 Absence is two different things
The honest-gauge rule (ADR-001) says a displayed value must come from vendor data. That
forces a distinction most schemas skip, because the HUD renders the two cases
differently:
| | what it is | how it is encoded | how it renders |
|---|---|---|---|
| **absent now** | the adapter can source this field, but there is no value for this session right now (Claude's `rate_limits` on an API-key login) | nil pointer **+** capability declared | `—` in the cell |
| **can't know** | the vendor exposes no such thing, ever | nil pointer **+** capability not declared | the column is dropped for that vendor |
So presence lives in two places on purpose: **the value** (a nil pointer, per session)
and **the capability** (a static declaration, per adapter). Neither alone is enough — a
column of dashes for a vendor that could never fill it is itself a small lie about what
was measured.
Every optional field is a pointer. There is no "unset" sentinel number anywhere in the
package: `0` always means the vendor said zero, and the zero `time.Time` is invalid input
(`Validate` rejects it) rather than a stand-in for "no timestamp".
### 4a.2 Fields
Required — an adapter that cannot produce these has no row to render:
| Go field | Type | Notes |
|---|---|---|
| `Vendor` | `VendorID` | stable lowercase id, matches the adapter package name; appears in config keys and fixtures |
| `ID` | `string` | opaque, unique within the vendor, stable for the session's life — the HUD matches rows across polls with `Vendor/ID` |
| `ObservedAt` | `time.Time` | when **this snapshot was read**, not when the session last did anything. The HUD marks rows whose snapshot has aged out because polling failed; showing an old snapshot as current is precisely the failure the honest-gauge rule exists to prevent |
Optional — each has a stable **field id** used by `Capabilities`, fixtures, and this doc:
| Field id | Go field | Type | Meaning |
|---|---|---|---|
| `name` | `Name` | `*string` | human label. Model-authored text: may contain U+2028/U+2029, so renderers must not assume one line (§4) |
| `model` | `Model` | `*Model` | `{ID, DisplayName}`; `Name()` falls back to the id, same rule as the statusline |
| `workspace` | `WorkspaceDir` | `*string` | absolute native-format path. `WorkspaceName()` gives the basename for display |
| `context_pct` | `ContextPercent` | `*Percent` | 0–100, as vendors report percentages. Convert once at the edge; nothing downstream rescales |
| `cost` | `Cost` | `*USD` | USD only. A vendor reporting another currency declares `CapNone` rather than converting at an unsourced rate |
| `quota` | `Quota` | `[]QuotaWindow` | labeled usage windows, see below |
| `last_activity` | `LastActivity` | `*time.Time` | last observable activity. An **input** to liveness, not a claim about it |
| `liveness` | `LivenessHint` | `*Liveness` | the adapter's own verdict; see 4a.4 for when you are allowed to set it |
| `subagents` | `Subagents` | `*int` | count of the session's recently-written sub-agent transcripts. **Zero is a measurement** (we looked and found none) and must survive as one; nil means the count could not be taken |
Three more per-snapshot annotations, all optional:
- `Derived FieldSet` — fields whose value in *this* snapshot the adapter computed rather
than read (summing transcript token counts into a context percentage, say). Must be a
subset of the adapter's declared `Capabilities.Derived`, and every marked field must
actually carry a value. The HUD renders these with an estimate marker; ADR-001 requires
inferred values be visibly marked, not silently mixed in with reported ones.
- `Degraded FieldSet` — fields the adapter tried to read and failed (a truncated JSONL
record, an unparseable number). Degraded fields must be absent. **Degraded and
plain-absent render identically as `—`**; the difference is diagnostic only, shown in
the detail pane. If "we failed to read it" got its own gauge glyph it would start to
read as data.
- `Diagnostics []string` — operator-facing notes explaining degradation. Never rendered
as values, and (public repo) they describe structure — `"record 41 truncated"` — never
transcript content.
`Extras []Extra` is the escape hatch for vendor-specific labeled strings, so an adapter
with something extra to show does not stuff it into a field that means something else.
Extras are display-only: no thresholds, no colors, no sorting, detail pane only. If an
extra deserves a gauge, it deserves a `Field` — propose one. *(v1.1: the detail pane
(§7.11) is that surface, and it is now the only place extras appear —
`TestDetailPaneIsTheOnlyPlaceExtrasAppear` asserts both halves. Both adapters populate
them: git branch, CLI version, Claude's context token count, Codex's plan and history
mode.)*
**Why `subagents` is a `Field` and not an Extra.** It fails the Extra test in both
directions: it is a *number* with an absent-versus-zero distinction the Extra type cannot
carry (an Extra is a string, and `""` would collapse "none running" into "could not
count"), and it renders as a gauge-adjacent mark in the grid rather than as a labelled
line in the pane. That is exactly the "if an extra deserves a gauge, it deserves a Field"
case, taken rather than dodged.
### 4a.3 Quota windows
Windows are a slice, not named fields, because the set is vendor-defined. Emit only the
windows your vendor actually has, in display order, shortest first. Presence works at two
levels and both are load-bearing:
- a window the vendor does not have is **absent from the slice**;
- a window that exists but has no usage figure yet is **present with a nil
`UsedPercent`** and renders `—`. Never `0%`.
Each window carries `ID` (stable snake_case key, e.g. `five_hour`), `Label` (short
display string, ≤ 4 cells — the statusline is character-budgeted), `UsedPercent`, and
`ResetsAt`. A nil `ResetsAt` hides the countdown rather than guessing one. A window whose
*length* the vendor did not report gets a positional label (`1st`, `2nd`) from the
adapter: calling it "5h" on a guess would be a duration claim with no source.
### 4a.4 Liveness: who decides
**The HUD decides.** It classifies from `LastActivity` against one shared
`LivenessThresholds`, so every vendor's rows are judged by the same rule and are
comparable side by side. Defaults (`DefaultLivenessThresholds`):
| age since last activity | class |
|---|---|
| ≤ 2 min | `live` |
| ≤ 15 min | `idle` |
| > 15 min | `stale` |
| no timestamp, no hint | `unknown` → renders absent |
2 minutes because a working agent can go a long single turn without writing anything to
disk, and a boundary shorter than the longest quiet stretch of real work flaps mid-task.
15 minutes because past that a session is nearly always one the user walked away from, so
it sorts to the bottom instead of competing for attention. Both are defaults, not
constants: the HUD may expose them and the eval harness pins renders against explicit
values. A `LastActivity` in the future (file mtime vs. local clock skew) clamps to age
zero rather than going negative — and the adapters degrade it to absent before it gets
that far, so the clamp is a floor rather than a render path.
`unknown` is a real state and renders as absent — **never as `stale`**. "Stale" is a
claim; "we have no activity signal" is not.
**The adapter's only input is `LivenessHint`, and it wins when set.** Set it *only* from a
positive vendor signal the HUD cannot see:
- ✅ a turn-started / turn-ended event from a hook or notify stream;
- ✅ the vendor has recorded the session as ended → `LivenessStale`, even though
`LastActivity` is seconds old (this is the strongest legitimate hint);
- ❌ anything computed from the age of `LastActivity` — that is the HUD's job, and
duplicating it per-adapter makes vendors incomparable at different boundaries;
- ❌ "a process with that name is running" — that is evidence a process exists, not that
the session is doing anything.
A hint the model cannot check is the one place an adapter can lie undetected, which is
why the bar for emitting one is a signal that actually separates working-now from
process-exists. **Neither v1 adapter emits one** (§3.1, §3.2).
### 4a.5 The adapter interface
```go
type Adapter interface {
Vendor() VendorID
Capabilities() Capabilities
Discover(ctx context.Context) ([]SessionRef, error)
Read(ctx context.Context, ref SessionRef) (*Session, error)
}
```
- **`Capabilities()`** is static: which normalized fields this vendor can source, and how.
It must not vary with what a particular session happens to contain — that is what nil
pointers are for. Callers may cache it.
```go
type Capabilities struct {
Reported FieldSet // read from vendor output verbatim (modulo unit conversion)
Derived FieldSet // computed by the adapter from something that isn't the value
}
```
The two sets are disjoint; a field in neither is `CapNone` — "can't know". Declaring a
capability is a promise about the *source*, not about any given session.
- **`Discover()`** must stay cheap: directory listing and `stat`, no parsing. The HUD
calls it every poll tick. It returns `SessionRef{Vendor, ID, Locator, LastActivity}` —
`Locator` is vendor-private (a path, a pipe name), opaque to the HUD, never rendered
(on a shared machine it can name another user's paths), and handed back to `Read`
unchanged. `SessionRef.LastActivity` is a scheduling hint (typically an mtime) so the
poll loop can skip unchanged sessions; the value the HUD *displays* comes from `Read`.
- **`Read()`** parses one session. **Partial failure is not an error**: a field you cannot
parse is left nil, added to `Degraded`, and explained in `Diagnostics` — the row still
renders with `—` in that cell. Return an error only when there is no session to report
at all.
- Errors the HUD handles by showing *less*, not by showing a banner:
`ErrVendorAbsent` (vendor not installed — the vendor disappears from the HUD entirely;
a user without Codex should not stare at a Codex error forever) and `ErrSessionGone`
(the session vanished between `Discover` and `Read` — the row drops silently). Any
other `Read` error drops that one row and nothing else; the Codex adapter uses this for
`ErrSubAgentThread`, a rollout that is a sub-agent's thread rather than a session and
cannot be identified before parsing.
- Implementations must be safe for concurrent use (the HUD polls vendors in parallel),
must not write to vendor state, and must not make network calls or read credentials.
- **JSONL adapters:** §4's framing rule is binding. Split on the `0x0A` byte only, use
`bufio.Reader.ReadBytes('\n')` (or `Scanner` with an enlarged buffer) and **check
`Err()`**, and treat a trailing partial line as not-yet-a-record. `internal/jsonl` is
the shared implementation; use it rather than re-deriving it.
### 4a.6 The validation gate
`(*Session).Validate(caps)` is the machine-checkable form of the honest-gauge rule. It
rejects a session that:
- carries a value for a field the adapter declared unsupported;
- marks a field derived that was not declared derived, or marks a field derived that
carries no value;
- reports a field as both present and degraded;
- has a percentage outside 0–100 (drop it and mark it degraded — a clamped value is
invented data), a negative cost, or a zero `time.Time` used to mean "absent";
- has no `Vendor`, `ID`, or `ObservedAt`, or duplicate/unlabeled quota windows.
The eval harness runs it over every fixture; run it in your adapter's own tests too. It is
not on the render path.
### 4a.7 Worked example: adding a Gemini CLI adapter
*(2026-08-02: this example became real — `internal/adapter/gemini`, seam verification
in §3.7. The sketch below is kept AS WRITTEN, wrong guess included, because the gap
between it and the shipped adapter is the section's whole lesson; see the postscript.)*
**Step 0 — verify the seam before writing a line of code.** ADR-001's live-doc rule
applies to third-party adapters too: check the vendor's *current* docs for what it
actually writes to disk. This repo's own Codex plan was falsified on first check (no
statusline hook), and the sketch below is deliberately a *shape*, not a claim — the file
locations and field names are placeholders until you have verified them. A capability you
cannot point at a documented source for is `CapNone`.
**Step 1 — write the capability table first**, in your ADR or PR description. It is the
honest inventory of what your vendor can actually answer, and it is the thing reviewers
argue with. Illustrative shape:
| field id | capability | source |
|---|---|---|
| `name` | reported | session file header |
| `model` | reported | session file header |
| `workspace` | reported | session file header |
| `context_pct` | **derived** | token counts summed from records — not a vendor-reported percentage |
| `cost` | none | vendor exposes no cost |
| `quota` | none | vendor exposes no quota window |
| `last_activity` | reported | timestamp of the last record |
| `liveness` | none | no turn-start/turn-end signal to read |
**Step 2 — implement.**
```go
// Package gemini adapts Gemini CLI's on-disk session data to model.Session.
// Source paths and field names verified against on .
package gemini
const Vendor = model.VendorID("gemini")
type Adapter struct{ root string } // e.g. filepath.Join(home, ".gemini", "sessions")
func (a *Adapter) Vendor() model.VendorID { return Vendor }
// Capabilities: context_pct is DERIVED — we sum token counts ourselves, so the
// HUD marks it as an estimate. Cost, quota and liveness are absent from the
// vendor's data entirely and are declared nowhere: the HUD drops those columns
// for gemini rather than printing dashes it can never fill.
func (a *Adapter) Capabilities() model.Capabilities {
return model.Capabilities{
Reported: model.NewFieldSet(
model.FieldName, model.FieldModel, model.FieldWorkspace,
model.FieldLastActivity,
),
Derived: model.NewFieldSet(model.FieldContextPercent),
}
}
// Discover stats the session directory only — no parsing on the poll path.
func (a *Adapter) Discover(ctx context.Context) ([]model.SessionRef, error) {
entries, err := os.ReadDir(a.root)
if errors.Is(err, fs.ErrNotExist) {
return nil, model.ErrVendorAbsent // not installed: hide the vendor
}
if err != nil {
return nil, err
}
var refs []model.SessionRef
for _, e := range entries {
info, err := e.Info()
if err != nil {
continue // racing the vendor's writer is normal, not fatal
}
refs = append(refs, model.SessionRef{
Vendor: Vendor,
ID: strings.TrimSuffix(e.Name(), ".jsonl"),
Locator: filepath.Join(a.root, e.Name()),
LastActivity: model.TimePtr(info.ModTime()),
})
}
return refs, nil
}
func (a *Adapter) Read(ctx context.Context, ref model.SessionRef) (*model.Session, error) {
f, err := os.Open(ref.Locator)
if errors.Is(err, fs.ErrNotExist) {
return nil, model.ErrSessionGone
}
if err != nil {
return nil, err
}
defer f.Close()
s := &model.Session{Vendor: Vendor, ID: ref.ID, ObservedAt: time.Now()}
// §4 framing rule, via the shared implementation: records are framed by
// 0x0A and nothing else, there is no 64 KiB cap, and a trailing partial
// line is not a record.
err = jsonl.Scan(f, func(line []byte) error {
var rec record
if json.Unmarshal(line, &rec) != nil {
// One bad record degrades the fields it fed, not the whole row.
s.Degraded = s.Degraded.With(model.FieldContextPercent)
s.Diagnostics = append(s.Diagnostics, "unparseable record skipped")
return nil
}
apply(s, rec)
return nil
})
if err != nil {
return nil, err
}
// Derived, and declared as such: this is a computed estimate, and the HUD
// renders it with an estimate marker rather than as a vendor-reported figure.
if pct, ok := estimateContext(s); ok && !s.Degraded.Has(model.FieldContextPercent) {
s.ContextPercent = model.PercentPtr(pct)
s.Derived = s.Derived.With(model.FieldContextPercent)
}
// Cost and quota are never set: not declared, so not knowable.
// LivenessHint is never set: no positive signal, so the HUD classifies by age.
return s, nil
}
```
**Step 3 — fixtures and the gate.** Add synthesized fixtures per state — healthy, empty,
degraded (a truncated final record), and whatever your vendor's equivalent of "logged in
a way that hides quota" is — and assert `Validate` passes plus the exact render for each.
Fixtures are **synthesized**: fake session ids, fake text, fake paths, realistic in shape
only. This repo is public and real transcripts carry private material; no real session
content enters `testdata/`, ever.
**Step 4 — write down what you could not source**, in your adapter's package doc. A field
declared `CapNone` with a one-line reason is a finished answer. A field quietly filled
with a plausible number is the bug this whole schema exists to make hard.
**Postscript — what Step 0 did to this very sketch.** The sketch above, written before
anyone read the vendor's source, guessed `context_pct: derived` ("token counts summed
from records"). The source read (§3.7) falsified it: Gemini's per-message token counts
are real, but no context-window size reaches disk — the CLI's own percentage divides by
a static table compiled into its binary, which is exactly the assumed denominator
decisions/001 forbids. The shipped adapter declares `context_pct: CapNone` and carries
the token reading as a display-only extra instead. The hypothetical also missed the two
things only the source could reveal: message records are upserts (same id re-appended),
and the writer deletes non-resumable sessions on exit. If you skip Step 0, those two
become bugs; the wrong capability guess becomes a fabricated gauge.
## 5. Eval harness
Fixture-driven, in-repo, CI-gating (`.github/workflows/ci.yml` runs `go vet ./...` and
`go test ./...` on `windows-latest`, then smoke-tests the built binary against a
statusline fixture). What it asserts today:
| Layer | Package | What is pinned |
|---|---|---|
| Statusline renders | `internal/statusline` | every segment against five stdin fixtures, including the API-key login that must render no quota |
| Statusline (agy) | `internal/statusline` + `internal/antigravity` | four agy fixtures + inline cases: full render exact, confirm? outranking the state word, ctx 0% as a reading beside hidden quota, bucket-without-reading hides, unknown state verbatim, reset_time fallback, U+2028 in a string value, the product routing marker |
| Framing rule | `internal/jsonl` | 0x0A-only framing, a 300 KiB record surviving the `bufio.Scanner` cap, read errors surfacing, torn tails held back, a seek fragment discarded |
| Schema gate | `internal/model` | `Validate` over every rejection case, liveness boundaries, presence semantics |
| Claude adapter | `internal/adapter/claudecode` | discovery filters, the `` trap, the `input_tokens` trap, torn tail invisibility, torn-only session, future mtime, capability table |
| Codex adapter | `internal/adapter/codex` | envelope + internally-tagged event parsing, derived context, quota window presence, null `rate_limits`, sub-agent rejection, capability table |
| SQLite reader | `internal/sqlite` | record decoding across every storage class, the zero-width serial types, the overflow-page chain (25 KiB blob against a 4 KiB page), missing-table-is-absence, and the WAL overlay: a committed sidecar value winning, a corrupt frame ignored with a note, a bad header rejected whole, mismatched salts ignored, a torn tail never assembled |
| SQLite reader (v1.2) | `internal/sqlite` | `Rows` streaming the same rows as `Table` and stopping when the callback says so; `Columns` splitting a CREATE statement across every quoting style, a parenthesized type, a comma inside a default and a table-level constraint, and yielding nothing rather than a guess on a statement it cannot read |
| Antigravity adapter | `internal/adapter/antigravity` | discovery that ignores the stale summary index and the sidecars, WAL overlay changing the reported model, overflow-spanning generation blob, dedup on the response id rather than the constant conversation UUID, invariant violation dropping the tokens with a diagnostic, zero tokens as data, absent workspace as absence, missing transcript as a typed sentinel, unreadable database degrading rather than dropping the row, Q8 fold, future mtime, transcript content never reaching a field, capability table |
| Cursor adapter | `internal/adapter/cursor` | discovery dropping the five row shapes that are not sessions (draft sentinel, `value.isDraft`, archived, sub-agent, empty-window), a store whose main file is one empty page reading entirely out of its sidecar, the vendor's own percentage reported unmarked, a computed one marked, neither-present as absence, `default` rendered literally, unpopulated zeros never becoming a cost or a token reading, missing workspace mapping as absence and an unparseable one as degradation, mixed epoch-ms/ISO-8601 timestamps, future-skew, all-timestamps-unreadable degrading, the stale `ItemTable` mirror losing to the table, an unrecognized schema erroring rather than reporting zero, capability table, and the allowlist: three planted markers (prompt text, credentials, plan entitlements) reaching no field, extra or diagnostic |
| Gemini adapter | `internal/adapter/gemini` | fixed-depth discovery (legacy `.json` and nested sub-agent files excluded), upsert last-wins, `$set` summary/lastUpdated, registry workspace lookup (verbatim, corrupt-degrades, absent-is-absent), sub-agent nest counting, `kind:"subagent"` rejection, torn tail invisibility, future mtime, Q8 fold, capability table |
| Claude adapter (v1.1) | `internal/adapter/claudecode` | the sub-agent count: recency boundary, future-mtime exclusion, non-transcript neighbours ignored, absent sidecar as a measured zero |
| HUD renders | `internal/hud` | every golden frame byte-for-byte at 52/72/80/120 columns (count enforced by TestEveryGoldenIsClassified, not restated here — a literal drifted once already), the §7.4 gauge table, the estimate marker, threshold colours, frame width/height invariants |
| HUD behaviour | `internal/hud` | vendor status words from adapter errors, key handling, one-scan-in-flight, spinner lifecycle |
| HUD behaviour (v1.1) | `internal/hud` | esc unwinding one layer at a time, find mode swallowing the keyboard, selection carried by session key across a re-sort, the pane closing when its session ends |
| Burn arithmetic | `internal/hud` | the minimum basis, the four refusals, least-squares slope against injected series, rollover detection vs. `resets_at` jitter, sample throttling and eviction |
| Fixture legality | `internal/hud` | every session behind every golden passes `model.Validate` against its vendor's declared capabilities — a golden may not pin a render of a state the schema forbids |
| Doc/code sync | `internal/hud` | every render pasted into `docs/design.md` §7.3/§7.11–§7.14 still matches its golden, and every golden is either embedded or explicitly exempted |
| Picture/code sync | `internal/hud` | `README.md`'s hero picture is re-emitted from the `readme` golden and byte-equal to the committed file, with the characters read back out of the emitted markup and diffed against the render — plus no dollar sign anywhere, and the estimate marker surviving into the picture |
| Picture/code sync | `internal/council` | the hero picture `README.md` and `docs/council.md` both show is re-emitted from the `activity` golden with its all-blank rows dropped, byte-equal and read back the same way, with exactly one seat wearing the focus mark |
| Fast path (ADR-002) | `internal/statusline` | the statusline's transitive import graph reaches no `charm.land/bubbletea` or `charm.land/lipgloss` package — the framework cannot be initialized on a path it does not import |
Rules that outrank convenience:
- Fixtures are **synthesized**. No real session content enters `testdata/`, ever.
- Golden renders use a pinned clock, an explicit terminal size and a plain style set, so
they never depend on the CI terminal.
- A failing render assertion fails the build.
- No number appears in README/launch material unless this harness generated it.
**Amendment, 2026-08-16: the fast path is gated structurally and reported numerically.**
ADR-002's fast-path rule — the statusline never initializes Bubble Tea — was the oldest
load-bearing claim in this document with nothing checking it. Any `import` added to
`internal/theme` or `internal/model` for one convenient helper would have compiled,
passed every golden and every smoke, and put the TUI framework's package init on a path
that runs on every prompt. Two things now exist, and they are deliberately unequal.
**The hard gate is structural.** `TestFastPathNeverReachesTUIFramework`
(`internal/statusline/fastpath_test.go`) asks the toolchain — `go list -deps` — for the
statusline's transitive imports and fails if any is under `charm.land/bubbletea` or
`charm.land/lipgloss`. It runs inside `go test ./...`, so it gates on both CI jobs. It
was verified against an injected violation before it was trusted: a bare
`import _ "charm.land/lipgloss/v2"` in `internal/theme` turns it red, naming the
package. A gate nobody has seen fail is a gate nobody has tested.
**The timing is reported, and holds only a gross-regression ceiling.** A `ci.yml` step
spawns the built binary 15 times against `full.json`, asserts the line rendered on every
sample, prints min/median/max, and fails only if the median exceeds **400 ms**. That
ceiling is not the budget and does not pretend to be. A shared runner's scheduling noise
is larger than the quantity ADR-002 cares about, so pinning single-digit — or even
double-digit — milliseconds there would buy a flaky build, not a guarantee.
**What the gate deliberately does NOT assert**, stated so a later reader does not credit
it with more than it does:
- **Not the single-digit-millisecond budget.** Nothing in CI measures that. The
measurement lives on the reference workstation (i7-7700K, Windows 11, 2026-08-16):
parse+render 14 µs, end-to-end median 26.6 and 29.0 ms over two runs of the CI
harness. The end-to-end figure is process start, not telltale's work.
- **Not "the binary does not link the framework".** It does — see §7.5's 2026-08-16
correction. One binary carries the HUD, so the module is linked and the structural
claim can only ever be about the statusline's own import graph.
- **Not a small constant regression.** The 400 ms ceiling catches a change in the SHAPE
of the path — a whole-corpus scan (§7.18 measured cold scans at 896–1200 ms), a
network round trip, a TUI init and its terminal-capability queries. A path that got
three times slower and stayed under the ceiling passes, and the printed median is the
only thing that would show it. That is a human's job, on purpose.
- **Not the other two gauges.** `internal/hud` and `internal/council` are TUI surfaces;
the rule does not apply to them and the gate does not look at them.
**Amendment, 2026-08-18: what the long-running modes cost after the first minute, measured.**
Everything above measures one call, and the 2026-08-16 timing step measures 15 of them.
Three modes do not stop after one call. `telltale events` and `telltale otel grok` are
listeners that an operator starts once and leaves. `telltale council` holds a room and one
long-lived child for each seat. A benchmark cannot answer the question these modes raise:
does an idle process hold a constant amount of memory across a day? Nothing in this
repository measured that. This amendment records the instrument and the first measurement.
**The instrument is `tools/soak.ps1`, and it is not a gate.** It has three modes. Residency
mode samples one process tree at an interval, and it records the working set, the private
bytes, the handle count, the live CPU and the child count. Per-fire mode starts a
short-lived process many times and records each run, because `telltale statusline` has no
residency to sample. Summary mode renders a finished JSONL file again, so a later reader
re-reads an arm without a second soak. The samples are the artifact, and a table is one
view of them. Re-read a finished arm with
`.\tools\soak.ps1 -Summarize "$env:TEMP\soak-events.jsonl"`.
Three decisions in that script are load-bearing, and each one comes from a measured
failure:
- **`Win32_Process`, and never `Get-Counter`.** The counter set names of `Get-Counter` are
localized, so `\Process(*)\Working Set` does not resolve on a Windows installation in
another language. Its instances also carry process names, so three `telltale.exe`
processes arrive as `telltale`, `telltale#1` and `telltale#2`, with no stable relation to
a pid and no parent data. A tree soak needs the parent data. One CIM query also serves a
whole sample at any tree size.
- **The summary prints a median-based drift beside the least-squares slope.** One outlier
moves a slope, and no outlier moves a median. Arm C below proves why that matters: its
slope and its drift disagree, and the drift is the honest figure.
- **An absent figure prints `--`, and never `0`.** `PeakWorkingSet64` reads 0 after a
process exits, and PowerShell evaluates `$null / 1MB` as 0. Both defects produced a
measured-looking zero in the first smoke arms. That is the zero-vs-absent collapse
ADR-001 forbids, inside the instrument that exists to find drift.
**Conditions.** The reference workstation: Intel i7-7700K (4 cores, 8 logical threads),
Windows 11, Windows PowerShell 5.1.26100.9168, Go 1.26.5. The binary came from `d024f03`.
Both listeners bound non-default loopback ports (14519 and 14318), and every arm ran with
`USERPROFILE` redirected to a scratch directory. The redirect is verified in both
directions: the listeners wrote their stores under the scratch home, and
`~/.telltale` held no file newer than the previous day when the arms ended. No arm read or
wrote a real store. Arms A and B ran concurrently, and arm C ran on the same machine while
they sampled.
**Arm A — `telltale events`, idle, 34.8 min, 210 samples at 10 s.**
| metric | min | median | p95 | max | slope/hour |
|---|---|---|---|---|---|
| working set | 9.14 MiB | 9.23 MiB | 9.23 MiB | 9.25 MiB | +0.11 MiB |
| private bytes | 45.32 MiB | 45.32 MiB | 45.35 MiB | 45.36 MiB | -0.02 MiB |
| handles | 128 | 128 | 128 | 128 | 0.0 |
| processes | 1 | 1 | 1 | 1 | 0.0 |
Robust drift: working set 0.00 MiB, private bytes 0.00 MiB, handles 0.0, processes 0.0.
CPU over 209 comparable intervals, 0 dropped: min, median and max all 0.000%, total 0.02 s.
**Arm B — `telltale otel grok`, idle, 34.8 min, 210 samples at 10 s.**
| metric | min | median | p95 | max | slope/hour |
|---|---|---|---|---|---|
| working set | 9.02 MiB | 9.11 MiB | 9.11 MiB | 9.13 MiB | +0.11 MiB |
| private bytes | 45.15 MiB | 45.15 MiB | 45.18 MiB | 45.20 MiB | -0.02 MiB |
| handles | 128 | 128 | 128 | 128 | 0.0 |
| processes | 1 | 1 | 1 | 1 | 0.0 |
Robust drift: working set 0.00 MiB, private bytes 0.00 MiB, handles 0.0, processes 0.0.
CPU over 209 comparable intervals, 0 dropped: min, median and max all 0.000%, total 0.00 s.
**Arm C — `telltale statusline`, 1000 fires.** Every fire exited 0, and every fire rendered
the `Opus` marker that `ci.yml` asserts on its own 15 samples.
| metric | min | median | p95 | max | slope/1k fires |
|---|---|---|---|---|---|
| wall ms | 20.0 | 20.9 | 26.0 | 2035.7 | -41.2 |
| cpu ms | 0.0 | 15.6 | 31.3 | 93.8 | 0.0 |
| peak working set | 9.32 MiB | 9.43 MiB | 9.61 MiB | 9.63 MiB | -0.01 MiB |
| peak private | 44.90 MiB | 45.24 MiB | 45.44 MiB | 46.25 MiB | -0.01 MiB |
Robust drift: wall -0.2 ms, cpu 0.0 ms, peak working set 0.00 MiB, peak private 0.00 MiB.
p99 is 55.23 ms. 11 fires passed 50 ms, and 6 of those passed 100 ms. The cpu figure is
quantized to the 15.625 ms scheduler tick, so only its distribution carries information.
**What the three arms found.**
1. **Neither listener leaks.** The handle count and the process count held exactly constant
across 210 samples of each arm. Both working-set slopes read +0.11 MiB/hour, and that
figure is smaller than the 0.11 MiB total spread of the same arm. The robust drift is
0.00 MiB on every metric. The slope is therefore the noise of the arm, and it is not a
trend. Each arm ran 34.8 min, so every hourly slope above is a 1.7x extrapolation.
2. **Idle CPU is under the measurement floor.** Every comparable interval read 0.000%, and
no interval was dropped. The two listeners spent 0.02 s and 0.00 s of CPU across 35
minutes. An idle listener costs nothing this instrument can measure.
3. **The statusline cost does not drift, and its tail is the finding.** 6 fires of 1000
blocked for 2019-2036 ms, while their CPU stayed at the usual 15-31 ms. The process
waited; it did not compute. The -41.2 ms per 1000 fires slope comes from those 6 samples
alone, and the median-based drift reads -0.2 ms across the same run. This is the
disagreement the drift figure exists to expose, and the drift is the honest one. **This
arm does not name a cause.** Two residency arms polled `Win32_Process` on the same
machine throughout, so the condition is recorded and the cause is not claimed. A
statusline fires on every prompt, so a 2 s stall earns a later arm on an idle machine.
**What this measurement does NOT assert**, stated so a later reader does not credit it with
more than it did:
- **Not a CI gate.** `tools/soak.ps1` is an operator instrument. No workflow runs it, and a
leak introduced tomorrow turns nothing red.
- **Not a multi-hour figure.** 35 minutes is the arm. A leak under this arm's noise floor
stays invisible, and every hourly slope is an extrapolation.
- **Not a loaded listener.** Both listeners sat idle, and no hook posted to either port. A
listener under traffic is a different measurement.
- **Not the council room.** That arm was owed when this list was written; the
dated payment below records its run (2026-08-18).
- **Not macOS.** The instrument uses PowerShell and `Win32_Process`, so it is Windows-only.
**Owed: the council arm, which an operator must run.** The room cannot be soaked
headlessly, for two reasons that belong to the mode rather than to the script. The room is
a TUI, so it needs a real terminal. Its seats are live vendor CLIs that read their
credentials from the real home directory, so this arm cannot use the redirected
`USERPROFILE` that isolated the other three arms. The room is also the arm that matters
most, because it is the only mode whose process count can move: each persistent seat is a
long-lived child, and residency mode counts the whole tree.
Terminal 1 opens the room and leaves it idle. `--read` seats every vendor and forbids every
write, which is what an unattended 35-minute arm needs:
```powershell
cd C:\Users\sanle\code\telltale
.\telltale.exe council --read
```
Terminal 2 resolves the room's pid, then samples the tree. `-Name telltale` is wrong here,
because it fails whenever more than one `telltale.exe` runs:
```powershell
cd C:\Users\sanle\code\telltale
$room = @(Get-CimInstance Win32_Process -Filter "Name='telltale.exe'" |
Where-Object { $_.CommandLine -match '\bcouncil\b' })
$room.Count # must print 1 before the next command runs
powershell.exe -NoProfile -ExecutionPolicy Bypass -File .\tools\soak.ps1 `
-Id $room[0].ProcessId -Label council-idle `
-Out "$env:TEMP\soak-council.jsonl" -IntervalSeconds 10 -DurationMinutes 35
```
The arm prints its own table when it ends. Read the same file again at any later time with:
```powershell
powershell.exe -NoProfile -ExecutionPolicy Bypass -File .\tools\soak.ps1 -Summarize "$env:TEMP\soak-council.jsonl"
```
**Paid, 2026-08-18. The owner ran the arm exactly as written above.** A `--read`
room with all five seats reattached sat idle for 34.8 minutes, 210 samples at
10 s, on the same box as the three headless arms (Intel i7-7700K, Windows 11).
One honest difference from those arms is stated up front: the room ran the
operator's own clone build (the binary a CI-gate run leaves in the repo root),
so the revision is not pinned, and the operator used the machine for other work
during the window — the sampler is tree-scoped, so only the room's own process
tree was counted, but the CPU percentages share a loaded box.
| metric | min | median | p95 | max |
|---|---|---|---|---|
| working set | 21.55 MiB | 587.95 MiB | 727.81 MiB | 1790.64 MiB |
| private bytes | 56.58 MiB | 648.42 MiB | 795.28 MiB | 2217.25 MiB |
| handles | 155 | 949 | 1500 | 6978 |
| processes | 1 | 5 | 9 | 34 |
Robust drift (median of the second half minus the first): working set
+2.44 MiB, private bytes −1.38 MiB, handles −3, processes 0. CPU over 205
comparable intervals: median 0.429%, max 1.779%, 91.16 s total.
Three readings, in the order they matter:
1. **The room is not the residency; the seats are.** The tree's first sample is
21.55 MiB and one process — the room alone, before reattach. Steady state is
~590 MiB, five processes and ~950 handles: one TUI plus four persistent
vendor children. An all-day room costs what its vendors cost, and the gauge
binary itself is the smallest process in its own tree.
2. **No leak at steady state.** The robust drift is +2.44 MiB of working set
and exactly zero processes across the halves. The large negative
least-squares slopes (−173 MiB/hour) are the reattach transient, not a
decline: the peaks — 34 processes, 1.79 GiB, 6978 handles — all belong to
the opening minutes when five conversations reattached at once, and the tree
settled from there. The same outlier-versus-median disagreement the
statusline arm showed, with the sign flipped.
3. **"Idle" seats are not quiescent.** 91 CPU-seconds over an idle
35 minutes, and bursts to nine processes mid-window: the vendor CLIs run
their own background machinery inside an idle room. That is vendor
behavior observed from outside, not attributed to any one vendor — a
per-seat attribution would need one arm per seat, which nobody has run.
What this arm still is not: one run, one build, one box, 35 minutes — and a
loaded box for the CPU column. The reattach transient it measured is also the
demo's opening beat, which is worth knowing before a stage.
## 6. Open design questions
1. ~~Language/stack~~ — **ANSWERED, ADR-002:** Go + Bubble Tea/Lipgloss, one binary,
two modes. Windows-first hardened; macOS/Linux deferred post-v1.
2. ~~Normalized session schema~~ — **ANSWERED, §4a:** `model.Session` + `Capabilities`,
pointers for absent-now vs. undeclared capability for can't-know, liveness classified
by the HUD from `LastActivity` with an adapter override only on positive vendor
evidence, and `Validate` as the machine-checked honest-gauge gate.
3. ~~HUD refresh model~~ — **ANSWERED, and the measurement it was waiting for has been
taken.** A 1 s `tea.Tick` poll, not a file watcher. The survey's inputs argued for it:
837 sessions across 33 directories, transcripts up to 7.7 MB, and a projects tree that
provably mutates mid-sweep. `Discover` is stat-only and `Read` is head+tail bounded,
which is what made polling affordable; a watcher over a mutating tree on Windows is a
larger correctness surface for a smaller win.
The open part was the cold cache, and `BenchmarkScan` closed it over a synthesized
1,400-session corpus: the warm scan went 798 ms → **82 ms** with the `(size, mtime)`
cache, and on the live corpus 1.84–3.37 s → **181–204 ms**. **The cold scan did not
move** (896 ms → 994 ms, unchanged within noise), and that is the answer rather than a
remaining gap: the first frame must genuinely read everything, which is what the
spinner exists for. The poll was never the cost; re-reading unchanged files was.
4. ~~Exact Claude/Codex on-disk data sources~~ — **ANSWERED, §3.1–3.3**, with Claude
verified live and Codex's first live pass run 2026-08-01 (§3.4; short remainder
itemized there).
5. ~~Distribution naming (`telltale-hud` on any registry; winget/scoop manifests) — at
packaging time. Go binary means npm is optional, not required.~~ — **ANSWERED
2026-08-08, §8 "Packaging decisions":** the bare name `telltale` was free on scoop
and winget, so the `telltale-hud` fallback goes unused; winget takes the
publisher-qualified `sanlee-ys.telltale`; npm stays skipped rather than renamed.
6. ~~HUD UI design section~~ — **ANSWERED, §7:** layout grid, colour/threshold tokens
shared with the statusline, motion rules, degraded-state renders. Written before HUD
build per ADR-002; every render in §7.3 and every row in §7.7 is a golden/fixture.
7. **Cross-vendor context percentage — ANSWERED PROVISIONALLY, wants a ruling.** Claude
and Codex context percentages are different statistics (§3.3), and Claude's is not
derivable from disk at all. Three options were on the table:
1. show raw token counts only, cross-vendor, no percentage anywhere;
2. show a percentage only where a vendor ships a denominator, accepting a ragged
column;
3. show each vendor's own formula, labelled as vendor-native and non-comparable.
**v1 ships (2).** Codex's `context_pct` is `CapDerived` and renders with an estimate
marker; Claude's is `CapNone` and renders absent; and when no visible row can fill the
column it is dropped entirely. (3) is rejected outright: two numbers in one column
that are not the same statistic is exactly the lie the honest gauge forbids. (1) is
the more conservative answer and remains the better one if the ragged column reads
badly in daily use — it needs a `context_tokens` field in the schema, which is an
additive change, and both adapters already carry the token count as an extra so
nothing has to be re-derived. **Decide after two weeks of dogfood, not before.**
8. ~~LastActivity source on Windows~~ — **RULED 2026-08-01: option (iii),
`max(mtime, newest record timestamp)`, implemented in both adapters.** NTFS defers
mtime while the writer holds the file (~100 s observed on a hot rollout; ~20 min on
a closing one, seen twice), so mtime alone under-reports on exactly the rows the HUD
exists to watch. The newest record `timestamp` (RFC3339, vendor-written on records
in both formats — verified live; some Claude housekeeping records omit it) comes out
of the tail window the adapters already read, so the fold is zero extra I/O. Rules:
each signal independently passes the future-skew guard or is excluded (a wrong
vendor clock cannot fresh-wash a row); the fresher valid signal wins; the field
degrades only when BOTH are unreadable. Still `CapReported` — the max of two
vendor-written stamps invents nothing. Pinned by
`TestLastActivityUsesNewestRecordTimestampOverStaleMtime` in each adapter; no HUD
golden changed (fixtures inject `LastActivity` directly).
## 7. HUD UI design
Written before the HUD was built, per ADR-002. This section is binding: every render
below is a golden-test target in `internal/hud/testdata/golden/`, and the degraded states
in §7.7 are eval fixtures like the statusline's. Library facts were verified against live
docs on 2026-08-01 and then against the compiled API: `charm.land/bubbletea/v2` **v2.0.8**
and `charm.land/lipgloss/v2` **v2.0.5**. Both are v2: `AdaptiveColor` and the global
`Renderer` are gone, `Style` is a plain value type, and `Model.View()` returns a
`tea.View` struct rather than a string.
> **The renders in §7.3 are generated by the build, not drawn by hand.** They are pasted
> from `internal/hud/testdata/golden/*.txt`, which `go test ./internal/hud -update`
> regenerates. If a render here and the code disagree, the code is right and this section
> is stale — fix it in the same change.
### 7.1 Principles
Five rules, in priority order. Where polish and a rule conflict, the rule wins.
1. **The honest gauge extends to pixels.** Absent data renders as absent. Specifically:
an *empty gauge track* means zero, so an absent gauge draws **no track at all** — the
field is blank and the number beside it is `—`. `0%` and "no data" must never produce
the same row of glyphs. This is the load-bearing render assertion of the whole HUD.
2. **Colour is always redundant.** Every distinction is carried by a glyph or a number
first; colour only reinforces it. This makes `NO_COLOR` degradation correct by
construction rather than by a second code path, and makes the HUD readable to
colour-blind users without a mode.
3. **telltale may animate its own work; it must never animate the vendor's.** See §7.6.
4. **Still by default.** This is a glanceable monitor. In steady state the only cell
permitted to change each second is the `AGE` of a session younger than one minute.
If more than that moves, it is a bug, not a flourish.
5. **Same product as the statusline.** Identical thresholds, identical palette, identical
`│` separator, identical `↻` countdown. The two surfaces share numbers through
`internal/theme` (§7.5), not by coincidence.
A sixth rule that is a consequence of #1 and worth stating on its own: **account-level
quota appears once per vendor, in the header, never per row.** `rate_limits` is a
property of the account, not the session; repeating it on every row would assert
per-session quota, which is false. If no source can honestly supply it, that vendor's
block is absent — not zeroed — so the block count is itself a measurement: the header
shows exactly as many vendors as telltale can speak for. Where the readings come from
and how the line fits them is §7.15.
### 7.2 Anatomy
```
header identity, session counts, account quota 1 line (2 below 100 cols)
rule ───────────────────────────────────────── 1 line
col header SESSION MODEL CONTEXT COST AGE 1 line
rows one per session, sorted, scrollable n lines
rule ───────────────────────────────────────── 1 line
footer key hints (left) · state notices (right) 1 line
```
Chrome is 5 lines at wide, 6 below 100 cols (the quota block wraps to its own line). The
column-header row is dropped when the body is not the grid — over the help overlay or the
empty state it would label columns that are not on screen.
**Column grid.** Widths are fixed; the `SESSION` column is the only flexible one and
absorbs all slack, which right-anchors the numeric block at every terminal width.
Offsets below are 1-based for the wide tier at 120 columns.
| Cols | Field | Width | Align | Notes |
|---|---|---|---|---|
| 1 | pad / **selection** | 1 | | blank, or `▸` on the selected row (§7.11) |
| 2 | state dot | 1 | | `●` live / `◐` idle / `○` stale / blank unknown |
| 4–5 | vendor | 2 | left | `CC` / `CX` |
| 7 | separator | 1 | | dim `│` |
| 9–67 | **session** | **W−61** | left | flexes; `…` truncation |
| 70–82 | model | 13 | left | normalized display name |
| 85–96 | context gauge | 12 | | see §7.4 |
| 98–103 | context % | 6 | right | 6, not 5: a derived value carries a `~` marker |
| 106–112 | cost | 7 | right | |
| 114 | separator | 1 | | dim `│` |
| 116–119 | age | 4 | right | |
| 120 | pad | 1 | | |
Only two `│` separators per row, deliberately. They cut the row into three zones —
**identity** (dot, vendor), **measurement** (name, model, gauges, cost), **time** (age).
A pipe between every column reads as a spreadsheet; two pipes read as structure.
**Session label content.** The session's own `name` if the vendor has one, else the
workspace basename, else the vendor session id. Then the sub-agent chip if the session is
fanning out (§7.13). Then, only if ≥14 cells remain free, two spaces and the parent
directory (left-elided with `…`). The parent path disambiguates same-named projects under
different roots and stops the wide tier from opening a dead gulf between the name and the
model. It drops out automatically as the terminal narrows.
The chip's width is reserved **before** the name is truncated, and the name loses the
character. A chip that vanished on a long project name would make the same session look
like a different kind of session at a different terminal width — a lie by omission — and
the name is the field that can afford to lose a character, because the parent path and
the detail pane both still carry the identity.
> **Deviation from the original spec, deliberate:** the `⌥worktree` mark is not rendered.
> `worktree.name` exists only on the statusline's stdin payload, which the HUD does not
> consume, so **no adapter can source it** — a cell no adapter can fill is a cell that
> should not be in the grid. The glyph and its ASCII form stay in `internal/hud/glyphs.go`
> for the day a vendor writes it to disk.
**Responsive tiers.** Breakpoints are on width only; the shedding order is fixed, so the
layout at any width is a pure function of the width.
| Tier | Width | Model | Gauge | Cost | `SESSION` width | At 120 / 80 / 72 |
|---|---|---|---|---|---|---|
| wide | ≥ 100 | 13 | 12 cells | shown | W − 61 | 59 |
| compact | 80–99 | 13 | 8 cells | **dropped** | W − 48 | 32 |
| narrow | 60–79 | 13 | **dropped** | dropped | W − 38 | 34 |
| floor | < 60 | — | — | — | — | one-line notice |
Cost sheds before the gauge, and the gauge before the model, because the gauge is a
redundant encoding of a number that stays on screen, and the model is identity — the
answer to "which of my agents is this?" — which nothing else supplies. `MODEL` is never
narrowed below 13: `gpt-5.1-codex` is exactly 13 columns, and truncating a model name to
`gpt-5.1-c…` destroys the one field a user scans for.
Height tiers: **H ≥ 9** full chrome; **6 ≤ H < 9** drops both rules and the column-header
row (header + rows + footer only); **H < 6** shows the floor notice. Row overflow is not
paginated — the footer gains `+3 more`, and the viewport **follows the selection**: `↑`/`↓`
move the cursor and the visible window slides to keep it on screen. That arithmetic lives
in `Render` rather than in `Update`, because `Update` does not know how tall the chrome is
this frame and a second copy of the height maths is a second thing to get wrong.
Floor renders, exactly:
```
telltale needs 60 columns (have 52)
```
```
telltale needs 6 rows (have 4)
```
**Column auto-hide.** A column that would render `—` for *every* visible row is dropped
entirely and its width returned to `SESSION`; a full column of dashes is noise, not
information. This applies to `CONTEXT` and `COST` only (never to `MODEL` or `AGE`), is
computed per frame from the visible rows, and is therefore deterministic. The help
overlay lists any column hidden this way and why.
### 7.3 Target renders
These are the golden-test targets, pasted from the generated files. All data is
synthesized.
**A — wide, healthy (120 cols).** The reference render. Rendered with a *synthetic*
vendor that declares every capability, so the whole grid is exercised; render I below
shows what the real v1 capability mix produces.
```
telltale │ 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL CONTEXT COST AGE
● CC │ telltale C:\src\code Opus 5 █████████▎── 84.2% $2.41 │ 12s
● CC │ acme-api C:\src\work Sonnet 4.5 ████▌─────── 41% $0.18 │ 48s
◐ CX │ notes-api C:\src\code gpt-5.1-codex — — │ 4m
○ CC │ learning-notes C:\src\code Haiku 4.5 ██████████▏─ 92.6% $11.07 │ 22m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys
```
Row 3 is the honest-gauge case in its normal habitat: a session whose adapter can source
a model and an age but not context or cost. Blank gauge field, `—` in both numeric
columns. Nothing about it looks like zero.
**B — compact (80 cols).** Cost gone, gauge halved, quota wrapped to its own line.
```
telltale │ 4 sessions │ claude 3 codex 1
claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h
──────────────────────────────────────────────────────────────────────────────
SESSION MODEL CONTEXT AGE
● CC │ telltale C:\src\code Opus 5 █████▉── 84.2% │ 12s
● CC │ acme-api C:\src\work Sonnet 4.5 ██▉───── 41% │ 48s
◐ CX │ notes-api C:\src\code gpt-5.1-codex — │ 4m
○ CC │ learning-notes C:\src\code Haiku 4.5 ██████▌─ 92.6% │ 22m
──────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage ? keys
```
**C — narrow (72 cols).** Gauge gone; the number it encoded stays. Vendor names shorten.
```
telltale │ 4 sessions │ cc 3 cx 1
claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h
──────────────────────────────────────────────────────────────────────
SESSION MODEL CTX AGE
● CC │ telltale C:\src\code Opus 5 84.2% │ 12s
● CC │ acme-api C:\src\work Sonnet 4.5 41% │ 48s
◐ CX │ notes-api C:\src\code gpt-5.1-codex — │ 4m
○ CC │ learning-notes C:\src\code Haiku 4.5 92.6% │ 22m
──────────────────────────────────────────────────────────────────────
q quit / find ? keys
```
**D — degraded rows (120 cols).** Four distinct failure shapes in one frame. Rows are
sorted by activity, so they do not appear in the order they are described.
```
telltale │ 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL CONTEXT COST AGE
● CC │ telltale C:\src\code Opus 5 ──────────── 0% $0.04 │ 3s
● CC │ a-really-long-project-name-that-overflows-the-label-column… Opus 5 ███████████─ 99.9% $340.50 │ 9s
◐ CX │ 4f2a9c81-1d3e-4a77-9b02-000000000000 — — │ 7m
CC │ acme-api C:\src\work Sonnet 4.5 — $1.02 │ —
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys
```
Row 1 is at exactly 0% and draws a full track. Row 2 is the label-overflow case,
truncated at `…` with the grid intact. Row 3 is a session discovered by filename whose
only record was torn, so nothing parsed — the label falls back to the session id, the
model cell is blank and every sourced field is `—`. Row 4 has a record timestamp in the
future (clock skew): its `AGE` is `—`, never a negative or a zero, and because its
liveness is `unknown` **its state dot is blank, not `○`** — unknown is not a claim of
staleness.
**J — zero versus absent (120 cols).** The single assertion the build exists to protect.
```
telltale │ 2 sessions │ claude 2
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL CONTEXT COST AGE
● CC │ at-zero C:\src\code Opus 5 ──────────── 0% $0.00 │ 5s
● CC │ no-source C:\src\code Opus 5 — — │ 6s
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys
```
`0%` is a full track of `────────────`; absent is whitespace. If these two rows ever
render the same, the build fails.
**E — stale scan (120 cols).** The scan has been failing for 47 seconds. Values are the
last ones actually measured; the whole row area renders `Muted` (invisible in a plain
golden, asserted separately), and the footer's right slot carries the notice. The header
is never used for notices — it holds identity and quota only, which keeps it from
overflowing at any width.
```
telltale │ 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL CONTEXT COST AGE
● CC │ telltale C:\src\code Opus 5 █████████▎── 84.2% $2.41 │ 12s
● CC │ acme-api C:\src\work Sonnet 4.5 ████▌─────── 41% $0.18 │ 48s
◐ CX │ notes-api C:\src\code gpt-5.1-codex — — │ 4m
○ CC │ learning-notes C:\src\code Haiku 4.5 ██████████▏─ 92.6% $11.07 │ 22m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
⚠ last scan 47s ago Access is denied.
```
**F — filter and sort active (120 cols).** The header count reads `3 of 4` so it cannot
contradict the per-vendor totals beside it. Non-default filter/sort is stated in the
footer, because a monitor that silently hides rows is a liar.
```
telltale │ 3 of 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL CONTEXT COST AGE
○ CC │ learning-notes C:\src\code Haiku 4.5 ██████████▏─ 92.6% $11.07 │ 22m
● CC │ telltale C:\src\code Opus 5 █████████▎── 84.2% $2.41 │ 12s
● CC │ acme-api C:\src\work Sonnet 4.5 ████▌─────── 41% $0.18 │ 48s
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys filter claude sort context
```
**G — empty (120 cols).** Distinguishes "watching, found nothing" from "vendor not
installed". Two different facts; two different words; never a fake row and never an
error dialog.
```
telltale │ 0 sessions
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
no active sessions
agy not detected %USERPROFILE%\.gemini\antigravity-cli
claude watching %USERPROFILE%\.claude\projects
codex not detected %USERPROFILE%\.codex
cursor not detected %APPDATA%\Cursor\User
gemini not detected %USERPROFILE%\.gemini\tmp
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys
```
The vendor status word is one of exactly four: `watching` (directory exists and is
readable), `not detected` (directory absent), `unreadable` (the vendor's data is there and
the adapter cannot read it — an OS refusal, or a store whose schema the adapter does not
recognize (§3.9); rendered `SevWarn` with the reason appended), `drifted` (the store
opened and read, and at least one session's read could not find the structure the adapter
was verified against — `internal/adapter/drift`; also `SevWarn`). On the dev machine today
the Codex line reads `not detected`, since `~/.codex` is absent (§3.2). The third word:
```
telltale │ 0 sessions
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
no active sessions
agy not detected %USERPROFILE%\.gemini\antigravity-cli
claude unreadable %USERPROFILE%\.claude\projects Access is denied.
codex not detected %USERPROFILE%\.codex
cursor not detected %APPDATA%\Cursor\User
gemini not detected %USERPROFILE%\.gemini\tmp
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys
```
`drifted` is the word the other three cannot say. Those three are all answers from
`Discover`, known before a single session is read. Drift is knowable only *after* the
read: it is the finding that a store which opened, listed and parsed no longer carries the
structure the adapter's readings hang off. Calling that `unreadable` would borrow the word
for "the OS refused" to describe a store the OS handed over intact — the same collapse
`zero` and `absent` are kept apart to avoid (§4a.1).
It renders with its **scope** appended: how many of the vendor's sessions reported drift,
out of how many this scan read. One of forty-one is a vendor mid-rollout; forty-one of
forty-one is a format that moved under the whole store, and the word alone cannot tell
those apart. The scope deliberately does **not** reuse the header's `n of m sessions`
sentence — the header counts visible-of-total across every vendor, this counts
drifted-of-read for one vendor, and the two land on the same screen. In the borrowed
grammar the vendor line would read as a claim about how many sessions are *showing*, which
the header directly contradicts; naming what the numerator counts is what keeps them
apart. The scope is the only part of the line that gives way when the width runs out, the
same way the grid sheds `COST`; the word never does.
This state needs sessions to exist and every one of them to be hidden — below, by the
8-hour idle cutoff — because a vendor cannot drift without having produced the sessions
that revealed it. The ordinary case is render M. The fourth word:
```
telltale │ 0 of 2 sessions │ codex 2
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
no active sessions
agy not detected %USERPROFILE%\.gemini\antigravity-cli
claude not detected %USERPROFILE%\.claude\projects
codex drifted %USERPROFILE%\.codex 1 drifted of 2 read
cursor not detected %APPDATA%\Cursor\User
gemini not detected %USERPROFILE%\.gemini\tmp
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys ⚠ codex drifted
```
**H — help overlay (120 cols).** Replaces the row area rather than floating over it; a
floating panel on a monitor obscures the thing being monitored.
```
telltale │ 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit (also ctrl+c)
↑/↓ move the selection (also j / k)
enter open the detail pane for the selected session
u what each vendor has left, and what it spent
w this week: the fleet's slow windows only
/ find: narrow rows by name or path
esc close the pane, or cancel the find, or quit
v vendor: all > claude > codex >
gemini > agy > cursor
s sort: activity > context > cost
a show all (include sessions idle > 8h)
r rescan now
? close this help
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
? close
```
**I — the v1 capability mix (120 cols).** What the real adapters actually render today:
Claude sources neither context nor cost from disk, Codex sources a **derived** context
percentage (marked `~`) and real quota windows. Nothing sources cost, so the `COST`
column auto-hides and its width returns to `SESSION`. The `AG` row carries no `Name`
at all — the only free text on agy's disk is prompt content (§3.8), so this field is
`CapNone` and the HUD falls back to the workspace basename, the same fallback a Gemini
row takes when it has no summary of its own (ruled 2026-08-12). Its `MODEL` cell
truncates because the vendor's display string is 23 characters against a 13-column
cell — both are what the HUD really shows. The `CU` row is the one to read next to the
`CX` row: both carry a context bar, and only one of them carries a `~`, because Cursor
persists its own `contextUsagePercent` and telltale reads it rather than computing one
(§3.9).
```
telltale │ 5 sessions │ claude 1 codex 1 gemini 1 agy 1 cursor 1 codex 5h ██████▎─ 88.4% ↻ 3h02m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL CONTEXT AGE
● CC │ telltale C:\src\code Opus 5 — │ 12s
● CU │ multi-vendor orchestration C:\src\code composer-2.5 ████▏─────── 37% │ 1m
● CX │ example-app C:\src\code gpt-5.1-codex ███████▋──── ~69.8% │ 1m
● AG │ example-app C:\src\code Gemini 3.6 F… — │ 2m
◐ GE │ glossary tooltips ⑂~2 c:\src\code gemini-3-pro — │ 3m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys
```
This render is the one to look at when judging §6 Q7. The ragged CONTEXT column is the
cost of option (2); it is honest, and whether it is *legible* is a dogfood question.
**K — every column hidden (120 cols).** No visible row reports context or cost.
```
telltale │ 3 sessions │ claude 2 codex 1
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL AGE
● CC │ telltale C:\src\code Opus 5 │ 12s
◐ CX │ notes-api C:\src\code gpt-5.1-codex │ 4m
○ CC │ learning-notes C:\src\code Haiku 4.5 │ 22m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys
```
**L — ASCII glyph mode (120 cols).** `--ascii`, `TELLTALE_ASCII=1`, or a non-terminal
output target. Absent renders `n/a`; the gauge loses its eighth-cell partials, which is a
real precision loss in the bar and acceptable only because the number beside it carries
the precision.
```
telltale | 4 sessions | claude 3 codex 1 claude 5h ###----- 42% ~ 2h13m 7d #------- 18% ~ 5d02h
----------------------------------------------------------------------------------------------------------------------
SESSION MODEL CONTEXT COST AGE
* CC | telltale C:\src\code Opus 5 #########--- 84.2% $2.41 | 12s
* CC | acme-api C:\src\work Sonnet 4.5 #####------- 41% $0.18 | 48s
o CX | notes-api C:\src\code gpt-5.1-codex n/a n/a | 4m
. CC | learning-notes C:\src\code Haiku 4.5 ##########-- 92.6% $11.07 | 22m
----------------------------------------------------------------------------------------------------------------------
q quit / find enter detail u usage w week v vendor s sort a all ? keys
```
**M — shape drift (120 cols).** A store that reads fine and no longer matches. Every row
here is render A's except the Codex one, whose read found no `session_meta` record — so
everything that record feeds is absent, and the row renders *exactly* as it would if Codex
simply had nothing to say. That is the failure: nothing in the grid can tell those two
apart, and the footer notice is the only thing on screen that knows.
It is also why the fourth vendor word needs a second home. The vendor line renders in the
empty state only, and a vendor cannot drift without having produced sessions — so the
screen drift actually happens on is this one, where the vendor line is not present at all.
`driftNotice` therefore renders under **every** body: grid, empty state, help overlay and
detail pane alike. A warning that came and went with whichever pane was open would be one
a reader could not trust to be there.
```
telltale │ 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL CONTEXT COST AGE
● CC │ telltale C:\src\code Opus 5 █████████▎── 84.2% $2.41 │ 12s
● CC │ acme-api C:\src\work Sonnet 4.5 ████▌─────── 41% $0.18 │ 48s
◐ CX │ 00000000-bbbb-4ccc-8ddd-000000000001 — — │ 4m
○ CC │ learning-notes C:\src\code Haiku 4.5 ██████████▏─ 92.6% $11.07 │ 22m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys ⚠ codex drifted
```
### 7.4 The gauge
Glyphs: fill `█` (U+2588) with eighth-block partials `▏▎▍▌▋▊▉` (U+258F–U+2589), track
`─` (U+2500). Full-height fill over a mid-height rule reads as a level above a baseline
and keeps the row visually quiet; a shaded `░` track reads as texture and fights the text.
Three rules, all of which exist to stop the gauge from lying:
1. **The last cell is reserved below 100%.** Fill is computed over `cells−1`, so a value
of 99.9% always leaves one visible track cell and only an exact 100% fills the bar. A
92.6% bar that renders as solid is a gauge claiming "full" when it is not.
2. **Any nonzero value draws at least one eighth.** 0.4% must not be pixel-identical to
0%.
3. **Absent draws nothing.** Not an empty track — nothing. (Principle #1.)
Verified scale at 12 cells (`TestGaugeScale` pins every row):
```
0% ──────────── 0% 84.2% █████████▎── 84.2%
0.4% ▏─────────── 0.4% 92.6% ██████████▏─ 92.6%
5% ▌─────────── 5% 99.9% ███████████─ 99.9%
25% ██▊───────── 25% 100% ████████████ 100%
50% █████▌────── 50% absent —
```
**Number formatting**, shared with the statusline via `internal/theme`:
- Percent: floored to one decimal, never rounded up (a usage gauge must not overstate);
whole numbers drop the decimal; `100%` has no decimal. Guarantees a 5-column field, and
the cell is 6 so a derived value can carry its `~`.
- Cost: `$0.00` under $1000, `$1234` at or above.
- Age: `12s` / `47m` / `2h` / `3d`, capped at 4 columns. Sub-hour precision is where a
monitor's value is; `2h13m` precision belongs to the quota countdown, not to row age.
- Countdown: `↻2h13m`, `↻47m`, `↻5d02h`. The days branch matters: without it a seven-day
window renders `↻120h00m`.
The statusline still uses its own `pct` and `shortDur`; see the divergence note in §2.
### 7.5 Colour and threshold tokens
Thresholds are the statusline's, unchanged and now literally shared:
**green < 60, yellow ≥ 60, red ≥ 85** from `theme.WarnPct` / `theme.CritPct`.
The palette is deliberately the terminal's own 4-bit ANSI palette rather than hex
truecolor. Reason: telltale then inherits whatever theme the user already chose, looks
native in Windows Terminal's default scheme and in a light-background scheme without a
second palette, and matches the statusline byte-for-byte in intent. Total palette: four
hues, one attribute, and the default foreground.
| Token | Meaning | ANSI | Statusline (raw) | HUD (lipgloss v2) |
|---|---|---|---|---|
| `Text` | primary values | default | *(unstyled)* | `NewStyle()` |
| `Muted` | chrome, labels, rules, de-emphasis | — | `\x1b[2m` | `NewStyle().Faint(true)` |
| `Identity` | model name, vendor tag | 6 | `\x1b[36m` | `Foreground(Color("6"))` |
| `SevOK` | value < 60; healthy notices | 2 | `\x1b[32m` | `Foreground(Color("2"))` |
| `SevWarn` | value ≥ 60; warning notices | 3 | `\x1b[33m` | `Foreground(Color("3"))` |
| `SevCrit` | value ≥ 85; error notices | 1 | `\x1b[31m` | `Foreground(Color("1"))` |
| `Track` | unfilled gauge cells | 7 / 8 | n/a | `Foreground(lightDark(Color("7"), Color("8")))` |
Semantic aliases, so intent is greppable rather than inferred: `Absent() = Muted`,
`Rule() = Muted`.
Hue owns exactly one meaning: **cyan is identity, the green/yellow/red ramp is severity,
faint is de-emphasis.** Nothing else gets a colour. In particular the state dot encodes
liveness by *glyph and intensity* (`●` Text / `◐` Text / `○` Muted), never by hue — green
already means "under 60%", and one hue meaning two things is how a colour system rots.
**Shared code, without dragging Lipgloss onto the fast path.** ADR-002 requires the
statusline to stay stdlib-only and never initialize Bubble Tea. So `internal/theme`
holds only numbers and names — `WarnPct`, `CritPct`, the ANSI indices, and the shared
format helpers — and no `Style` type at all. `internal/statusline` maps those indices to
escape codes as it does today; `internal/hud/style.go` maps them to `lipgloss.Style`
values. One source of truth for the thresholds, zero coupling of the statusline's
**import graph** to the TUI stack.
**"Import graph", not "binary" — corrected 2026-08-16.** This paragraph used to claim
zero coupling of the statusline *binary*, and the shipped artifact refutes it:
`go version -m telltale.exe` lists `charm.land/bubbletea/v2 v2.0.8` and
`charm.land/lipgloss/v2 v2.0.5`, because §1 ships ONE binary and `telltale hud` is in
it. §9.8 had already measured the consequence (the 14 MB binary costs ~54 ms per gated
call, against 36.2 ms for a small one) without this sentence being brought into line. What ADR-002 actually buys,
and all it buys, is that the code `telltale statusline` reaches never touches the
framework: no renderer is constructed, no program is started, neither module's package
init runs. That is now gated rather than asserted — see §5's 2026-08-16 amendment.
**Light and dark backgrounds.** Lipgloss v2 removed `AdaptiveColor` and the global
renderer, so adaptation is explicit: `Init()` lifts `tea.RequestBackgroundColor()` into a
`Cmd`, `Update` handles `tea.BackgroundColorMsg` and calls `msg.IsDark()`, and the style
set is rebuilt with `lipgloss.LightDark(isDark)`. Only `Track` consumes it (light gray on
light backgrounds, dark gray on dark). Terminals that never answer the OSC query leave
the default: assume dark. Because exactly one token depends on it and no layout does,
golden layout tests are unaffected by which branch is taken — and
`TestBackgroundColorRebuildsTheStyleSetWithoutMovingTheLayout` enforces that.
**NO_COLOR.** `colorprofile` caps the profile at `Ascii` when `NO_COLOR` is set; Bubble
Tea v2 downsamples internally, so no telltale code path is involved. Under `Ascii` every
`Foreground` disappears while `Faint` survives — chrome still recedes. Nothing is lost,
because by principle #2 colour was never the sole carrier of any distinction. No
`--no-color` flag of our own: one mechanism, the standard one.
**ASCII glyph mode** is a *separate* switch from colour — `--ascii`, or
`TELLTALE_ASCII=1`. For legacy consoles and non-UTF-8 code pages:
| Unicode | ASCII | | Unicode | ASCII |
|---|---|---|---|---|
| `●` `◐` `○` | `*` `o` `.` | | `─` (light rule / track) | `-` |
| `━` (heavy rule) | `=` | | | |
| `█` + eighths | `#` (no partials) | | `│` | `\|` |
| `—` (absent) | `n/a` | | `…` | `>` |
| `↻` | `~` | | `⌥name` | `(name)` |
| `⚠` | `!` | | spinner | `-\|/` rotation |
### 7.6 Motion
**The rule: telltale may animate its own work; it must never animate the vendor's.**
Everything follows from it. A spinner on a session row would assert "this agent is
working right now" — a claim telltale cannot source, since the adapters read files on
disk and know a last-write timestamp, not liveness. That is a narrated animation, which
is the honest-gauge violation in motion form. A tweened gauge is worse: every
intermediate frame displays a value no vendor ever reported.
**Animates — one thing.** A braille spinner `⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏` at 10 fps in the header, only
while the *first* scan is in flight, and only once that scan exceeds 250 ms. It reports
telltale's own I/O, which telltale is entitled to describe. After the first successful
scan it never appears again; later slow scans surface as staleness (§7.7), not motion.
`TestSpinnerStopsForeverAfterTheFirstScan` pins that.
**Never animates.** Numbers. Gauge fills. Colour pulses or flashes. Row insertion and
removal. Sort transitions. Per-row spinners or activity indicators of any kind.
**Cadence.** Poll every 1000 ms via `tea.Tick`. Rows re-sort only when the underlying
values change, and never as a tie-break wobble — the sort comparator falls back to
session id so equal keys hold a stable order frame to frame. Steady-state churn budget is
principle #4: the `AGE` cell of sessions younger than 60 s, and nothing else. The
footer's scan-freshness notice appears *only when abnormal* (> 3 s), precisely so the
healthy screen has no ticking element in it.
The 1 s cadence is affordable because the poll is stat-first: `Discover` lists and stats
only, and `Read` touches a bounded head and tail rather than a whole transcript. Any
further tail-read optimization must honour §4 — a backward seek can land mid-record, so
the first partial record after a seek is discarded, not parsed. `internal/jsonl.Tail`
already does this.
**Bubble Tea v2 implications** (verified against v2.0.8):
- `Model` is `Init() Cmd`, `Update(Msg) (Model, Cmd)`, `View() View`. `View` is a struct,
so alt-screen and cursor are view state, not program options: return a `tea.View` with
`AltScreen: true`, `Cursor: nil`, `WindowTitle: "telltale"` (`--no-title` suppresses
the title). There is no `WithAltScreen` in v2.
- `tea.RequestBackgroundColor()` is a `Msg`, not a `Cmd`; it has to be lifted into a
`func() tea.Msg` before `tea.Batch` will take it.
- **Scanning never happens inside `Update`.** The tick dispatches a `tea.Cmd`; the scan
returns a `scanResultMsg`. At most one scan is in flight and ticks arriving during a
scan are dropped. This is a Windows correctness requirement, not tidiness: a `stat`
against a disconnected network path blocks, and a blocked `Update` freezes input
including `q`.
- **`View()` is pure and never calls `time.Now()`.** The model carries `now`, stamped when
the tick arrives — the same discipline as `statusline.Options.Now`. This is what makes
the renders in §7.3 testable at all.
- Do not raise the renderer FPS. Frames are byte-identical between data changes, so the
framerate cap is not the thing limiting redraws — the data is.
- Render nothing until the first `tea.WindowSizeMsg` arrives. One blank frame beats one
frame of wrong layout.
### 7.7 Degraded and empty states
Each row below is an eval fixture. Fixtures are synthesized — fake session ids, fake
paths, fake text — never copied from real transcripts.
| Fixture | Condition | Where it is asserted | Assertion |
|---|---|---|---|
| `zero-vs-absent` | one session at 0%, one with no context source | golden J + `TestAbsentGaugeIsNotAnEmptyTrack` | The two rows differ. The build fails if a gauge cannot tell "no data" from "zero". |
| `degraded` row 3 | adapter sources model + age only | golden D | Blank gauge field, `—` in `CONTEXT` and `COST`; no `0`, no stale carry-over. |
| torn tail | file ends mid-record, no trailing `\n` | `TestTornTailChangesNothing` (both adapters) | The partial line is held, never parsed (§4). A torn tail changes nothing on screen. |
| torn-only | the file's *only* record is torn | `TestTornOnlyRecordStillListsWithEverythingAbsent`, golden D row 3 | Session still listed (discovered by filename), label falls back to session id, every sourced field `—`, age from file mtime. |
| clock skew | record timestamp in the future | `TestFutureMtimeDegradesRatherThanClampingToZero`, golden D row 4 | `AGE` = `—`. Never a negative duration, never `0s`. Liveness is `unknown`, so the dot is blank. |
| `stale-scan-47s` | scan older than 3 s | golden E + `TestStaleScanDimsTheRowArea` | Values are retained (they were true at the displayed `AGE`) but nothing renders at full intensity. |
| `stale-scan-90s` | scan failing over 60 s | golden | Notice escalates to `SevCrit` and the header quota goes `Muted` too. Quota is as stale as everything else and must not look fresh. |
| `empty-watching` | vendor dirs readable, no sessions | golden G | `watching` + the path actually checked, home-redacted. |
| vendor missing | `~/.codex` absent | golden G, `not detected` + `TestVendorAbsentBecomesNotDetected` | No fake row, no error state; the other vendor still renders. |
| `empty-unreadable` | dir exists, OS refuses | golden + `TestUnreadableVendorKeepsTheOSMessage` | Third word, distinct from the other two, with the OS message in `SevWarn`. |
| `quota-absent` | API-key login, no `rate_limits` | golden + `TestNullRateLimitsYieldNoQuotaWindows` | Header quota block absent. Mirrors the statusline's load-bearing test — never `5h 0%`. |
| `degraded` row 2 | label longer than the column | golden D | Truncation at `…`, grid intact. |
| `column-hidden` | `CONTEXT` and `COST` absent for every visible row | golden K | Columns dropped, width returned to `SESSION`; help overlay names them. |
| `floor-width` / `-height` | 52 cols / 4 rows | goldens | One line, no partial grid. |
| gauge scale | the §7.4 table | `TestGaugeScale` | Exact glyph string per value: the reserve-last-cell and min-eighth rules. |
| separator injection | a session name containing U+2028/U+2029 | `TestSessionNameSeparatorsCannotTearTheGrid`, `TestDetailPaneSanitizesModelAuthoredText`, `TestFindQueryCannotTearTheFooter` | The character never reaches the frame — grid, pane or footer — and no line exceeds the terminal width. |
Added in v1.1:
| Fixture | Condition | Where it is asserted | Assertion |
|---|---|---|---|
| `detail-pane` | pane over a Claude row, real capability table | golden + `TestDetailPaneSeparatesCantKnowFromAbsentNow` | Fields Claude declares `CapNone` get **no line**; they are named once on `not sourced`. |
| `detail-degraded` | pane over a session whose records did not parse | golden + `TestDetailPaneShowsDegradedFieldsAndDiagnostics` | Degraded field names and every diagnostic are on screen; a declared-but-empty quota is `—`, never `0%`. |
| clean session | no degraded fields, no diagnostics | `TestDetailPaneStatesTheAbsenceOfProblems` | The honesty block says `—` rather than going blank; a blank block is indistinguishable from a pane that forgot to render it. |
| measured zero fan-out | `Subagents = 0` | `TestDetailPaneStatesAMeasuredZeroFanOut`, `TestSubagentChipOnlyAppearsForANonzeroCount` | Grid draws **no chip**; the pane says `~0 recent`. |
| uncountable fan-out | sidecar unreadable | `TestDetailPaneRendersAnUncountableFanOutAsAbsent` | `—`, never `0`. |
| selection vanishes | the selected session ends mid-poll | `TestASelectedSessionThatVanishesClosesThePane`, `TestDetailPaneSaysSoWhenItsSessionIsGone` | The pane closes rather than retargeting; an out-of-range cursor says "no longer listed". |
| re-sort under the cursor | a bottom row becomes the newest | `TestSelectionFollowsTheSessionNotTheIndex` | The selection follows the **session key**, not the index. |
| `row-grammar` | selection mark + fan-out chips | golden + `TestSelectionIsAGlyphNotAHighlight` | Selection is a glyph in the pad column, not reverse video. |
| chip vs. truncation | a 73-character session name | `TestSubagentChipSurvivesLabelTruncation` | The chip survives at every width; the name gives way. |
| `burn-forecast` | 7 samples over 18 min on one window, a near-flat second window | golden + `TestForecastArithmeticIsPinned` | Exact projected time and basis; the slow window renders **nothing**. |
| below basis | < 3 samples, or a span < 5 min | `TestForecastRefusesToProjectBelowTheMinimumBasis`, `TestNoForecastRendersWithoutABasis` | Nothing renders. Not a placeholder, not a dash — the header cell simply ends. |
| window rollover | usage drops, or `resets_at` jumps a window forward | `TestUsageDropClearsTheSamples`, `TestResetsAtJumpClearsTheSamplesButJitterDoesNot` | The buffer clears; three seconds of `resets_at` jitter does not clear it. |
| `find-active` | find mode with a query typed | golden | The footer becomes the query line and says how to leave. |
| `find-applied` | query applied, mode left | golden + `TestAnAppliedQueryAlwaysAnnouncesItself` | Header reads `2 of 4`; footer keeps naming the query. |
| query hides everything | a query matching no row | `TestAnEmptyResultNamesTheQuery` | The empty state names the query rather than saying "no active sessions". |
| over-long query | 156 characters at 60–120 cols | `TestALongQueryIsTruncatedNotDropped` | Truncated with `…`, never pushed off the footer — a query that vanished while still filtering is the silent row-hiding the footer exists to prevent. |
| query with a trailing space | `"acme "` | `TestTheDisplayedQueryIsTheQueryBeingMatched` | The string on screen is the string being matched; the display is not trimmed. |
Added with shape-drift reporting:
| Fixture | Condition | Where it is asserted | Assertion |
|---|---|---|---|
| `shape-drift` | a read reports drift; every row still renders | golden M + `TestDriftIsVisibleOnTheGridNotOnlyInTheDetailPane` | The grid is unchanged and the footer carries `⚠ drifted`. A healthy frame never mentions drift, and the notice survives `--ascii` and `NO_COLOR` as a word. |
| `empty-drifted` | sessions exist, all past the idle cutoff | golden + `TestTheDriftScopeCannotBeReadAsTheHeaderCount` | The fourth vendor word, with its scope in the slot `unreadable` gives to the OS message — and in a grammar the header's own count cannot be mistaken for. |
| partial drift | 1 of 41 sessions reports drift | `TestOneDriftedSessionDriftsTheVendor`, `TestTheVendorLineStatesHowMuchOfTheStoreDrifted` | **Any** drifted session drifts the vendor, and the counts travel with the word. Every row still renders. |
| drift under a failed `Discover` | vendor absent, or the OS refuses | `TestTheDiscoverTierStillWinsOverDrift` | `not detected` and `unreadable` are untouched: the roll-up only runs where `Discover` succeeded, so the ordering is structural rather than a comparison. |
| reworded drift note | `drift.Watch` changes its wording | `TestDriftIsRecognizedFromTheNoteTheAdapterLayerActuallyWrites` | The HUD reads drift off `Diagnostics` text, which the compiler cannot check. The test folds a real `drift.Watch` so a rewording fails the build instead of silencing the vendor line. |
| notice pile-up | 60 cols with a 24-char query, a filter, a sort, a stale scan **and** drift | `TestADriftedFrameStillFitsEveryTier`, `TestTheFooterGivesUpItsCheapestNoticesFirst` | Every line fits the terminal. Whole notices are dropped, cheapest first, and `…` says so. |
Freshness escalation, stated once: **≤ 3 s** normal; **> 3 s** row area `Muted` + footer
notice in `SevWarn`; **> 60 s** notice in `SevCrit` and the header quota goes `Muted` too.
Retained values are not "presented as fresh" in any of these, because the age of the
measurement is on screen next to them — that is the condition the honest-gauge rule
actually imposes.
Notice priority, stated once: the footer cannot always hold every notice — `joinEnds` has
no truncation path, so a block that does not fit runs off the end of the terminal. The
block is therefore fitted first, by dropping **whole** notices cheapest-first and
prefixing what survives with `…`. The order is `sort`, `+N more`, `filter`, `find`, the
stale-scan warning, drift. That is the rule `joinEnds` already applies between the key
hints and the notice block, asked one level down: what survives is what the reader cannot
find out anywhere else on this screen. `sort` hides nothing at all; `+N more` sits above a
row area the reader can see is full; a filter and a query hide rows silently, but the
header's `N of M sessions` still declares *that* rows are hidden, so only the cause is
lost; a stale scan re-announces itself every tick and clears the moment a scan succeeds;
drift does neither, and is the last to go. A single notice wider than the whole line is
truncated rather than dropped — an ellipsis on a warning still says a warning is there,
and a footer that dropped its last one would quietly claim nothing is wrong.
#### The zero-config first frame — measured 2026-08-15, then narrowed
The empty states above are all about a machine telltale already lives on. This subsection
is about the frame before that: what a stranger sees on the first run, with nothing
configured and no vendor store anywhere.
**It was measured before anything was built.** A clean profile was made by pointing
`HOME`, `USERPROFILE`, `APPDATA` and `LOCALAPPDATA` at an empty directory and running the
real binary. The isolation was verified rather than assumed — `telltale snapshot` reported
`vendors_not_detected: 6` and `sessions: 0`, so no real session leaked into any reading
below.
| Mode | What a stranger actually saw | Verdict |
|---|---|---|
| `telltale` (bare) | all 203 lines of `usageText`, on **stderr**, exit **2** | **fail** — true and useless. Eight modes, no start-here, and a failure code for typing the binary's own name |
| `telltale hud` | `no active sessions`, then six vendors as `not detected` with the path checked for each | **partial** — every word true and complete, and silent on the reader's next question |
| `telltale doctor` (vendors present) | five seats, binary path, version, and `not checked` said out loud for auth and network | **pass** on truth, no next step |
| `telltale doctor` (bare `PATH`) | five `FAILED` rows under `0 checks passed, 5 failed` | **pass** on truth, and the frame most likely to read as telltale being broken |
| `telltale council` | the room opens; unseatable seats fold out and `collapsedNotice` names each one and why | **pass** — not the zero-config entry point |
The frame the HUD draws on that profile, generated by the build like every render in
§7.3 — this is `empty-nothing-detected`, and the last line is what this subsection added:
```
telltale │ 0 sessions
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
no active sessions
agy not detected %USERPROFILE%\.gemini\antigravity-cli
claude not detected %USERPROFILE%\.claude\projects
codex not detected %USERPROFILE%\.codex
cursor not detected %APPDATA%\Cursor\User
gemini not detected %USERPROFILE%\.gemini\tmp
grok not detected %USERPROFILE%\.grok\sessions
pi not detected %USERPROFILE%\.pi\agent\sessions
telltale doctor checks the vendor binaries; this screen reads their stores
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys
```
**The HUD's gap was not the empty state; it was one question the empty state cannot
answer.** The HUD reads STORES. A vendor CLI that is installed and has never been run has
no store, so `not detected` six times is the same picture on a bare machine and on a
machine holding five unopened vendors. `telltale doctor` resolves the BINARIES, which is
exactly the missing measurement — so the frame now names it.
Three changes, and the narrowing is the design:
1. **A bare `telltale` prints a short first frame on stdout and exits 0.** A no-argument
run is a first run, not a usage error. An unknown SUBCOMMAND keeps stderr and exit 2,
because that one is an error and the manual is its correction. `telltale help` is new
and reaches `usageText` — before this, every route to the manual was an error path, so
the frame could not point at it without inventing a command.
2. **The HUD's empty state carries one pointer line, in one case only:** every vendor
`not detected`, nothing watching, nothing drifted, nothing unreadable. A watching store
needs no remedy; a drifted one needs the adapter re-verified; a store the OS refused
wants a permission fixed, and doctor would report that vendor's binary `ok` and lead
the reader away from the answer. `empty-watching`, `empty-drifted` and `empty-unreadable`
are byte-identical after this change, which is the check that the narrowing held.
3. **`doctor` closes with a `What runs next` paragraph**, branching on the seat count the
summary above already printed and claiming nothing more. The no-seat branch says the
report is working rather than telltale failing, and still leaves a command that runs.
**Nothing invented, at either end (ADR-001).** `firstFrameText` prints before `main` has
stat'd a store or resolved a binary, so it asserts nothing about the machine at all — it
names modes and points at the one that measures, and `TestTheFirstFrameClaimsNothingAboutThisMachine`
keeps it that way. `nextStep` stays pure over its `Report` and never touches auth or
network, which are `not checked` on every seat, always.
| Fixture | Condition | Where it is asserted | Assertion |
|---|---|---|---|
| `empty-nothing-detected` | six vendors, every one `not detected` | golden + `TestTheZeroConfigEmptyStateNamesDoctor` | The frame names `telltale doctor` — the measurement this screen cannot make. |
| the other three empty states | watching / drifted / unreadable | `TestOnlyTheNothingDetectedFrameNamesDoctor` | None of them names doctor. A pointer on every frame is a pointer meaning nothing. |
| no vendor looked for | an empty vendor slice | `TestAnEmptyVendorListIsNotNothingDetected` | Not the same fact as every vendor missing (§4a.1, one level up from the gauge). |
| narrow terminal | the pointer at `MinWidth`–120 cols | `TestTheDoctorPointerIsShedRatherThanTearingTheFrame` | Shed whole, never wrapped and never past the frame — council's `collapsedNotice` rule, for its reason. |
| first frame | `firstFrameText` | `TestTheFirstFrameIsShortAndNamesTheModeThatMeasures` | Under 30 lines, and it names every mode it sends a reader to. |
| doctor next step | no seat ready, and one ready | `TestTheReportSaysWhatToRunNext` | Both branches name a command that runs on the machine just described. |
### 7.8 Keyboard
Minimal, and every key earns its place.
| Key | Action |
|---|---|
| `q`, `ctrl+c` | quit |
| `esc` | close the detail pane → close help → clear the find query → **then** quit |
| `↑`/`↓`, `j`/`k` | move the selection (scrolls the help overlay while it is open) |
| `enter` | open the detail pane for the selected session; close it if open |
| `/` | find: type-to-filter on name or path |
| `v` | vendor filter cycle: all → claude → codex → gemini → agy → cursor → all |
| `s` | sort cycle: activity → context → cost → activity |
| `a` | toggle show-all (default hides sessions idle > 8 h) |
| `r` | rescan now |
| `?` | toggle help |
In **find mode** the keyboard belongs to the query: only `esc` (clear and leave), `enter`
(keep and leave), `backspace` and `ctrl+c` are commands, and everything else is text.
That is why the mode takes over the whole footer — a mode that silently changes what `q`
means without saying so is how a read-only monitor surprises someone.
`--vendor all|claude|codex|gemini|agy|cursor` sets the starting filter; the cycle takes
over from there. `antigravity` is accepted as a synonym for `agy` and `composer` for
`cursor`; the short forms are the ids the footer and the header counts print.
Cycles, not multi-select menus: with six vendors and three sorts, a cycle is one keystroke
and no mode. Non-default filter, sort or query is always visible in the footer. The help
overlay writes the cycle with `>` rather than `->` for one reason worth recording: the
fourth vendor pushed that line past the 60-column floor, and a golden test at that width
is what caught it. The **sixth** vendor exhausted that trick, and the cycle now wraps onto
a continuation line indented under the first hop — shortening the vendor names instead
would have made the overlay teach a name the footer does not print.
> **Reversed in v1.1, deliberately.** v1 said: *"There is no selection cursor — the
> default sort puts the interesting sessions on top, and a cursor invites drill-down,
> which is a different product."* The roadmap (§8) then decided drill-down **is** the
> product: the schema already carried `Diagnostics`, `Degraded` and every `Extra` with no
> surface to show them on, and that machinery is the thing this project is actually
> about. The original objection is answered rather than ignored — the cursor starts at
> **no selection** and the mark appears the first time the user asks for it, so the
> steady-state monitor frame is byte-identical to v1's.
Anything that changes *which rows are visible or in what order* (`v`, `s`, `a`, a new
query) **drops the selection** and closes the pane. The cursor is an index into the
visible rows, so a different row set makes the old index point at a different session.
Between polls the selection is carried by **session key**, not by index, because the
activity sort re-orders rows as sessions write — holding the index would silently move
the selection, and with the pane open would relabel one session's diagnostics with
another's.
Show-all deliberately does **not** hide a session with no activity timestamp: "we have no
signal" is not evidence that a session is old.
Deliberately absent: mouse support, fuzzy/regex/embedding search (the query is displayed
literally, and a syntax that can mean something other than what it looks like is a filter
that hides rows without saying so), and configuration UI. And one invariant that outranks
all future feature requests: **the HUD is strictly read-only. No keybinding may ever
mutate vendor state or send anything to a running agent.** The HUD is a telltale.
That invariant is scoped to the observation surfaces — `hud` and `statusline` — and it does
not weaken. `telltale council` (ADR-008, §9) is a separate subcommand that *does* dispatch to
vendor CLIs; it is a dispatch room, not a gauge, it is entered deliberately, and it says so on
screen. Nothing in §7 may reach for it.
### 7.9 Golden tests
- `Render` is pure over `State` — `(sessions, vendors, now, width, height, filter, sort,
showAll, help, scroll, scanning, thresholds)`. Tests construct the state directly and
compare against `internal/hud/testdata/golden/*.txt` — no terminal, no program loop.
`go test ./internal/hud -update` regenerates them.
- Two families. **Layout goldens** render with `PlainStyles()`, a style set in which every
`Render` is the identity, at widths 120 / 80 / 72 / 52 and at the height floor — so they
never depend on the CI terminal's colour profile. **Style assertions** render with
`NewStyles(true)` and check one escape code per severity band, mirroring
`TestThresholdColors` in `internal/statusline/render_test.go`.
- Width is measured with `lipgloss.Width`, never `len()` — the label column carries
arbitrary project names. `TestNoLineExceedsTheTerminalWidth` sweeps seven widths with
and without the help overlay.
### 7.10 Known limitations
- Every glyph in the visual language — `● ◐ ○ ─ │ █ ▏▎▍▌▋▊▉ … — ↻ ⚠ ▸ ⑂ ·` — is
East-Asian-**Ambiguous** width. Windows Terminal, the reference environment, renders
ambiguous as narrow, which is what the grid assumes. A terminal configured to render
ambiguous glyphs double-width will shear the layout; `--ascii` is the escape hatch.
Stated here rather than discovered later.
- `⑂` (U+2482 OCR FORK) is the least-common glyph in the set and the most likely to miss
from a font. It appears **only** on a session that is fanning out, so a font gap costs
a tofu box on a minority of rows rather than a broken grid — and `--ascii` renders it
`Y`. It was chosen over a second `│` or a bracket because both already mean something
here (`│` separates zones; `]` is the ASCII selection mark).
- The detail pane does not scroll. A pane taller than the row area is clipped, and on a
terminal shorter than about 16 rows a long extras list can run off the bottom. The
arrows are spent on moving between sessions, which is the more valuable binding while
the pane is open; a scrollable pane needs a second axis and is deferred.
- The `█` fill and `─` track differ in glyph height by design. Verified legible in
Cascadia Mono; other fonts may render the step more harshly.
- Fill resolution is one eighth of a cell (1.04% at 12 cells). The number beside the bar
carries the precision; the bar carries the glance.
- ~~The account quota block is sourced from one session (§7.1). A second quota-bearing
vendor needs a per-vendor block.~~ **Closed in two steps.** §7.15 (2026-08-07) gave the
*header* a block per vendor, from the statusline relay alongside the transcript reading.
§7.17 (2026-08-09) gave the per-vendor block its own surface: `u` opens a body with one
block per vendor, the gauge at 20 cells instead of the header's 8, and — the part the
header has no room for at all — a stated reason wherever a vendor has nothing to say.
What remains is not this limitation but §7.17's own: an aged-out relay reading and one
that never arrived render alike.
- ~~The 1 s poll has not been measured on a cold cache over an 837-session tree (§6 Q3).~~
**Measured** — see §6 Q3 and the `BenchmarkScan` table. The cold scan is the half that
did NOT improve, and that is the ruling rather than the residue: the first frame has to
read everything, and the spinner is what covers it.
- The burn forecast's sampling history lives in the process and dies with it. Restarting
the HUD restarts the basis at zero, and for the first five minutes of every run there
is no forecast at all. Persisting samples would mean writing to disk, which "telltale
never writes" forbids, so this limitation is load-bearing rather than an oversight.
### 7.11 The detail pane
**The problem it solves.** v1 carried `Diagnostics`, the `Degraded` field set and every
`Extra` from adapter to renderer and displayed **none of them**. The grid can only draw
one kind of nothing: a dropped column and an em dash both read as "no value here", and
§4a.1 insists there are two different facts underneath. The pane is where the difference
gets said in words. It is the honesty machinery becoming product rather than plumbing.
`enter` opens it on the selected row; `enter` or `esc` closes it. It **replaces** the row
area rather than floating over it, for the same reason the help overlay does — a panel
covering the thing being monitored is a monitor you have to move to read.
**Layout.** Line one is literally the selected row's identity zone (dot, vendor, `│`,
label), so the pane opens where the row was. Everything below hangs off the `SESSION`
column at offset 8: a 12-column muted label, two spaces, then the value. Field order
mirrors the row's three zones — identity, measurement, time — then extras, then the
honesty block, separated by one blank line because it is a different kind of statement.
```
telltale │ 4 sessions │ claude 3 codex 1
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
● CC │ telltale ⑂~2 C:\src\code
session 00000000-aaaa-4bbb-8ccc-000000000001
workspace C:\src\code\telltale
model Opus 5
subagents ~2 recent
activity live · 12s ago
branch main
cli 2.1.219
ctx tokens 215k
degraded —
diagnostics —
not sourced context_pct, cost, quota
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
esc close ↑/↓ session
```
Read the last three lines together, because they are the whole point:
- **`not sourced`** is "can't know" — the fields this vendor declared `CapNone`. They get
no line of their own at all, exactly as the grid drops a column no visible row can
fill. This line is the answer to "why is this row's CONTEXT cell empty?", and it is the
first surface in the product that answers it.
- **`degraded`** is "we tried and failed", named field by field. §4a.2 requires degraded
and plain-absent to render identically in the grid — otherwise "we failed to read it"
starts to look like data — and this is the one place that difference is legible.
- **`diagnostics`** is why. One line per note, structure only, never transcript content.
A clean session prints `—` on both rather than going blank: a blank honesty block is
indistinguishable from a pane that forgot to render one.
Degraded and absent under real failure:
```
telltale │ 4 sessions │ claude 3 codex 1
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
◐ CX │ 4f2a9c81-1d3e-4a77-9b02-000000000000
session 4f2a9c81-1d3e-4a77-9b02-000000000000
workspace —
model —
context —
quota —
activity idle · 7m ago
degraded workspace, context_pct
diagnostics 2 unparseable records skipped
no turn_context record in the read window
not sourced name, cost, subagents
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
esc close ↑/↓ session
```
Every `—` above is a field Codex **can** source and has no value for right now; every
name on `not sourced` is one it never could. Same glyph in the grid, two different facts,
and this is where they separate.
**`activity` is the one line that never renders `—`.** It reports the liveness *class*,
which the HUD can always produce, with the age only as the evidence behind it. A session
with no timestamp reads `unknown`, not `—`: the em dash would say "no value" where the
truthful statement is "no basis for a claim" (§4a.4).
**Selection.** `▸` in the row's leading pad column — the column that was already blank,
so selection costs the grid no width. A glyph rather than reverse video, because §7.1
rule 2 says every distinction is carried by a glyph or a number first and a highlight-only
cursor disappears under `NO_COLOR`. The mark is jammed against the state dot (`▸●`) on
purpose: it reads as a pointer at the row's state, and the alternative is a dedicated
column on every row forever to serve a mark that is off most of the time.
```
telltale │ 4 sessions │ claude 3 codex 1
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL AGE
▸● CC │ telltale ⑂~2 C:\src\code Opus 5 │ 12s
● CC │ acme-api C:\src\work Sonnet 4.5 │ 48s
◐ CX │ 4f2a9c81-1d3e-4a77-9b02-000000000000 │ 7m
○ CC │ learning-notes ⑂~5 C:\src\code Haiku 4.5 │ 22m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys
```
That frame is also §7.13's: row 2 measured **zero** sub-agents and therefore draws no
chip, and the CONTEXT and COST columns are auto-hidden because these are real Claude rows.
### 7.12 The burn-rate forecast
**What makes this ours.** The incumbents in this lane project a burn line against a plan
budget nobody publishes. That is the exact fabrication decisions/001 exists to forbid, and
§8's "deliberately rejected" list names it. telltale instead samples the vendor's own
`used_percentage` **over its own runtime**, reports the slope it measured, marks it
derived, and states the sampling window beside it. The number is telltale's measurement of
telltale's own observations, which is the one kind of computed figure this product is
entitled to show.
Rendered in the header beside the window it describes, never per row (the §7.1 corollary:
quota is a property of the account):
```
telltale │ 4 sessions │ claude 3 codex 1
claude 5h ███───── 42% ↻ 2h13m ~13:27 · 18m basis 7d █▎────── 18% ↻ 5d02h
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL CONTEXT COST AGE
● CC │ telltale C:\src\code Opus 5 █████████▎── 84.2% $2.41 │ 12s
● CC │ acme-api C:\src\work Sonnet 4.5 ████▌─────── 41% $0.18 │ 48s
◐ CX │ notes-api C:\src\code gpt-5.1-codex — — │ 4m
○ CC │ learning-notes C:\src\code Haiku 4.5 ██████████▏─ 92.6% $11.07 │ 22m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys
```
Both windows in that frame have the same 7 samples over the same 18 minutes. The 5h window
is moving fast enough to project; the 7d window renders **nothing at all** — not a dash,
not a placeholder, the cell simply ends. That contrast is the feature.
**The four refusals.** A forecast renders only when all of these hold, and each one exists
to prevent a specific lie:
| Condition | The lie it prevents |
|---|---|
| ≥ **3 samples** spanning ≥ **5 minutes** | Two samples fit a line through themselves and cannot disagree; the third is the first one that can. Five minutes because a stepped percentage sampled over ninety seconds measures the step, not the rate. |
| **positive slope** | "You will never run out" is not a time. A flat or falling window renders nothing rather than an infinity. |
| exhaustion **before the window resets**, when `resets_at` is known | Projecting past the reset describes a window that will not exist. |
| exhaustion **within 24 h** | The render is a wall clock with no date on it. `~04:12` sixteen hours out is misleading, not informative. |
**The arithmetic**, pinned by `TestForecastArithmeticIsPinned`: least-squares slope over
the retained samples, projected from the **last observed value**. Least squares rather
than a first-to-last difference because vendor usage percentages move in steps and a
two-point slope is dominated by whichever endpoints straddle a step. Anchored to the last
observed reading rather than to the fitted line so the projection starts from the number
printed next to it — a forecast that quietly starts from 44% while the cell says 42% is a
small lie in the place this product is least allowed one.
**Sampling.** One sample per *completed* scan (a failed scan contributes nothing rather
than a repeat of the last reading, which would flatten the slope with data we did not
measure), throttled to one every 15 s, bounded to 30 minutes and 128 entries. A window
with a nil `UsedPercent` this scan is a **gap, not a reset** — the history stands.
**Rollover clears the buffer**, on either of two signals: usage dropping (monotonic within
a window, so a drop is a rollover), or `resets_at` jumping forward by more than a minute
(a rollover moves it a whole window; jitter does not). Fitting a line across a rollover
reports a negative rate or a wild one, and every sample before it describes a window that
no longer exists.
**Amendment to §7.1 rule 4** ("still by default"), stated rather than quietly taken: the
forecast cell may change when a new sample lands, which is at most once every 15 s and
only when the measurement itself moved. That is a measurement changing, not an animation —
the §7.6 rule is about telltale never animating the *vendor's* state, and this is telltale
reporting its own arithmetic on a new reading.
### 7.13 The sub-agent chip
`⑂~2` after the session label on any row whose adapter counted recently-written
transcripts in that session's `subagents/` sidecar. Sourced by a stat pass (§3.1), Claude
only.
**Why the `~`.** The count is exact — telltale listed the directory. What is *inferred* is
the 15-minute recency boundary that turns "written lately" into "a fan-out is running
now", and ADR-001 requires the inferred part be visible. So the chip carries the same
estimate marker the CONTEXT column does, and it means the same thing: this number was
computed by telltale, not reported by the vendor.
**Zero draws nothing.** The absence of a chip is not a claim, and a `⑂0` on every Claude
row would be noise asserting a fact nobody asked for — the same reasoning as an absent
gauge drawing no track. The measured zero is not discarded, though: the detail pane says
`~0 recent`, where there is room to distinguish "we counted none" from "we could not
count". A sidecar the OS refuses renders `—` there, never `0`.
Styling: the chip renders in `Text`, not `Muted`. `Muted` is this palette's "chrome or
absent" (§7.5, `Absent() = Muted`), and rendering real measured data in it would put a
sourced number in the same visual class as a missing one.
### 7.14 Type-to-filter
`/` opens the query; typing narrows rows by case-insensitive substring; `enter` keeps the
query and hands the keyboard back; `esc` clears it.
```
telltale │ 2 of 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL CONTEXT COST AGE
● CC │ acme-api C:\src\work Sonnet 4.5 ████▌─────── 41% $0.18 │ 48s
◐ CX │ notes-api C:\src\code gpt-5.1-codex — — │ 4m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
/api_ esc clear enter apply
```
and once applied, with the mode left:
```
telltale │ 2 of 4 sessions │ claude 3 codex 1 claude 5h ███───── 42% ↻ 2h13m 7d █▎────── 18% ↻ 5d02h
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL CONTEXT COST AGE
● CC │ acme-api C:\src\work Sonnet 4.5 ████▌─────── 41% $0.18 │ 48s
◐ CX │ notes-api C:\src\code gpt-5.1-codex — — │ 4m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys find "api"
```
Four rules, all of them the same rule — **a monitor that hides rows must say so**:
1. The header count reads `2 of 4`, so the headline can never contradict the per-vendor
totals beside it. (Same mechanism as the vendor filter; they compose.)
2. An applied query keeps announcing itself in the footer after the mode is gone. A filter
the user has forgotten about hides rows just as silently as one they cannot see.
3. If nothing matches, the empty state says `no sessions matching "zzz"` rather than
`no active sessions` — naming the thing that emptied the list.
4. The match is a literal substring, displayed literally. No globs, no regex, no fuzzy or
embedding search: a syntax that can silently mean something other than what it looks
like is a filter that hides rows without saying so.
**What it matches:** the vendor's session name, the workspace path, and the session id.
The id is in the set because a torn-record row is *labelled* by its id (§7.3 render D row
3), and matching only the "title" would make the one piece of text on that line unable to
find it.
`/` is a mode, and it is the product's only one — which is why it takes over the whole
footer instead of quietly changing what an unmodified key does.
### 7.15 The quota relay — every vendor the header can honestly speak for
Added 2026-08-07, San's ruling. The header's quota block was one vendor (Codex) because
Codex is the only vendor whose quota exists on disk where a passive reader can see it:
Claude's `rate_limits` arrive **only** on its statusline stdin payload (§3.1 — the live
corpus was grepped, nothing quota-shaped reaches the transcripts), and agy's named
buckets exist only in its statusline payload the same way (§3.8). Cursor's store holds
plan-entitlement constants that must never render as usage (§3.9), and Gemini has
nothing — so those two vendors have no quota **anywhere**, relay or not.
**The mechanism: the statusline relays what it just rendered.** After the line is on
stdout, `telltale statusline` writes the payload's quota windows to
`~/.telltale/quota/.json` (`internal/quotacache`), and the HUD's scan reads
every surviving entry alongside the vendor stores. This is a deliberate, scoped
amendment to "the gauges never write" (§1, CLAUDE.md):
- **numbers only, never content** — vendor id, timestamp, window ids/labels,
percentages, reset instants. The same keys-not-content standard as council's
`room.json`, pinned by a test that walks the serialized form field by field.
- **atomic and best-effort** — temp + rename in the same directory, error ignored
after the render is delivered; the cache can never cost a statusline frame or a
torn read.
- **self-expiring** — the reader drops a window whose reset has passed (its
percentage is not stale, it is *false*), and whole entries past 24h or stamped
from the future beyond clock-jitter tolerance.
- **age travels with the reading** — past 5 minutes a relayed block carries
`· 2h ago` at every dress level, the §7.12 basis rule applied to time: shedding
the age would re-present a stale number as fresh. Past `quotaAgeWarn` it stops
being muted chrome and escalates to `· ⚠ stale 19h ago`; §7.17 as amended
argues the threshold and owns both surfaces' wording.
**One block per vendor, transcript outranks relay.** A vendor sourced from its own
store (Codex) is re-measured every scan; its relay entry, if one ever exists, is as old
as the last statusline render. The scan-fresh reading wins and the vendor renders once.
Only the transcript-sourced block may carry a burn forecast — window ids collide across
vendors (Claude and Codex both have a `seven_day`), and re-reading an unchanged cache
file is not a new observation, so a forecast on a relayed block would be one vendor's
slope pinned to another's account.
**The line fits by shedding decoration, never fact.** Dress levels, tried in order
until one fits: full (names, gauges, countdowns, forecasts) → drop forecasts → names
to two-letter tags → drop gauges (the percentage beside each bar says the same thing)
→ drop countdowns. Vendor, window label, reading, and a stale reading's age survive
every level. If even the barest level overflows, whole trailing blocks are dropped and
an ellipsis says so — the footer's dropping-is-never-silent rule. The generated render
(`quota-fleet` golden):
```
telltale │ 1 session │ codex 1
ag gemini-weekly 38% ↻ 3h00m │ cc 5h 42% ↻ 2h13m 7d 6% ↻ 5d00h · 2h ago │ cx 7d 79% ↻ 22h48m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL AGE
◐ CX │ notes-api C:\src\code gpt-5.1-codex │ 4m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys
```
At 120 columns three vendors and four windows already shed to tags-without-gauges;
the full dress needs ~145. That is the honest trade as measured, not a bug: the bars
went first precisely so every fact could stay.
What the relay does **not** change: the statusline's own display (it renders from
stdin as before, the write happens after), the HUD's read-only posture toward *vendor*
files, and the absence rule — a vendor whose statusline never fires simply never
appears, and one that stops firing ages out.
### 7.16 The token relay — what Cursor cost, from a seam with no network call
Added 2026-08-08. §3.9 declared Cursor's `cost` and `quota` **ABSENT**, and re-verifying
it that day made the verdict harder rather than softer: `usageData` was `{}` in 19 of 19
blobs, `tokenCount` was zero in 1,622 of 1,622 message rows, 78 `turn_ended` records and
51 transcripts carried status and no numbers, and the only account figures anywhere on
disk were Statsig experiment values stamped `is_user_in_experiment:false`. Nothing about
consumption reaches the store as a byproduct of a turn.
So the number was fetched rather than found: **Cursor Hooks**, the vendor's own
documented and versioned contract (cursor.com/docs/hooks), whose `afterAgentResponse`
step hands a command hook the turn's token counts on stdin.
**Why the hook and not print mode — the derived-`inputTokens` trap.** `cursor-agent -p
--output-format json` prints a `usage` block, and reaching for it would have been the
obvious move. It is the wrong one: its `inputTokens` is **not** the raw count. The CLI
publishes `max(raw − cacheRead − cacheWrite, 0)`, measured printing **24,076 where the
un-derived input was 48,012**. Rendering that under the label "input tokens" would be
telltale repeating a vendor's arithmetic as if it were a reading — the ADR-001 violation
this project exists to refuse, and a *quieter* one than usual, because the number looks
perfectly plausible. The hook payload carries the vendor's own `tokenUsage` fields
untouched, which is the entire reason it wins.
Source-read at **cursor-agent 2026.08.04-aaa8809** (`8674.index.js`,
`./src/after-agent-hooks.ts`), where the payload is assembled as
`{conversation_id, generation_id, model, text, input_tokens: tokenUsage?.inputTokens,
output_tokens: …, cache_read_tokens: …, cache_write_tokens: …}` and then enriched by the
executor (`190.index.js`) with `hook_event_name`, `cursor_version`, `workspace_roots`,
`session_id`, `transcript_path` and **`user_email`** before it reaches stdin. Both
Windows transports (`argv_heredoc`, the default, and `windows_temp_file`) deliver that
JSON on the command's **stdin**, under PowerShell.
**The payload is the reason the allowlist is a struct.** This is the first telltale seam
where the numbers arrive in the same object as the model's full reply *and* the user's
email address. `internal/cursorhook` decodes into a four-field struct of integer
pointers; `encoding/json` discards everything with no destination, so no content field
can reach the cache unless someone adds a field on purpose — the technique
`internal/adapter/cursor` already uses against a store that keeps OAuth tokens beside
session state (decisions/007), pointed at a payload that keeps PII beside numbers. A test
plants markers in every content-bearing field of a real payload shape and asserts none of
them survives, at the parser AND again on the serialized cache file.
**Three fields were left out on purpose, and they are not content.** `model` and
`generation_id` are per-turn facts and the entry is a TOTAL — naming one turn's model
beside a sum invites reading the sum as that model's. `conversation_id` names a
cursor-agent **CLI** conversation, and the HUD's Cursor rows come from the **IDE's**
Composer store (§3.9); the CLI keeps a separate one. Storing it would dangle a join that
does not exist, which is also why this reading is not rendered on a session row.
**Amended 2026-08-29: the `conversation_id` ruling HOLDS and its reason no longer does.**
The HUD draws CLI rows now, out of `~/.cursor/chats///meta.json` (§3.9's
2026-08-29 addendum), so the join has something to join to for the first time. It is still
not built and the field is still not stored, because whether this `conversation_id` IS that
directory's session uuid was never measured. A key stored on the assumption that two ids
match is how a relay begins attributing one session's tokens to another, and §7.16's whole
argument is that the tokens must not be attributed to a row on a guess. Measure the two ids
against each other first; nothing about the held display changes either way.
#### The accumulation ruling: a total, and never without its window
A hook fires once per agent response, so the file is either the last turn's numbers or a
running total. It is a **running total**, and the price of that choice is that the window
is not optional:
- a single turn's counts answer a question nobody asks — the turn you just watched
finish — and go stale the instant the next one starts. A *counter* is the thing a token
figure wants to be.
- but a sum over an unbounded window is a different and much weaker claim than a reading.
So the entry carries `since` and `turns`, both travel to the screen, and the renderer
may never print the sum without them. "48k" is a number pretending to be a state;
"in 48k · out 1.2k · 14 turns over 12m" is a measurement with its scope attached, the
same §7.12 basis rule the burn forecast and the relayed quota block already follow.
- the window's boundaries are mechanical rather than chosen. Accumulation continues onto
any entry a *reader* would still accept, and opens a fresh window otherwise — first
turn ever, first turn after a day of silence, first turn after a corrupted or
clock-skewed file. `internal/usagecache.readEntry` is shared by `Add` and `ReadAll`
precisely so those two can never disagree; without that, a sum could silently span a
week-long gap and still call itself a total.
- **a partial turn is refused, not part-counted.** A payload missing any of the four
counts is not accumulated at all. Summing the three that arrived and treating the
fourth as zero would leave the total wrong by an amount nothing on screen could name,
while it kept looking like a total; refusing makes the counter go quiet, and a visible
absence is the failure mode §7.7 prefers every time. Every count in the file is
therefore a sum of complete readings, and `turns` says exactly how many.
Everything else is §7.15's mechanism copied deliberately, function for function: one file
per vendor under `~/.telltale/usage/`, atomic temp+rename in the same directory,
best-effort, self-expiring at 24h and on future-skew, and the reading's age travelling
with it past five minutes. `internal/usagecache` is a **sibling package** rather than a
second store inside `internal/quotacache` because the two share their mechanism exactly
and their schema not at all — quota is windows with percentages, resets and a
"reset has passed, so this window no longer exists" rule that means nothing to a counter
— and folding them together would put one keys-not-content test in charge of two
unrelated formats.
#### What it rendered, and what stayed absent
*This subsection describes the display as built on 2026-08-08. It was retired on
2026-08-09 — see the amendment below — and is kept in the past tense because the rules it
worked out still bind the one spend line that remains (§7.17).*
**Tokens spent are not quota, and the render may never blur that.** There is no
denominator anywhere in this reading, so no percentage, no gauge, no countdown, no bar —
any of them would invent a ceiling out of nothing, the same class of error as filling a
`CapNone` field with a plausible guess. The spend block therefore got **its own header
line, never shared with quota at any width**, and carried a verb: `cursor spent`. The
verb is a word, not a glyph or a colour, so `--ascii` and `NO_COLOR` lose none of the
claim. It rendered as:
```
telltale │ 2 sessions │ codex 1 cursor 1 codex 7d █████▌── 79% ↻ 22h48m
cursor spent in 48k · out 1.2k · cache read 1.9M · cache write 62k · 14 turns over 10m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL CONTEXT AGE
```
Codex's quota was on the quota line; Cursor's was nowhere, and Cursor's spend was on a
line of its own. The shed cascade followed §7.15's grammar — cache pair first, then the
turn count, then full names down to the two-letter tags — and vendor, verb, `in`, `out`
and the window survived every level, with the ellipsis saying so whenever anything went.
`theme.Tokens` floors at every step for the same reason `theme.Percent` does: this is
what a machine *spent*, and rounding 47,950 up to "48.0k" invents fifty tokens nobody was
billed for.
#### The display is retired; the relay is not (owner's ruling, 2026-08-09)
**Ruled by the owner: the Cursor spend line comes off every surface. The seam, the hook,
the cache and the HUD's read of it all stay.** The reason was not that the number was
wrong — it was measured end to end and it was right. It was that *it bought no decision*.
A running token count for a vendor with no ceiling anywhere answers a question nobody was
asking, and it was answering it from a header line the header does not have to spare:
§7.15's whole design is a shed cascade fighting for one or two rows, and this was
permanently occupying a third.
That is a product judgement, not an honesty one, and it is worth naming which because the
two have different consequences. Nothing here was retracted. §7.16's measurements, its
vocabulary rules and its accumulation ruling all still stand and all still bind — the
fleet usage view's remaining spend line is held to them (§7.17), and the amendment below
is the first place they were applied to a sum of a different shape.
What changed, exactly:
- **removed:** the header's spend line, and the usage view's Cursor spend row. Nothing
renders a `usagecache.Total` anywhere.
- **kept, deliberately untouched:** `telltale hook cursor`, `internal/cursorhook`,
`internal/usagecache` and every test either owns; `~/.cursor/hooks.json` on this
machine; and the HUD's own read — `Snapshot.Spend` is still filled by every scan and is
read by nothing. `internal/hud/state.go` says so on the field, because a reader who
finds an unused field will otherwise correctly conclude it is dead and delete it.
Reinstating the display is a call site, not a re-plumb, and the accumulating file means
the day it comes back it has history in it rather than starting from this minute.
- **pinned:** `TestTheCursorSpendDisplayIsRetiredEverywhere` renders the same fixture that
produced the old block — relayed total still in the snapshot — at five widths, in both
glyph sets, with the usage view open and closed, and fails on the verb or on either of
the counts appearing anywhere in the frame.
`TestTheRetiredDisplayStillHasItsRelayUnderneath` is the other half, and it is the one
that catches "retired" being implemented as "deleted".
`TestCursorIsStillGivenNoQuotaBlock` survives from the old pair: losing its spend line
is exactly the moment a renderer would be tempted to find Cursor a home on the quota
one.
The same fixture now renders (`cursor-without-spend` golden) — a two-line header where
there were three, Cursor's row still on the grid, and its quota still visibly nowhere:
```
telltale │ 2 sessions │ codex 1 cursor 1 codex 7d █████▌── 79% ↻ 22h48m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL CONTEXT AGE
● CU │ Multi-vendor orchestration C:\src\code composer-2.5 ████▏─────── 37% │ 1m
◐ CX │ notes-api C:\src\code gpt-5.1-codex — │ 4m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys
```
#### The boundary amendment
This is the **third** bounded write on a gauge path, and §1 and `CLAUDE.md` name it
alongside council's `room.json` and the statusline's quota relay. It meets the same bar:
under `~/.telltale/`, numbers and keys only, atomic, best-effort, and pinned by a test
that walks the serialized form field by field. It is also the first one written by
neither gauge — `telltale hook cursor` is its own mode, because a hook's stdout is parsed
by the vendor as a hook result, so this is the one path in the binary where printing
nothing is the contract and every exit is clean.
#### Live verification, 2026-08-08 — and the half of it that did not hold
Measured, on Windows 11 at cursor-agent 2026.08.04-aaa8809:
- the relay path end to end, driven by the real binary from a real turn's real counts
(`input 23,941 / output 33 / cache 0 / 0`, from a print-mode `result` whose zero cache
figures make its derived `inputTokens` equal to the raw one). Two turns relayed;
`~/.telltale/usage/cursor.json` read back
`{"vendor":"cursor","turns":2,"input_tokens":47882,"output_tokens":66,…}` — accumulated,
with no `text`, no `user_email`, no `conversation_id` and no `model` anywhere in it —
and the HUD rendered `cursor spent in 47.8k · out 66 · cache read 0 · cache write 0 ·
2 turns over 6s` from that file.
- **the vendor never invoked the hook, on any surface reachable from a script.** Marker
hooks were installed at `~/.cursor/hooks.json` on both `beforeSubmitPrompt` and
`afterAgentResponse` and neither fired for: `-p --output-format json`, a non-TTY run
without `-p`, or an **ACP** turn driven over JSON-RPC. The ACP result is corroborated by
source rather than resting on one capture — the ACP chunk (`8096.index.js`) contains
**zero** references to `hookExecutor` — and the only call site of the
`afterAgentResponse` helper sits in the React/ink agent-app module, beside the protobuf
agent-server dispatcher the IDE talks to. The config itself was not the problem: `type`
defaults to `"command"` in the bundle's own validator, and an invalid config produced no
`[hooks]` warning on those paths either, which is consistent with the subsystem not
being initialized there at all.
So the seam is real and its payload is verified by source read at a pinned version, and
the **invocation is verified for no surface yet**. What is unverified, itemized: that a
true-TTY interactive `cursor-agent` session fires the hook; that the Cursor IDE's agent
server does; and therefore that any figure ever appears without being fed in by hand. The
hook is left wired at `~/.cursor/hooks.json` so the first session that does fire it
captures a total. Worth stating plainly because it bears on the roadmap: **council's
Cursor seat runs on ACP (§9.36), and ACP does not carry hooks** — so the seat that most
wanted per-turn cost is, at this version, the one surface that provably cannot supply it.
**AMENDED 2026-08-15: the last sentence above is too wide, and the ACP path DOES carry
hooks.** The paragraph is kept whole because its two measurements still stand — neither
`afterAgentResponse` nor `beforeSubmitPrompt` fired, and the ACP chunk really does hold no
`hookExecutor` reference. What is wrong is the generalization from those two events to the
subsystem. **The hook subsystem is live on the ACP path, per event.** Measured at the SAME
build the paragraph above was measured at, `cursor-agent` **2026.08.04-aaa8809**, on Windows
11, driven from a PowerShell parent (`PARITY.md`'s launch-parent trap), with the handshake
`internal/council/vendors/cursoracp.go` builds: `initialize`, `session/new`, `session/prompt`,
and any `session/request_permission` answered `allow-once`. The observable is the filesystem.
| arm | `~/.cursor/hooks.json` | marker created | trials |
|---|---|---|---|
| baseline | telltale's `afterAgentResponse` only, unchanged | **yes** | 1/1 |
| deny | plus one `beforeShellExecution` returning `permission: "deny"` | **no** | 2/2 |
**The hook's own breadcrumb is the proof, and the vendor stamps its build into it.** The
payload arrived carrying `"hook_event_name":"beforeShellExecution"`,
`"command":"mkdir cursor-deny-marker"` and `"cursor_version":"2026.08.04-aaa8809"`. So a
`~/.cursor/hooks.json` hook runs on a turn driven entirely over the ACP wire, and its denial
holds: the directory was never created, on either trial.
**Three records disagreed and each was right about its own event.** §7.16 measured
`afterAgentResponse` and `beforeSubmitPrompt` and found neither firing. `cursoracp.go` and
`PARITY.md` recorded a hook blocking an ACP tool call, which is `beforeShellExecution`. The
research read that found a hook executor in chunk 4414 was reading a chunk the ACP path
reaches, not the ACP chunk itself: at this build `8096.index.js` holds **zero**
`hookExecutor` references and `4414.index.js` holds one, while `beforeShellExecution` appears
in `190.index.js`, `3384.index.js` and `index.js`. A per-event answer reconciles all three
without withdrawing any of them.
**The token relay's own conclusion is unchanged, and it was re-checked rather than assumed.**
Telltale's `afterAgentResponse` entry stayed wired through all three ACP turns above.
`~/.telltale/usage/cursor.json` still reads `"turns":2` with a `written_at` of 2026-08-08, so
that event still does not fire on ACP. The seat that most wanted per-turn cost still cannot
supply it; what changes is only the reason, which is the event rather than the protocol.
**A trap for the next hook written against this vendor.** The payload cursor writes to the
hook's stdin begins with a UTF-8 **BOM**, so a plain JSON parse of it fails. `telltale hook
cursor` is not exposed to this today, because it is the one event that does not fire here.
#### Known limitations
- **Accumulation is a read-modify-write and takes no lock.** Two hook processes finishing
in the same instant can lose one turn. Accepted rather than locked: the loss is bounded
and self-consistent (the `turns` count drops with the counts it names, so the total
never disagrees with its own window), and the alternative is a lock file on a path a
vendor's turn is waiting on — a gauge that can hang a turn is strictly worse than one
that undercounts.
- **`~/.cursor/hooks.json` is un-versioned machine state.** It names an absolute path to
the binary, so it does not travel; versioning it belongs in the dotfiles repo, not here.
- The counter says nothing about *which* conversation or model spent the tokens, by the
design ruling above. If a future seam makes a CLI conversation joinable to a HUD row,
that is a new claim and needs its own section.
#### The statusline seam, measured per surface (2026-08-16)
A third Cursor seam, measured on the same build the two above were measured on —
**cursor-agent 2026.08.04-aaa8809**, Windows 11. `~/.cursor/cli-config.json` accepts a
top-level `statusLine` object using the stdin-JSON/stdout-text contract telltale already
implements twice. §2.2 is the design record and the renderer shipped; this is the
measurement.
**The §7.16 lesson was applied before the fact.** "A config that validates is not a
subsystem that initializes" is what the deny-composition block above cost to learn, so no
surface was reasoned about — each was driven with a marker command that tees its stdin to
a file and exits 0.
| surface | how it was driven | marker fired |
|---|---|---|
| interactive TUI, real console | hidden-window `cmd /c cursor-agent.cmd`, known workspace | **yes** (2 invocations inside 0.4s) |
| ACP | `cursor-agent acp`, `initialize` + `session/new` — the handshake `cursoracp.go` builds | **no** |
| print mode | `cursor-agent -p`, one real turn | **no** |
| interactive argv, piped stdio | no TTY | **no** — the TUI never mounts at all |
| interactive, workspace never opened before | same as row 1, in a fresh git worktree | **no**, 2 trials totalling 55s |
The source read agrees and explains the shape of it: the invocation is a React hook
(`./src/hooks/use-status-line.ts`) called from the ink chat component in `./src/ui.tsx`,
and `statusLine` appears in exactly one bundle chunk (`8674.index.js`). No headless surface
mounts that component. The last row is recorded as measured and its mechanism was not
chased — the fresh worktree had no `~/.cursor/projects/` entry and the workspace that fired
did.
**So this seam repeats §7.16's conclusion rather than relieving it: council's Cursor seat
runs on ACP (§9.36), and this seam is interactive-only.** Two seams now, both real, both
invisible to the one surface that wanted them. The reason is different each time — an
unfired event there, an unmounted component here — and neither generalizes to the other,
which is the mistake the 2026-08-15 amendment above exists to correct.
**No vendor marker, so routing is a flag.** Every captured payload was checked. There is no
`product`, no `hook_event_name`, and no other vendor name; `version` carries the CLI build
string, a value rather than a marker. The payload is Claude-shaped by the vendor's own
statement of intent. §2.1's affirmative-marker scheme therefore cannot be extended, and
`telltale statusline --vendor cursor` is the routing (§2.2 argues why a structural guess
was refused).
**The BOM is NOT on this seam — and the check is now in the code anyway.** The trap note
above records that cursor's HOOK payload begins with a UTF-8 BOM, which breaks a plain JSON
parse. The statusline payload does not: both captures began `7B 22 73` (`{"s`).
`cursorstatus.Parse` strips one regardless, and `TestCursorBOMIsStripped` pins it. One
vendor writing two seams with two encodings is the case where three bytes of caution are
cheaper than rediscovering the difference.
**The ACP `usage` field is parsed and deliberately goes nowhere.** The response schema in
`8096.index.js` declares session/prompt's result as `{stopReason, usage?}`, with usage as
`{inputTokens, outputTokens, totalTokens, cachedReadTokens?, cachedWriteTokens?,
thoughtTokens?}`. The thirteen live arms of §9.36 saw no `usage` on any turn, and that
finding stands — the schema says optional and this build never sent one. `acpTurnEnded` now
parses it into `acpUsage` and **nothing reads it, relays it or renders it.** The reason to
stop there is on the record twice already: print mode's `result.usage` published a CLI-
computed `inputTokens` under that exact name, and the statusline seam does the same thing
today with `total_input_tokens`. Whether ACP's `inputTokens` is raw or derived is an open
question a live capture at a pinned version must close, and displaying it first would be
the violation rather than the discovery of one.
**Hand-back — the one measurement this session could not take.** Every captured payload
came from a session that had made no API call, so `context_window` arrived with all six
keys `null`. The populated shape is therefore documented and synthesized, never observed:
`used_percentage` has been read from the bundle as vendor-sourced, but its live magnitude
and precision have not been seen. Closing it needs a human at a real TTY, which is the one
thing an agent cannot drive here. Roughly a minute:
1. Add to `~/.cursor/cli-config.json`:
`"statusLine":{"type":"command","command":"telltale statusline --vendor cursor"}`
2. Run `cursor-agent` in a workspace it has opened before, and send any one prompt.
3. Read the line above the input. `ctx N%` appearing after the reply is the confirmation;
the segment staying hidden means `used_percentage` is null for longer than assumed.
**Paid, 2026-08-17, at cursor-agent `2026.08.11-e8db854`.** The operator ran the three
steps above. The environment was an interactive session in Windows Terminal, from a
PowerShell parent, in this repository's own workspace. After the first reply the line
read:
```
Composer 2.5 Fast │ ctx 12.7% │ telltale
```
The segment drew, so `used_percentage` populates on a real session. Three things are now
observed rather than assumed. The magnitude is live: `12.7` is a reading, not a null and
not a zero. The precision is **one decimal**, which is the precision the synthesized
fixture already assumed — so the fixture shape stands unchanged. The field populates
**after the first reply**, which answers step 3's alternative outcome; nothing waited
longer than one turn.
**The build is newer than this section's own pin, and that is stated rather than
smoothed.** Every other measurement in this section came from `2026.08.04-aaa8809`,
including the bundle arithmetic that ruled two fields derived. This capture ran at
`2026.08.11-e8db854`, the build §9.46 pinned. So read the populated shape as confirmed at
the newer build, and the derived-field ruling as measured at the older one. The two do not
substitute for each other.
**One limit, named.** This capture confirms that the field carries a number. It does not
re-measure `remaining_percentage` or `total_input_tokens`, because §2.2 keeps both out of
the structs and the line renders neither. A capture cannot observe a field the parser
deletes.
### 7.16a The OTLP collector — what grok spent, pushed rather than hooked (2026-08-10)
§3.9a closed grok's quota question three times over and ended on a seam: the vendor's one
designed-for-reporting surface is an external OpenTelemetry stream that carries **spend and
no quota at all**. This section is that seam, spent. `telltale otel grok` is a loopback-only
OTLP/HTTP listener; grok's own exporter pushes to it, and each `grok_code.api_request`
event's four token counts are folded into `~/.telltale/usage/grok.json` — the same cache,
accumulation ruling and refusal gates as §7.16, fed by a push instead of a hook. The
measured export shapes, capture environment and grok version are pinned in §3.9a's export
addendum; `internal/grokotel`'s package doc carries the same facts beside the code.
**Why a listening socket does not breach §4a.5.** The adapter contract's "no network
calls" protects two things: a gauge that can stall on a wire, and a gauge that can *reach
out* — toward an endpoint, with credentials, spending from the pool it reads. The
collector does neither, and the direction of the arrow is the whole argument: telltale
opens a socket on 127.0.0.1 and **the push is grok's**, exactly as §3.9a recorded when it
named the seam. The gauges still make no network calls and read no credentials; they read
the FILE this mode writes, exactly as they read the hook relay's. It is its own mode for
the same reason `telltale hook cursor` is (§7.16's boundary amendment): its I/O contract —
a foreground server holding a port — belongs to neither gauge. The bind refuses any
non-loopback address at startup, mechanically: a collector reachable off-box would be an
open door wearing a gauge's name.
**One source, chosen over a redundant second — and there turned out to be a third.** The
rule below is scoped to the two envelopes on the wire, which is all that was known here in
2026-08-10. A 2026-08-29 re-measure found the same per-turn counts on grok's DISK as well
(§3.9a's `usage` re-measure), so the rule now has a wider job; the amendment at the end of
this section states it. The stream carries the same counts twice —
per-request on `api_request` events, aggregated on the `token.usage` metric — and §3.9a's
capture measured them value-for-value equal. The collector reads the EVENTS and
acknowledges `/v1/metrics` without reading it: one record is one claim, an event carries
all four counts atomically (so §7.16's complete-or-refused gate maps onto it unchanged),
and reading both envelopes would be two chances to count one number. The 200 on the
unread path matters — an unacknowledged export is retried, and making the exporter loop
on a signal nobody reads would spend grok's batches on nothing.
**The window unit is the api request, and the entry says so.** grok's counts arrive per
API call, not per turn (`turn_completed` carries no counts — measured), so the cache
entry's window count is `requests`, a new sibling of `turns` in the §7.16 schema. The
same amendment gave the schema `reasoning_tokens` and made two fields *optional with
their absence meaning something*: a cursor entry carries `cache_write_tokens` and no
`reasoning_tokens`, a grok entry the reverse, because each vendor's file may only claim
the counts its vendor keeps — §4a.1's zero-versus-absent rule, applied to the serialized
form. `TestTheGrokShapedEntryCarriesItsOwnKeysOnly` and the cursor keys test pin both
shapes field by field.
**What the wire carries and what survives.** Every record arrives with `session.id`,
`user.id`, `team.id`, `model` and timing beside the counts; with a content gate open it
would carry prompt text. Four counts survive. The extraction is an allowlist the same way
`internal/cursorhook`'s struct is — an attribute key with no case in the parser falls
through unread — and `TestNothingFromTheWireReachesDisk` plants content markers on a real
record shape (plus a gate-open `user_prompt` event) and asserts nothing but the numbers
reaches the file. `session.id` and `event.sequence` are read into collector *memory* for
one purpose: the exporter retries unacknowledged batches, and a total that counts a
retried batch twice is overstated by an amount nothing can name. A replayed
(session, sequence) pair is refused; the guard is never written to disk.
**The display is held, and it is the owner's own ruling applied.** §7.16's amendment
retired the cursor spend line because a running count for a vendor with no ceiling
anywhere buys no decision — and grok is *more* ceiling-less than Cursor, not less: §3.9a
swept its disk twice, probed the free network half and read the vendor's own monitoring
schema, and no quota exists anywhere. §7.17's Declined already refused grok a spend line
sourced from disk ticks. So this relay ships exactly as the cursor one now stands: write,
cache and the HUD's read of it wired (`Snapshot.Spend` carries the entry; nothing renders
it), display a call site away, and the accumulating file means the day a display is ever
ruled in it has history rather than starting from that minute.
`TestTheGrokSpendRelayRendersNowhere` pins the hold at every width, in both glyph sets,
with the usage view open and closed.
**Wiring it on a machine** (the enable is machine-local config, deliberately not in this
repo):
```toml
# ~/.grok/config.toml — grok's double opt-in, pointed at the default local endpoint
[telemetry]
otel_enabled = true
otel_logs_exporter = "otlp"
```
then leave the collector running while grok runs:
```
telltale otel grok
```
It listens on 127.0.0.1:4318 (OTLP's http default, so the zero-flag pairing finds
itself; `--addr` moves it, loopback only) and prints one line per counted request. The
content gates (`otel_log_user_prompts`, `otel_log_tool_details`) stay off; the collector
keeps nothing they would add, and the planted-marker test is the proof, but a
content-free wire is strictly better than a filtered one. Verified end to end on
2026-08-10: a config-driven `grok -p "hi"` (grok 1.0.0 (3cd0d0cbce), no env overrides,
default batch intervals) against the running collector produced
`{"vendor":"grok","requests":1,"input_tokens":23767,"output_tokens":96,
"cache_read_tokens":1408,"reasoning_tokens":81,…}` — real numbers, keys only.
**Amended 2026-08-16 — the collision on 4318, measured and then named.** 4318 is OTLP/HTTP's
registered port. That is why this mode defaults to it, and it is also why every other local
OTLP receiver takes it: Jaeger, the OpenTelemetry Collector and the vendor agents all default
there. So the most likely startup on a working machine is the one that cannot bind, and
nobody had measured what that looked like. **Measured 2026-08-16**, Windows 11, `main` at
`4e0cf6b`, with a throwaway listener holding 127.0.0.1:4318: `telltale otel grok` printed one
line on stderr and exited 1.
```
telltale otel: listen tcp 127.0.0.1:4318: bind: Only one usage of each socket address (protocol/network address/port) is normally permitted.
```
So the failure was already loud and already correctly coded. Nothing pretended to collect,
and nothing hung. What the line did not carry is what to do next, and there are three parts
to that: the likely holder is another OTLP collector, `--addr` moves this side (the flag
already existed and was already documented), and moving this side ALONE counts nothing,
because grok's exporter goes on posting to 4318. A collector listening on a port nobody
pushes to reads exactly like "grok spent nothing", which is the failure §7.7 rates worst.
The message now states all three and keeps the bind error verbatim underneath it. On a port
the operator chose it names no likely holder, because telltale cannot know who took 4444, and
it says that instead of guessing. `TestAHeldPortSaysWhoLikelyHasItAndHowToMove` and
`TestTheDefaultPortCollisionNamesTheOtherCollectors` pin both branches;
`TestAMovedPortBindsAndCountsARequest` pins that the way out works, and
`TestAMovedPortIsStillLoopbackOnly` keeps the loopback bind absolute across the flag.
The redirect the message prescribes is `OTEL_EXPORTER_OTLP_ENDPOINT=http://` in grok's
own environment, beside the `[telemetry]` pair above. **That redirect is NOT measured, and
the message says so on the line that prescribes it.** The capture pinned the exporter as
OTel-OTLP-Exporter-Rust/0.32.0 posting to the default endpoint with no variable set; nothing
here re-ran an export with the variable moved. A named knob marked unverified beats no knob
at all, and marking it is §4a.1's estimate rule applied to a sentence instead of a number.
One Windows detail earns its line, because the portable-looking version of it is wrong.
`errors.Is(err, syscall.EADDRINUSE)` is FALSE on Windows for a real collision: the bind
returns errno 10048 (`WSAEADDRINUSE`), while Windows builds define `syscall.EADDRINUSE` as
one of Go's synthetic `APPLICATION_ERROR` constants, 536870914 on this box (measured, go
1.26). Windows is the primary target (ADR-002), so a one-arm check would have detected the
collision on the two platforms CI does not run and missed it on the one it does. The
detection carries both arms and cites the measurement beside them.
#### Known limitations
- **The collector must be running to hear the push.** grok's exporter retries briefly and
then drops a batch; spend accrued while the collector is down is not counted later. The
counter goes quiet rather than drifting — the §7.7-preferred failure — but "quiet"
here can also look like "nothing spent", and only the window's `since` says how long
the file has been accumulating.
- **The replay guard is memory-only.** A batch retried across a collector restart is
counted twice; bounded by one batch, and accepted for the same reason §7.16 accepted
its write race — a guard file would be a second store keyed on session ids.
- **A record without `session.id` or `event.sequence` is counted unguarded** rather than
refused: both ids exist on every measured record, and if a later grok drops them the
honest failure is a counter exposed to duplicate retries, not one that silently stops.
- **The schema is the vendor's alpha (`grok_code.schema.version = v1`)** and the
collector does not read the version attribute. A rename lands as quiet non-counting —
visible as a counter that stops moving, and §3.9a's capture is the shape to re-measure
against.
- **The endpoint redirect is prescribed, not measured** (2026-08-16 amendment). Moving the
collector with `--addr` is measured; moving grok's exporter to meet it rests on the
exporter library's documented variable, not on a re-run capture. The startup message and
this section both say so, so nobody quotes it as verified later.
- The capture behind every claim here is one machine, one day, one grok version, one
signed-in account. The §3.4 discipline applies: re-measure before extending any claim.
**Amended 2026-08-16 — who may push here.** The listener above took any loopback POST, and
"loopback" was carrying more weight than it could hold: measured the same day, a web page on
another origin planted a forged `api_request` in `usage/grok.json` from a real headless
Chrome, with no local code running at all. §7.24 is the measurement and the fix. Two things
changed here: `/v1/logs` and `/v1/metrics` now refuse a request carrying `Origin`, and both
require `Content-Type: application/x-protobuf` — the media type this section's own capture
pinned on grok's exporter, so a correctly configured grok is unaffected. A local *program*
is still trusted completely and deliberately, because it can write the cache file directly;
§7.24 states that boundary rather than pretending a token would move it.
**Amended 2026-08-29 — this listener is no longer the only way to read what grok spent, and
the one-envelope rule is what that costs.** §3.9a's `usage` re-measure at grok 1.0.5 found
`inputTokens`, `outputTokens`, `totalTokens`, `cachedReadTokens`, `cacheCreationTokens`,
`reasoningTokens`, `modelCalls` and `apiDurationMs` sitting beside `costUsdTicks` on every
`turn_completed` record on disk — present since 1.0.0, and missed only because §3.9a's cost
sweep was spelled `"[a-z_]*cost[a-z_]*"` and could not match them. **Nothing in this section
is falsified by that.** The opening sentence claims the OTLP stream is the vendor's one
*designed-for-reporting* surface and that is still true; a session-update log persisted
verbatim is not a reporting surface. What is no longer true is the unstated corollary a
reader would draw, that the push is the only way telltale **could** obtain grok's per-turn
counts. It is not. The disk is a second source and a passive one: nothing to run, no double
opt-in, no port to hold, and no batch lost while a collector is down — which is this
section's first Known limitation, answered by a seam it did not know it had.
Three constraints bind any future disk reader, and they are recorded now precisely because
nothing is being built here.
1. **One envelope per turn, and the replay guard cannot enforce it across seams.** The
guard above refuses a repeated `(session.id, event.sequence)` pair, in collector memory,
over OTLP records. A disk reader keys on a file and an offset and would be invisible to
it, so a machine running both would fold one turn into `~/.telltale/usage/grok.json`
twice and the file would be overstated by an amount nothing could name. That is the
failure §7.7 rates worst, arriving through the front door. **A disk reader and this
listener must never both count the same turn.** Which of the two is the source is a
design question, and this amendment does not answer it — it only forbids the answer
"both".
2. **`input` does not mean the same thing on the two seams.** Measured on one turn: the wire
reports `input_tokens: 22516` beside `cache_read_input_tokens: 256`, and the disk reports
`inputTokens: 22772` for that same turn, which is the sum. The wire excludes the cache
read and the disk includes it. This cache's `input_tokens` field currently holds the wire
sense, because this collector is its only writer. A disk reader writing `inputTokens`
into the same field would silently change what the column means, and history already in
the file would not convert.
3. **The read budget is unmeasured and it is the deciding question.** `updates.jsonl`
reached 818 KB in one observed session and is append-only, so whether a complete read is
bounded — not whether the fields exist — is what decides a disk reader. §3.9a's cost row
ruled a tail-window sum a lower bound, and a lower bound accumulated into a total is the
derived-number refusal in a second unit.
**The display stays held** by the owner's ruling above, so none of this has a consumer
today, and the reader is deliberately NOT built in the lane that measured it. Recorded as
the seam, not spent — the same way §3.9a recorded this listener's own seam before it was
spent.
### 7.16b The Claude statusline's token block — measured, modelled, relayed nowhere (2026-08-16)
The token relay has two writers (§7.16, §7.16a) and an obvious-looking third. Claude Code's
statusline payload grew token fields, and the Claude statusline is **already a relay writer**
— it writes the quota it just rendered (§7.15). A writer fed from a payload the gauge is
holding anyway would cost one function and no new seam.
**It was measured first, and the measurement refused it.** This section is the record of a
seam that was NOT spent, which is worth as much as one that was: the fields are modelled and
parsed, nothing renders them, and `~/.telltale/usage/claude.json` is never written.
#### What was measured, and how
Pinned at Claude Code **2.1.233** — `GIT_SHA f8d57569aaf350fe25dc4dfa10cad59db8ea4d45`,
`BUILD_TIME 2026-08-14T17:21:48Z`, `DD_SOURCEMAP_GROUP win32`, the build installed at
`~/.local/share/claude/versions/2.1.233` and the one `claude --version` reports.
**The method was a source read of the shipped bundle, not a live capture, and that is a
weaker instrument in one specific way.** A capture shim was installed at the dispatcher
(`~/.claude/statusline.sh`, backed up and restored byte-identically) and it produced **zero
real payloads** in a fifteen-minute window: the harness session that installed it renders no
statusline, and an idle session does not re-fire one. Rather than leave a modified dispatcher
on the owner's machine, the shim came out and the bundle was read instead. `CLAUDE.md` admits
both instruments ("a live run, a source read at a pinned version") and §7.16's own Cursor
payload rests on exactly this one. What a source read buys here that a capture could not is
the **arithmetic** — a capture shows numbers, and the numbers are not the problem.
The whole block comes from one function:
```js
function TAw(e,t){let r=wMo(e,t);return{
total_input_tokens: e ? e.input_tokens + e.cache_creation_input_tokens
+ e.cache_read_input_tokens : 0,
total_output_tokens: e?.output_tokens ?? 0,
context_window_size: t, current_usage: e,
used_percentage: r.used, remaining_percentage: r.remaining}}
```
and `e` is `TDr(messages)`, which walks the message list **backwards and returns on the first
usage it finds** — the single most recent assistant message:
```js
function TDr(e){for(let t=e.length-1;t>=0;t--){let r=e[t],n=r?QUe(r):void 0;
if(n)return{input_tokens:n.input_tokens,output_tokens:n.output_tokens,
cache_creation_input_tokens:n.cache_creation_input_tokens??0,
cache_read_input_tokens:n.cache_read_input_tokens??0}}return null}
```
The vendor's own doc comments in the same bundle say it in words, and they agree with the
code: `current_usage` is "Token usage from last API call (null if no messages yet)",
`total_input_tokens` is "Input tokens currently in the context window (incl. cache
reads/writes)", `total_output_tokens` is "Output tokens from the most recent API response".
#### What `current_usage` actually is — and why the write half dies
**It is one API call, and the two totals are a LEVEL rather than a counter.**
`total_input_tokens` is not a session spend; it is what is sitting in the context window
right now, which is why it includes cache reads. It is also **derived** — the vendor sums
three fields of the same single call — so it is `§7.16`'s derived-`inputTokens` trap wearing
a different label, and a plausible-looking number is the quiet kind of wrong.
`internal/usagecache` accumulates **per-turn counts taken as-is**. Nothing in this payload is
one. Three routes to a spend figure exist and all three are refused:
- **sum `current_usage` across fires.** Every fire reports the same last call, and the
statusline is debounced at 300ms and re-renders on state changes, so this counts one call
an unbounded number of times.
- **difference successive renders.** This is deriving a number and presenting it as a
reading — the ADR-001 refusal this project exists for — and `usagecache.Delta`'s contract
says the vendor must report the count, not telltale reconstruct it.
- **treat `total_input_tokens` as a total.** It is an occupancy level. It goes **down** after
a `/compact`, and a "total" that decreases is not one.
So the write half is dead on the measurement, not on taste. **The one thing that would revive
it is the vendor reporting a per-turn count** — and the bundle shows it already computes
something close (`U5d` sums input + cache_creation + output across messages, de-duplicated by
id) and does **not** put it in the statusline payload. If a future version does, this ruling
is re-openable against that field and nothing else.
#### What was built
- **`internal/claude`**: `context_window` gains `total_input_tokens`, `total_output_tokens`
and a `current_usage` object (four counts); `StatuslineInput` gains `prompt_id`. All are
pointers or omitempty, because a pre-2.1.233 CLI sends no such key and "this CLI does not
report it" must stay distinguishable from the measured zero `TAw` emits before the first
message (§4a.1). `CurrentUsage` is the allowlist for its block, the `internal/cursorhook`
technique reused.
- **nothing else.** No renderer, no relay, no `usagecache` converter, no HUD field.
- **pinned three ways**: `TestTheTokenCountsParseAtTheMeasuredVersion` asserts the fields land
*and* that the fixture still models `TAw`'s arithmetic — the moment those numbers stop
agreeing, the fixture has begun claiming a cumulative total nobody measured;
`TestAnOlderPayloadLeavesTheTokenCountsAbsent` is zero-vs-absent on a schema that grew;
`TestTheTokenCountsAreParsedAndNeverRendered` guards the render, and CI asserts on the
**built binary** that no count reaches stdout and that `usage/claude.json` does not exist.
#### Two claims from the research brief that the measurement corrected
- **`prompt_id` is real, and it is not on the documentation page.** It arrives via the
vendor's shared session-basics helper (`py`), as `pr.requestJournal.promptId()`, not via the
statusline's own assembly. `internal/claude`'s package doc now says which of its fields are
documented and which are measured-only.
- **`aborted` and `api_retry` appear in NEITHER payload.** Every `aborted` in the bundle is
`AbortSignal`/undici machinery or a telemetry reason string. `api_retry` is a **stream event
type** — "Emitted when an API request fails with a retryable error and will be retried after
a delay" — alongside `assistant`, `result` and `compact_boundary`. Neither is a statusline
field at this version. Nothing shelved was reopened to establish this; it is this session's
own read.
#### Known limitations
- **A live capture now backs the field list — closed 2026-08-16, the same day.** The
dispatcher wore a tee for one interactive session and came out again; the host was the same
pinned 2.1.233. Seven fires landed, all inside one prompt (one `session_id`, one
`prompt_id`). Every fire satisfied `TAw`'s arithmetic: `total_input_tokens` equalled the
three `current_usage` input fields summed (e.g. 2 + 1122 + 66355 = 67479), and
`total_output_tokens` equalled `current_usage.output_tokens`. Two refusals above were also
witnessed rather than argued: two adjacent fires repeated one call byte-for-byte on the
token block (a summer counts that call twice), and `total_output_tokens` fell 169 → 3
between fires inside the one prompt — a "total" that decreases, with no `/compact`
involved. The host sent 16 top-level keys, a strict subset of what the bundle can assemble:
`vim`, `pr`, `remote`, `agent_type`, `workspace.repo` and `worktree` never appeared (the
session's cwd was not a repo), and `permission_mode` stayed absent — §2's exclusion holds
on the wire, not only at the call site. The capture also found one field the source read
missed: `effort` (`{"level": ...}`), present on every fire; it joins the
deliberately-unmodelled list in the next limitation. Still open: one session on one
machine, `prompt_id`'s rotation across prompts unobserved, and the occupancy level's
decrease after `/compact` unwitnessed (the output half's decrease above is the same
property on the cheaper field).
- **The payload has grown fields this struct still ignores** — measured present at 2.1.233 and
deliberately unmodelled, because nothing needs them: `exceeds_200k_tokens`, `fast_mode`,
`output_style`, `thinking`, `vim`, `pr`, `remote`, `agent_type`, `workspace.added_dirs`,
`workspace.repo`, an expanded `worktree`, and three more `cost` fields
(`total_api_duration_ms`, `total_lines_added`, `total_lines_removed`). The live capture
added `effort` to this list — the source read missed it entirely. Unknown fields are
ignored by design (§2), which is why that growth was a non-event.
- **§2's "permission mode is not in the payload" survives, for a changed reason.** `py` CAN
emit `permission_mode`, but the statusline calls it with two arguments, so the field is
`undefined` and never serialized. The exclusion is now a property of the call site rather
than of the schema, and a future version that passes the argument would ship it.
### 7.16c The cache hit ratio — the vendor computes it, so telltale may render it (2026-09-04)
§7.16b is the record of a seam that measured out to nothing. This is its mirror image on the
same payload: a number arrived that telltale is allowed to display, and the reason it is
allowed is the whole content of this section.
**The question was whether any surface telltale already reads passively carries prompt-cache
counts.** Two do, and the difference between them decided where the work went.
#### What was measured, and how
- **The transcript on disk carries the raw counts and not the ratio.** A read-only pass over
the owner's corpus on 2026-09-04 — 300 transcripts sampled from 1,508, 65,589 records —
found `message.usage.cache_read_input_tokens` and
`message.usage.cache_creation_input_tokens` on 32,416 records, written by six CLI builds
from 2.1.209 to 2.1.258, plus the same two keys nested under `message.usage.iterations[]`
and under `toolUseResult.usage`. No `prompt_cache` and no `hit_ratio` key occurs at any
depth. `internal/adapter/claudecode` has parsed both counts since v1 and folds them into
`contextIn()`.
- **The statusline payload carries the ratio itself, computed by the vendor.** Source read of
the shipped executable at **2.1.260** (the same field names are present in the 2.1.251
build on the same machine, and 2.1.251 is the floor the vendor documents). One function
assembles the block and returns an empty object — so the whole key is absent — while
`requests` is 0:
```js
function WZt(w=Date.now()){let x=JUt(void 0,w);
if(x.requests===0||x.lastRequest===null)return{};
return{prompt_cache:{warm:x.warm,caching_observed:x.cachingObserved,ttl:…,
requests:x.requests,misses:x.misses,expected_rebuilds:x.expectedRebuilds,
hit_ratio:x.hitRatio,cache_write_tokens:x.cacheWriteTokens,…}}}
```
and the ratio comes from the accumulator's `summary()`:
```js
let n=this.cacheReadTokens+this.cacheCreationTokens+this.inputTokens;
… hitRatio: n>0 ? this.cacheReadTokens/n : null
```
#### The ruling
**A ratio derived from the transcript is refused; the reported ratio is rendered.** They are
the same formula and they are not the same claim. Dividing the adapter's own counts would be
arithmetic telltale invented, which is the ADR-001 §4a.1 refusal, and it would additionally
be wrong in scope: the adapter reads a head+tail window of a file that reaches 7.7 MB, so a
"session" ratio built there is a ratio over whatever the windows happened to cover, silently.
The vendor's figure is taken over every main-conversation request of the session. telltale
reads that quotient and multiplies by 100 — the unit conversion §2.1 already permits for
Antigravity's `remaining_fraction` — and computes nothing else.
So the adapter gains nothing and stays exactly as it was. That is the answer to "add a cache
ratio to the Claude adapter": the honest place for it was the other Claude seam.
#### What was built
- **`internal/claude`**: `StatuslineInput` gains `PromptCache`, modelling four of the
block's fourteen fields. `HitRatio` renders. `CachingObserved` gates it. `Warm` and
`Requests` are parsed and rendered by nothing, each for a reason stated on the field.
The struct is the allowlist for the block, the `internal/cursorhook` technique reused.
- **`internal/statusline`**: one segment, `cache 91%`, placed directly after `ctx` because
the two answer one question together — how full the window is, and how much of what fills
it came from cache.
- **No threshold colour, and this is a rule rather than a preference.** Every other
percentage on that line is a consumption, so `pct()` paints high values red. A hit ratio
inverts that, and nothing here has measured the ratio at which a cache becomes bad, so
inventing an inverted scale would be inventing a judgment. The value renders unpainted and
the word `cache` carries the distinction, which is what §9's accessibility rule asks of
every distinction anyway.
- **Three absences hide the segment and none renders as zero**: no block (a CLI older than
2.1.251, or a session before its first API response), a null `hit_ratio` (the vendor's own
`n>0` guard), and `caching_observed` false (the provider or gateway reports no cache tokens
at all). The last is the sharp one: it is an unread field, not a 0% hit rate. A ratio of 0
**with** caching observed is a reading and renders `cache 0%`.
- **Pinned on both halves**: the parse tests cover the measured shape, both flavours of an
absent block, and the `caching_observed: false` case; the render tests pin the line, prove
the segment reads the reported ratio rather than the per-call counts sitting beside it in
the same fixture (they would yield 75%, not 91%), and assert the value carries no ANSI. CI
asserts on the **built binary** that the 2.1.233 fixture grows no cache segment and that
the 2.1.260 fixture renders `cache 91%` and nothing else from the block.
#### Known limitations
- **This is a source read with no live capture behind it**, which is weaker than §7.16b
ended up being. The formula, the field names and the absence semantics come from the
2.1.260 executable and the vendor's documentation page agreeing with each other; no
interactive session was teed to watch a real `prompt_cache` block arrive on the wire, and
no cold-cache or `caching_observed: false` payload has been observed. A capture would
close all three at once and is the obvious next measurement.
- **The docs page's field table is not the whole block.** The same source read found two
fields it does not list: `last_miss_cause` (an object carrying `causes`, `tools_added`,
`tools_removed`, `system_char_delta`) and `miss_causes`. That is §7.16b's `prompt_id`
finding repeating, and it is recorded rather than modelled — absence of need is a result.
- **Eight documented siblings are deliberately unmodelled**: `ttl`, `expires_at`, `misses`,
`expected_rebuilds`, `cache_write_tokens`, `miss_recache_tokens`, `last_miss_at`,
`recache_tokens_if_cold`. `warm` is the one to reopen first — a cold prefix makes a healthy
ratio unusable, and a cold mark on the segment is a second claim that needs its own ruling.
- **The statistics cover the main conversation only.** The vendor excludes subagent requests,
so a fan-out's cache behaviour is not in this number and the segment does not imply it is.
- **§7.16b's relay refusal is untouched.** Nothing here writes `usage/claude.json`; a ratio
is not a count, and the two CI relay assertions still stand beside the new ones.
### 7.17 `u`: the fleet usage view — two claims, and never one
Added 2026-08-09; amended the same day by the reading pass below (the models census, the
title's rule weight, and the age escalation the 19-hour incident forced).
§7.15 gave the header a block per vendor and §7.16 gave it a spend line,
and between them they filled the one or two lines the header has. §7.10's last open
limitation said the quiet part: *"the account quota block is sourced from one session. A
second quota-bearing vendor needs a per-vendor block."* It has one now — but not in the
header, because the header is answering a different question.
**Glance and read are different jobs.** The header answers *am I about to run out?* in the
time it takes to look up from an editor, and its whole design is a shed cascade that spends
decoration to keep facts on one line (§7.15). It cannot also answer *what can telltale
actually say about each of my five vendors, and where it says nothing, why?* — that answer
is a paragraph per vendor, and a paragraph per vendor is a body, not a header. So `u` opens
a third body over the row area, on the detail pane's precedent (§7.11): it replaces the
grid rather than floating over it, because a panel covering the thing being monitored is a
monitor you have to move to read.
**The header was left unchanged, and that was wrong for this one body** — see *the header
stops repeating the page* below, which reverses it. Over the GRID the two surfaces render
the same readings from the same assembly and the duplication is the point, one for glancing
at and one for reading. Over this body there is nothing left to glance at: the reader is
already looking at the read surface.
#### The organizing insight, and it came from the measurement
"Usage" is **two different claims**, and the view may never blur them.
| | **Quota** | **Spend** |
|---|---|---|
| what it is | a reading against a limit the vendor published | a count of tokens with no denominator anywhere |
| sources today | Codex (its own store, scan-fresh); Claude and agy (statusline relay) | agy (summed from the conversations this scan read, §3.8) |
| may render | gauge, percentage, reset countdown, severity hue | a verb, the counts, and the accumulation window |
| may never render | — | a gauge, a percentage, a countdown, a bar, or the sum without its window |
The right-hand column is not a style preference. There is no ceiling anywhere in a token
count — no vendor here publishes an account limit a passive reader can see (§3.8, §3.9,
§3.9a) — so a bar or a percentage would **invent** one, which is the same class of error
as filling a `CapNone` field with a plausible guess.
`TestUsageSpendBorrowsNoneOfQuotasVocabulary` pins it, and since the header's spend line
was retired (§7.16's amendment) that test is the *whole* of the guarantee rather than the
second half of a pair. It matters most here anyway: this surface puts the two measurements
four lines apart under one vendor name instead of on separate header rows, and
**proximity is what makes this the riskier render of the two**.
#### The layout
One block per vendor, in **fixed fleet order** — claude, codex, gemini, agy, cursor, grok —
never sorted by usage. Position is the navigation: a vendor moving must mean a vendor was added
or removed, not that another vendor's percentage crossed it. `fleetOrder` is now one
variable, walked by both the header's per-vendor counts and this view, so a vendor sits in
the same place on both surfaces.
Each block is a heading that states the **quota seam** — where the reading came from, or
why there is none — and then one line per fact under it, hanging off the detail pane's
label column so the two bodies read as one product. The generated render (`usage-fleet`
golden, and the shape it took after the 2026-08-09 amendment below):
```
telltale │ 7 sessions │ codex 1 gemini 1 agy 3 cursor 1 grok 1
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
fleet usage quota is a reading against a limit; spend is a count with none ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
claude quota relayed by the statusline · 2h ago
5h ████████──────────── 42% ↻ 2h13m
7d █▏────────────────── 6% ↻ 5d00h
codex quota read from its own store, this scan
models gpt-5.1-codex
7d ███████████████───── 79% ↻ 22h48m
gemini no quota reaches disk anywhere telltale can read
models gemini-3-pro
agy quota relayed by the statusline
models Gemini 3.6 Flash (High)
gemini-weekly ███████▎──────────── 38% ↻ 3h00m
spent uncached in 1.2M · out 13.1k · summed across 2 sessions on disk, this scan
cursor no quota anywhere · its store holds experiment values, not usage
models composer-2.5
grok no quota anywhere · no window, no ordinal, no reset time on its disk
models grok-4.5
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
esc close ↑/↓ scroll
```
Five things in that frame are doing work:
- **The gauge is 20 cells, not the header's 8.** That is the concrete reason this view
exists rather than a wider header: one window per line buys the bar room to be *read*
rather than merely glanced at. Fill resolution goes from 1.6% to 0.66%. It sheds to 12
cells at the compact tier and disappears below 80 columns, on §7.2's own breakpoints
rather than on new ones — and it is allowed to shed only because the number beside it
stays.
- **The heading names the source, and §7.15 makes that load-bearing.** A transcript-sourced
block is re-measured every scan; a relayed one is exactly as old as the last statusline
render, and only the first may ever carry a burn forecast. A reader deciding how much to
trust a percentage needs to know which they are looking at. The relayed reading's **age
survives every dress level** — the phrase around it gives way instead — because shedding
the age would re-present a stale number as fresh.
- **`agy` carries both kinds of claim, and they arrive from two different seams.** Its
percentage is relayed by the statusline; its token counts are summed out of the
conversations this scan read. The heading always speaks about quota and never about
spend, deliberately: the spend line explains itself in its own vocabulary (a verb and a
window), while an absent or relayed reading explains nothing at all unless something says
it out loud.
- **`gemini` is one line, and it is a line rather than a row of dashes.**
- **`cursor` and `grok` have a sentence each and no numbers, and the sentences differ**
because the measurements behind them do.
- **The header above it is identity only.** It carried the quota strip when this view
shipped; the amendment below took it off, because every figure on it is restated
underneath with room.
#### Three kinds of nothing, and the one that had to collapse
§4a.1's rule is that the kinds of absence stay distinct. On this surface there are three,
and they are the whole reason the absence line carries a reason rather than an em dash:
| | Vendors | What it renders | Why that wording |
|---|---|---|---|
| **structurally absent** | gemini, cursor, grok | `no quota reaches disk anywhere telltale can read` / `no quota anywhere · its store holds experiment values, not usage` / `no quota anywhere · no window, no ordinal, no reset time on its disk` | There is no seam to fire. Cursor's verdict was re-measured 2026-08-08 and came back harder: the only account figures on its disk are Statsig experiment values stamped `is_user_in_experiment:false`, never consumption (§7.16). grok's was measured the same way on 2026-08-09: a rate/limit/quota sweep of the whole store matched tool-configuration keys and nothing else (§3.9a). Naming an action here would send someone to enable a thing that does not exist. **Each vendor gets its own sentence rather than sharing one** — the verdicts are the same shape and different measurements, and lending grok Cursor's wording would claim something about grok's disk that nobody looked for there. |
| **seam exists, never seen** | claude, agy | `no quota relayed yet · the telltale statusline writes it` | This is the one absence a user can act on, so it **names the statusline**. The reading turns up as soon as the gauge runs in that vendor. An absence with an action behind it that does not say the action is just a shrug. |
| **aged out** | any relayed vendor | *renders as never-seen* | `quotacache`'s reader drops a window whose reset has passed and any entry over 24h old before the HUD ever sees it (§7.15's self-expiry). Telling the two apart would mean **holding numbers §7.15 calls not stale but FALSE** so this view could display them. Losing one distinction is the cheaper of those two trades — and it is recorded as a limitation below rather than left to be discovered. |
Codex is in none of the three: its quota comes from its own store, so an absence there is a
statement about what this scan read (`no quota in the sessions read this scan`) rather than
about a relay that never fired. Borrowing either of the other two sentences for it would
name the wrong seam.
**A vendor with no sessions, no quota and no total does not appear at all.** The view is a
report on what is running, not a checklist of every adapter that was compiled in — and an
absence line is for a vendor that is *here* and silent, not for one that is not here. When
nothing anywhere has anything to say, the body is a sentence rather than five blocks of
dashes, because a table of nothing is a table asserting it measured five things
(`usage-empty` golden):
```
telltale │ 0 sessions
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
fleet usage quota is a reading against a limit; spend is a count with none ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
no vendor on this machine has reported a quota reading or a token count
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
esc close ↑/↓ scroll
```
#### The reading pass (amended 2026-08-09): it worked, and it did not read like a page
The view above shipped the same day it was designed and it was *correct* — every claim in
it is measured, every absence names its seam. Asked whether it was the best the product
could do, the answer was no, and for three separable reasons: it did not say **which
models did the work** (half of the original ask, dropped in the first cut), it had **no
visual nesting** (a body title at the same weight and the same column as its own entries),
and an **old reading looked exactly like a fresh one** apart from four muted characters.
The third one had already cost something real, which is why it is the longest item here.
**The census: which models actually did the work.** Each vendor block now carries a
`models` row naming the model display names this scan saw under that vendor.
```
telltale │ 6 sessions │ claude 4 codex 1 gemini 1
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
fleet usage quota is a reading against a limit; spend is a count with none ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
claude no quota relayed yet · the telltale statusline writes it
models Haiku 4.5, Opus 5, Sonnet 4.5
codex quota read from its own store, this scan
models gpt-5.1-codex
7d ███████████████───── 79% ↻ 22h48m
gemini no quota reaches disk anywhere telltale can read
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
esc close ↑/↓ scroll
```
It is a **third kind of line**, and it belongs to neither of the two claims this view is
built around. A census has no limit and no total — nothing to compare it against — so it
borrows neither vocabulary and simply lists what was there. What governs it is the same
rule everything else here obeys:
| Rule | Why it is the honest-gauge rule again |
|---|---|
| **Only this snapshot** | Never a remembered list and never the vendor's catalogue. A name surviving from a previous scan would be a claim about the past presented as the present — the defect the relay's age exists to prevent, arriving through a different door. Four Claude sessions above; take them away and the row goes, even though the vendor still has a quota reading. |
| **The grid's own normalization** | Through `DisplayModel`, so `claude-opus-5` reads `Opus 5` here and in the MODEL column, and a reader never has to work out that two spellings are one model. |
| **Deduped, sorted** | Four sessions, three names. Alphabetical rather than by recency or by session count: ranked ordering reshuffles every time a turn lands, and §7.1 rule 4 budgets the movement on this screen at one cell. |
| **Absent renders absent** | A vendor whose sessions carry no model gets **no row** — `gemini` above has a session and no census. Not an em dash: the dash would claim telltale looked at a model and could not name it, when what happened is that the adapter sources no model at all. |
| **Overflow is announced** | `+3 more`, never a clipped list. The cap is the *room* rather than a magic number, and names are dropped whole rather than cut, with the marker's own width reserved before the last name is accepted. At the 60-column floor: ` models Haiku 4.5, Opus 5 +3 more`. An ellipsis there would leave the reader unable to tell whether one name went missing or nine. |
**It is its own row rather than part of the heading**, and that was the one real layout
fork in this pass. The obvious placement is beside the vendor name —
`claude · Opus 5, Sonnet 5 quota relayed by the statusline · 2h ago` — and it is wrong
for two reasons that point the same way. First, the heading has one job and §7.17 already
made it load-bearing: it states the **quota seam**, and the relayed reading's age may
never shed from it. A second variable-length vendor-supplied fact on that line puts a
census in competition with the one thing on the surface that is not allowed to give way,
and at 60 columns the census wins by being at the front. Second, the census is a *session*
fact aggregated per vendor while the heading speaks about the *account* — putting them on
one line blurs exactly the distinction the block's shape exists to draw. In the label
column it instead joins `5h`, `7d` and `spent` as a fourth labelled fact about one vendor,
which is what the column is for.
**Within the block it comes first**, before the quota windows and before spend: it names
who did the work and the rows under it say what that work cost, subject before predicate.
It does not tear the heading from its evidence, because the heading states a provenance
("relayed by the statusline") rather than a number.
#### The title carries the room's second rule weight
`fleet usage` and ` claude` started in the same column at the same weight, so the body's
**title read as a peer of one of its entries** — §9.23's finding one surface over, where a
turn page's outline whispered while its entries shouted. The HUD had one rule weight and
was asking it to be both the frame's edge and a gauge's empty track.
`RuleHeavy` is `━` and `=`, **council's own pair** rather than a second one: §7.1 principle
5 is that these are one product, and a second heavy-rule character would be a second
alphabet. `=` is the one unclaimed mark left in the HUD's reduced set — `-` is the light
rule and the gauge track and the fact separator and a spinner frame, `#` the gauge fill,
`|` the separator, `>` the ellipsis, `~` the reset, `!` the warning, `]` the cursor, `Y`
the fan-out, `*`/`o`/`.` the state dots, `_` the caret — and
`TestTheHeavyRuleHasAnUnclaimedASCIIPartner` enumerates that list so the next glyph cannot
be added without meeting it.
**The rule goes ON the title, not under it**, and that is council's ruling rather than a
preference: §9.11 spent a whole item removing a heading followed by a horizontal rule, on
the finding that such a rule says nothing the heading had not, and ruled that a heading
carries its own. Here it also avoids a specific defect — the frame's own light full-bleed
rule sits one row above, and a *heavier* line three rows inside a lighter outline is
§9.26's hierarchy argument inverted. It costs **zero rows**, which is what makes it
affordable on a body with a line budget, and it yields to the legend rather than the other
way round: the note is fitted first and the rule takes what is left, because a legend is a
statement and a rule is chrome. At the 60-column floor that leaves ten cells; below
`usageRuleMin` the line simply has no rule.
**The vendor headings deliberately get no rule, not even the light one**, and
`TestOnlyTheUsageTitleDrawsTheHeavyRule` asserts the weight as a *count* on the rendered
frame — one line, one run, and zero on every other body. §9.26's argument is that a second
weight is worth exactly what it is scarce; a rule on every block would spend it five times
a screen to restate what an indent, a blank row and the identity hue already say.
#### Air and alignment: what was already right, and what is now pinned
The columns were already shared — one `usageLabel` cell and one `usageGap` for every fact
row, so `claude`'s `5h` gauge starts where `codex`'s `7d` gauge and `agy`'s
`gemini-weekly` gauge start — and the blocks were already one blank row apart. This pass
**pinned that rather than built it**, which is the honest description:
`TestUsageFactsShareOneColumnGridAcrossVendors` now walks the rendered body at four widths
and fails if the label column, the gauge, the percentage or the reset countdown lands in
two different columns across vendors, and if any label runs into the value column. Before
it, the models row could have been added with a layout of its own and nothing would have
noticed.
Two things were considered and **declined**. A second blank between blocks: §9.11's
threshold is that every deliberate blank this product draws is exactly one row, and a
one-row gap is a boundary placed between two things meant to be kept together — two rows
is nothing the design asked for. And a light rule under each vendor heading, for the
scarcity reason above; air is the boundary strength this body can afford, and it already
has it.
> **The first of those was reversed on 2026-08-09** — narrowly, and only where the rows are
> otherwise going to waste. See *the page stops trailing off*. The rule under each vendor
> heading stands declined.
#### An old reading has to look old (the 19-hour incident)
**What happened, 2026-08-09.** A Claude relay entry written nineteen hours earlier
reported 15% of the seven-day window. It rendered at full confidence beside a live gauge
and was read as current. The account was at 44%.
Nothing in it was dishonest. The age was on screen — `· 19h ago`, exactly as §7.15
requires — and every part of the state is one the product genuinely reaches: the five-hour
window was gone because `quotacache` drops a window whose reset has passed, the entry
survived because it was inside the 24h ceiling, and the reset it reported was still four
days out so nothing upstream had any reason to touch it. **The age was present and it was
not loud.** A muted four-character suffix is the same weight as every other piece of
chrome on that line.
```
telltale │ 7 sessions │ codex 1 gemini 1 agy 3 cursor 1 grok 1
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
fleet usage quota is a reading against a limit; spend is a count with none ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
claude quota relayed by the statusline · ⚠ 19h ago · older than the fleet's shortest quota window
7d ██▉───────────────── 15% ↻ 4d07h
codex quota read from its own store, this scan
models gpt-5.1-codex
7d ███████████████───── 79% ↻ 22h48m
gemini no quota reaches disk anywhere telltale can read
models gemini-3-pro
agy quota relayed by the statusline
models Gemini 3.6 Flash (High)
gemini-weekly ███████▎──────────── 38% ↻ 3h00m
spent uncached in 1.2M · out 13.1k · summed across 2 sessions on disk, this scan
cursor no quota anywhere · its store holds experiment values, not usage
models composer-2.5
grok no quota anywhere · no window, no ordinal, no reset time on its disk
models grok-4.5
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
esc close ↑/↓ scroll
```
**Five hours, and the number is argued rather than tuned.** `quotaAgeWarn` is the shortest
quota window telltale has measured anywhere in the fleet — Claude's `five_hour` (§3.1) —
and therefore the shortest span over which a vendor is known to reset a limit *wholesale*.
A relayed reading older than that has outlived the fastest-moving quota this product knows
about: whatever window it reports, an entire window of the shortest kind could have opened
and closed since it was taken, so the reader may no longer assume the number describes
now. Below five hours the reading is old and still bounded by something; above it there is
no window short enough to bound it. It pairs with `quotaAgeShown` (5m, where the age starts
rendering at all) as the two boundaries of a relayed reading's life on screen.
Three thresholds it deliberately is **not**:
- **Not the per-block shortest window**, which was the first candidate. `model.QuotaWindow`
carries no duration — a label, a percentage and a reset time — so a per-block rule would
have to infer a length by parsing `"5h"` or `"seven_day"`, and a threshold derived from a
display string is the class of guess §4a.1 rejects. It would also have failed on the very
reading that prompted this: the expired 5h window was already gone, so the surviving
block reported only `7d` and a per-block rule would have stayed silent at nineteen hours.
- **Not a second reason to drop a reading.** `quotacache` owns expiry (§7.15) and this view
renders everything it is handed, changing only how loudly. Dropping earlier than the
reader does would hide a measurement telltale holds.
- **Not a freshness gauge.** There is no denominator for "how fresh", which is the spend
line's argument in a different costume.
**Word first, glyph second, hue third**, in that order, because §7.1 rule 2 is that colour
never carries a distinction alone. Past the threshold the heading's dress gains the warning
glyph beside the age *and* the reason in words — `· ⚠ 19h ago · older than the fleet's
shortest quota window` — and the whole statement renders in `SevWarn`, which §7.5 already
defines as the token for warning notices and which the footer's own `⚠ last scan 1m ago`
already uses. In the reduced set that line reads
`claude quota relayed by the statusline - ! 19h ago - older than the fleet's shortest quota window`.
It says "the fleet's" rather than "its" on purpose: five hours is the shortest window in
the fleet, not necessarily in this block, and an agy weekly reading six hours old has not
outlived its own window. Trading one over-confident render for one over-stated warning is
not a trade.
The shed cascade is unchanged in grammar — the age is the fact and never sheds, the
sentence explaining it is decoration and does, so the barest level is `⚠ 19h ago`. What a
narrow terminal loses is the argument, not the alarm. And the **reading itself is
untouched**: the 15% keeps its own severity hue, because that is a statement about the
account and this is a statement about the measurement.
**There is exactly one step.** §9.26's lesson is that a second level is worth what it is
scarce, and the only boundary above this one that is not invented is `quotacache`'s 24h
drop — which is a disappearance rather than a louder warning.
#### Amended 2026-08-09: the header escalates too
The pass above deliberately left the header alone, and recorded that as unfinished work
rather than as a boundary. It was the wrong half to fix first: the reading that was acted
on was read off the **glance** line, so the product ended up loud on the surface a reader
opens on purpose and quiet on the surface they merely look at. The header now escalates on
the same threshold, in the same order, off the same constants — `quotaAgeWarn` and the
reason string are read from §7.17's declarations rather than restated, because two copies
of `5 * time.Hour` is how the two surfaces would come to disagree about one reading.
Past the threshold the header's age suffix becomes `· ⚠ stale 19h ago`, in `SevWarn`, with
`· older than the fleet's shortest quota window` beside it at the most dressed level. The
reduced set renders `- ! stale 19h ago`. The reading itself is untouched — 15% keeps its
own severity hue, for §7.17's reason: that is a statement about the account and this is a
statement about the measurement.
**The header keeps a WORD where the view's barest level does not**, and that is the one
place the two surfaces deliberately differ. §7.17's shed bottoms out at `⚠ 19h ago`,
because by then the reader has opened a body on purpose and the sentence above it is still
on screen. The header has no sentence anywhere and is read at a glance, so the level that
survives every shed has to state the verdict without the reader knowing what `⚠` is
supposed to cost them. `stale` is that word: it names a state rather than restating the
duration (`19h old` would be the number a second time), and it is the word this codebase
already spends on a reading that has outlived its currency — `DotStale`, and the footer's
own `⚠ last scan 1m ago`.
**What sheds is the argument and only the argument.** `ageReason` joins the dress ladder as
a sixth level above the existing five and is the first thing dropped, ahead of forecasts:
it is by a wide margin the longest clause on the line, and it is the only part of the
escalation whose absence costs a reader something they cannot otherwise see. Everything
below it — glyph, word, age — survives to the barest level, so a narrow terminal loses why
the reading is distrusted, never that it is. Same grammar as the view, one surface over.
#### Everything else is inherited, not invented
- **`u` opens and closes it; `esc` closes it.** The esc chain gains a step and now reads
usage → detail → help → clear the query → quit. One body at a time: opening any of the
three closes the others, enforced in `Update` rather than in `Render`, because a pane
that appears only because it won an ordering is a pane nobody can predict. Find mode
closes it too — but only in that direction, since once the mode has the keyboard `u` is
a letter (§7.8).
- **`u usage` joins the footer hints** beside `enter detail`, and sheds on the same tier
boundary for the same reason: below 80 columns the footer keeps only the keys nothing
else can teach. It is on the help overlay's keys page at every width.
- **The drift notice renders under this body too** (§7.3), like every other. A warning that
comes and goes depending on which pane is open is one a reader cannot trust to be there.
- **Overflow scrolls** with the help overlay's vocabulary — `↑`/`↓` move the body, bounded
against its own rendered length — not the grid's `+N more`, which counts sessions.
- **The 60-column floor and the height tiers are unchanged**, and nothing here animates
(§7.1 rule 4). At the floor the bars are gone and every fact survives, including the
relayed reading's age and the spend total's window.
- **The view is not narrowed by `v` or by the find query.** Those narrow the *session*
list, and nothing on this surface is a session fact — filtering an account reading by a
session filter would be the per-row quota §7.1 forbids, arriving by the back door.
#### Declined
- **Trend sparklines.** The caches hold one reading per vendor, so a trend line would be
drawn from data that does not exist. The burn forecast (§7.12) is the honest version of
this and it is confined to the one scan-fresh block that can support it.
- **Sorting by usage.** Position is the navigation; see the fleet-order note above.
- **Per-row quota.** An account fact is not a session fact — §7.1's sixth rule, and the
reason this view exists at all.
- **A fabricated fleet total.** Different units (percentages of unrelated windows, and raw
token counts), different accounts, different vendors. Any single number across them would
be arithmetic telltale invented, which is the ADR-001 violation this whole product is
built to refuse.
- **A spend line for grok**, added 2026-08-09 and the closest call in this section. grok
is the only vendor here that writes real money to disk, so it is the one block where a
reader might expect a figure — and it gets none. The only cost on its disk is
`usage.costUsdTicks` on each `turn_completed` record: per-turn, not cumulative,
in an append-only file that reached 818 KB in one session (§3.9a). A tail-window sum is
a **lower bound**, and a lower bound rendered next to the word "spent" is a derived
number wearing a read one's clothes. The last turn's cost is already a labelled Extra in
the detail pane, where its label says which turn it belongs to. Nothing on grok's block
says any of this: the heading speaks about quota and never about spend, and buying one
vendor an exception would cost every other block its meaning.
#### Amendment, 2026-08-09: the owner's ruling on which vendors this speaks for
Ruled by the owner: **Cursor's spend display goes; agy and grok come on.** The retirement
half is §7.16's amendment. This is the half that adds.
**agy's spend line, and why its window reads differently.** The source is the agy
adapter's measured per-conversation token counts — `gen_metadata`'s `#1.#4.#2` (uncached
input) and `#1.#4.#3` (output), guarded by the `thinking + answer == output` identity §3.8
requires and the adapter asserts. Those were already on screen per row as display-only
extras; what is new is `model.Session.Tokens`, the same two numbers as integers, summed
per vendor by the view. The integers exist because a sum of pre-rounded display strings is
a sum of roundings; the adapter sets both in the same branch from the same variables, so a
row and the fleet total cannot disagree about what was counted.
**This is a scan, not a meter, and the wording has to carry that.** Cursor's total was a
file that only ever went up. agy's is a sum over the conversations that are on disk at
this moment, so *deleting a conversation makes it smaller*. §7.16's rule — the sum never
prints without its window — therefore binds against a different window:
- the wording is `summed across 2 sessions on disk, this scan`, and it never says "since
". A "since" is a meter's claim and this is not a meter.
- the shed cascade drops words and never facts: `summed across N sessions on disk, this
scan` → `summed across N sessions on disk` → `across N sessions on disk` → `N sessions
on disk`. The **count** survives every level because it *is* the window, and **"on
disk"** survives every level because it is the difference between this sum and a
monotonic one. "summed" and "this scan" are what give way, in that order.
- the count is of the sessions that **contributed a measured reading**, not of the vendor's
sessions. The fleet fixture's agy has three conversations and the line says two: the
third has not called a model yet, carries no counts, and is in neither the sum nor its
window. Folding it in as a zero would put a session in the denominator that contributed
nothing to the numerator — §4a.1's rule applied to a window instead of to a cell.
- a generation that **failed its self-check** is dropped by the adapter and named in that
row's Diagnostics, which is where a reader finds out a total is over fewer generations
than the conversation ran. That is inherited behaviour, not new, and it is the one place
this surface is quieter than the row it summed: the count says how many sessions, not how
many generations inside them were refused. Recorded as a limitation below.
- the label is **`uncached in`**, not `in`. §3.8 marks the cache-read component's field
number lower-confidence and the adapter refuses to fold it into a rounder total;
labelling the number `in` would quietly promote a partial figure to a whole one, and the
fleet line is the one place a reader could not catch that.
**And it wraps rather than sheds, which nothing else on this surface does.** Every other
line here sheds decoration — a gauge that re-states a number still on screen, a phrase
around an age. This line is facts only: two counts, a label that has to say "uncached",
and a window that may not go. At 60 columns those do not fit on one row and there is
nothing left to spend. The header solved that class of problem by dropping whole vendor
blocks, because a header has a hard one-or-two-line budget; **this is a body and it
scrolls**, so it can pay a second row, and a second row costs a reader nothing next to a
sum that has stopped saying what it summed. The window hangs under the counts in the same
indent and re-runs its own cascade there — a row to itself buys back dress the shared row
could not afford, so the *narrow* render says more about the window than a cramped
single-line one would have. It carries no leading glyph: the mid dot separates facts on a
line everywhere else in this product, and giving it a second job as a continuation mark is
the failure §9.26 is a whole section about. `usage-floor` (60 columns):
```
agy quota relayed by the statusline
models Gemini 3.6 Flash (High)
gemini-weekly 38% ↻ 3h00m
spent uncached in 1.2M · out 13.1k
summed across 2 sessions on disk
```
**grok's block** is the absence table's third structural row, above. It qualifies for a
block on sessions alone, and it would have rendered the un-surveyed fallback
(`no quota telltale can read`) — honest, and a step down from a sentence that names what
was measured. It now says `no quota anywhere · no window, no ordinal, no reset time on its
disk`, from §3.9a's sweep. It has no spend line, for the reason in Declined above.
Added by the reading pass:
- **A freshness gauge.** No denominator for "how fresh", so a bar would invent one — the
spend line's argument, one field over.
- **A second escalation step for the relay's age**, and a per-block threshold parsed out of
a window label. Both above.
- **A light rule under each vendor heading.** The air-and-alignment note above. (A second
blank row between blocks was on this list too, and came off it — see below.)
#### Amended 2026-08-09: the header stops repeating the page
**The defect, from driving the real thing.** With the usage body open the header still drew
the full quota strip — two cramped rows of
`ag 3p-5h 0% ↻2h10m 3p-weekly 0% … cc 5h 40% ↻2h17m 7d 62% … cx 7d 37% … cursor spent in
47.8k out 66 …`. Every fact on them is stated properly in the blocks four rows below, with
a label column, a 20-cell gauge and the provenance sentence the strip has no room for. It
was the same measurement twice, once cramped and once legible, with the cramped copy on
top, and it was the ugliest thing on the screen.
**Over the `u` body the header collapses to identity and session counts** —
`telltale │ 144 of 1380 sessions │ claude 961 codex 299 …` — and the quota strip does not
render. The grid keeps today's header exactly. Two rows come back to the body.
This **reverses #163**, which ruled the header untouched when this view landed. #163 was
right for the grid and wrong for this page, and the principle worth keeping is the one that
distinguishes them:
> A **glance** surface may not repeat the **read** surface it sits above.
The glance line earns its rows by being the only statement of a fact. Over the grid it is:
account quota appears nowhere else, so the strip buys a fact that is otherwise off screen.
Over the page built to state those same facts at length, it buys nothing and costs two of
the rows that page is short of. Identity survives the collapse for exactly that test — the
session census is not restated anywhere below, because the usage blocks speak about
accounts and never about sessions.
One consequence worth naming: at the 60-column floor the recovered row is enough for
`grok`'s census row to fit, which it did not before. The fix pays for itself in content on
the narrowest terminal telltale renders on.
#### Amended 2026-08-09: the page stops trailing off
**The defect, same session.** On a tall terminal the content stopped around 40% of the way
down and the rest was blank. Nothing was wrong with any line; the page simply had no bottom.
A surface built to be *read* trailed off like a truncated file.
Council solved this exact problem twice (§9.23's contiguous rails, §9.11's boundary-strength
grammar) and neither answer ports: the HUD has no rails, and importing them would give this
one body a vocabulary no other surface in the product speaks. The fix is in the HUD's own
language, and it is two moves that only work together.
**1. The closing rule hugs the content.** The frame's bottom rule was already being drawn —
sixty rows below the last vendor block, where it reads as the terminal's edge rather than as
the page's. Moved to where the content stops, the same rule makes the body a bounded region
with a visible bottom edge, and the leftover rows fall *outside* it: unused terminal, not
unfinished page. This is §9.11's boundary-strength grammar with the weight the frame already
owns, moved — no new glyph, no new hue, no second rule. The grid is untouched, because its
body is a list and a list that ends early has ended; a hard edge under a row area still open
for more rows would claim something false.
**2. The blocks breathe into two-row gaps.** This is the reversal of the air-and-alignment
ruling above, and it is narrow. That ruling was made under a **line budget**, on the
assumption that a row spent on air is a row taken from a fact. On a tall terminal the
assumption is false — the rows are there, unspent — and the real choice is air versus void.
So the second row is never a constant and never a fiat: it appears only when the widened
gaps consume fewer rows than the page would otherwise leave blank at the bottom, which makes
it a **redistribution of air the page already has**. Short terminal, tight page, exactly as
before. The gap before the *first* block never grows, so the page still starts where the
page starts — the §7.x anchor rulings hold and this is not vertical centring.
**It stops at two, and the cap is the argument.** Distributing all the surplus — justifying
the blocks down to the closing rule — was the obvious alternative and is declined twice
over. It makes gap height a function of terminal height and block count, so the distance
between two vendors would encode nothing while looking like it encoded something. And it
puts every gap on a variable, so one grok session appearing reshuffles the vertical position
of every block below it, against §7.1 rule 4's one-cell churn budget. Air that says "these
are separate things" has to be a constant to say it.
Also declined: **inventing content to fill the space** (a fleet total is §7.17's own
rejected list; anything else would be a number nobody measured), and **centring the block
vertically**.
The `usage-tall` golden pins both halves at 52 rows:
```
telltale │ 7 sessions │ codex 1 gemini 1 agy 3 cursor 1 grok 1
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
fleet usage quota is a reading against a limit; spend is a count with none ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
claude quota relayed by the statusline · 2h ago
5h ████████──────────── 42% ↻ 2h13m
7d █▏────────────────── 6% ↻ 5d00h
codex quota read from its own store, this scan
models gpt-5.1-codex
7d ███████████████───── 79% ↻ 22h48m
gemini no quota reaches disk anywhere telltale can read
models gemini-3-pro
agy quota relayed by the statusline
models Gemini 3.6 Flash (High)
gemini-weekly ███████▎──────────── 38% ↻ 3h00m
spent uncached in 1.2M · out 13.1k · summed across 2 sessions on disk, this scan
cursor no quota anywhere · its store holds experiment values, not usage
models composer-2.5
grok no quota anywhere · no window, no ordinal, no reset time on its disk
models grok-4.5
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
esc close ↑/↓ scroll
```
#### Per-vendor hues: ratified, and exactly as far as council's went
Answered 2026-08-09 (San), and it is a **yes**. The question §7.17 left open — whether
council's seat-hue exception (§9.28) extends to this surface's vendor names — turned on
whether the HUD has the concept the exception was granted for. It does, on this body and
nowhere else: **six vendor blocks stack in one column**, each a heading with a paragraph
under it, so position answers nothing about which vendor a reader is looking at. That is
the exact condition §9.28 named. The grid does not qualify and does not get it — a row's
vendor is already answered by a two-letter tag in a fixed column.
It arrives on the terms the open question itself set, and every one of them is council's
rather than a new argument:
- **Names only.** The vendor NAME in each usage-block heading, and nothing else. Not the
seam sentence beside it (chrome, and past `quotaAgeWarn` a warning that must keep
`SevWarn`), not the gauges, percentages or countdowns (severity owns those), not the
spend line, not the models census (theme's identity hue, the same token the grid's
`MODEL` column spends), not the grid rows, and not the header's quota block — the glance
surface names its vendors by tag in a fixed order, so the question the hue answers is not
the one it is asking. `TestTheVendorHueIsSpentOnlyOnTheVendorName` walks the rendered
frame and fails if a hue reaches the header, a fact row or the footer.
- **The assignments are council's, matched by literal value** — claude `5`, codex `6`, agy
`4`, cursor `12`, grok `14`, gemini falling back to the identity hue. A reader who learned
in the room that magenta is Claude must not meet a second colour for Claude one keypress
away. `TestVendorHuesMatchCouncilsSeats` writes council's numbers out and fails when one
copy moves without the other — the shape `TestStripTagsMatchTheHUDSpelling` already uses
for the two-letter tags, and for the same reason: the seam between the two surfaces is the
normalized session model and `internal/theme`'s numbers, and reaching across it for a
rendering detail is the coupling that seam exists to prevent.
- **The map does NOT go into `internal/theme`**, and the stdlib rule is not why. These are
plain strings and would compile there. The reason is theme's own contract — one hue, one
meaning, across every surface that imports it — and `internal/statusline` has no vendor
blocks. Two packages holding the same map is the honest cost of keeping that contract
intact, and the parity test is what makes the cost bounded.
- **4-bit indices, severity and chrome off limits.** `1`/`2`/`3` and their bright twins
`9`/`10`/`11` are the ramp this very surface draws its percentages in — a vendor heading
wearing red would read as an account in trouble — and `0`/`7`/`8`/`15` are the gauge track
and the terminal's own fore/background. `TestNoVendorHueIsASeverityOrChrome` fences it.
- **`PlainStyles`-identity by construction, and zero golden churn.** One `retint` helper
returns the base style untouched when `Plain` is set, so no golden on this surface can see
the feature exist; **any golden diff on this change is a bug**, and none was produced.
`NO_COLOR` needs nothing new — `colorprofile` downsamples inside Bubble Tea exactly as
§7.5 describes.
The **honest weakness is council's too**: `4`/`12` and `6`/`14` are two pairs of one hue at
two intensities, and some schemes render each pair close. Council carries the distinction on
the two-letter tags; here the thing being tinted is the vendor's full name, spelled out at
the head of its own block, so this surface has more carrying it than the room does.
Declined inside the ratification, and the list is closed:
- **The hue on the grid rows and the header.** Position and the two-letter tag already
answer "which vendor" on both, so it would be the circus row §9.28 refused for council's
column headers, spent on a question the layout had already settled.
- **A hue on the vendor tag as well as the name**, wherever the two appear together — double
the ink for a distinction the name already carries. Council's own declined list.
- **Gemini given a hue of its own.** The legal set has one index left (`13`), and spending
it to complete a set is a palette entry with no argument behind it. Gemini takes the
identity-hue fallback and looks exactly as it always did.
#### Known limitations
- **An aged-out relay reading is indistinguishable from one that never arrived**, by the
trade in the table above. Both render as "no quota relayed yet".
- **The header and this view still say some of the same things twice — just never at the
same moment.** The 2026-08-09 amendment above removed the on-screen overlap by collapsing
the header while this body is open, so no reading appears in two places in one frame any
more. What survives is the *maintenance* half of the limitation: two renderers still speak
for the same measurement, and a future change to one has to be made in both.
`quotaVendors` being shared by both surfaces is what keeps them from disagreeing about
*which* source speaks for a vendor; nothing yet keeps them from diverging in tone.
- **The void fix is decided per frame, so a resize can change the gaps.** `usageAir` widens
the blocks apart only when the taller layout still fits, which means dragging a terminal
across the threshold moves every gap by a row at once. It is a resize, not a tick — §7.1
rule 4 budgets *frame-to-frame* churn on a still screen and this cannot fire on one — but
it is a visible jump at the boundary rather than a smooth one, and smoothing it would mean
the variable-height gaps the amendment declines.
- **The absence sentences are per-vendor literals**, so a seventh vendor arrives with the
fallback wording (`no quota telltale can read`) until someone measures its seam and gives
it a sentence. That is the honest default — it claims nothing about a seam nobody has
looked at — but it is a step down from the five that name theirs. **`grok` was that case
live** for one day: it landed as the fleet's sixth vendor (#183), flowed into the block
layout and the shared column grid with no change to either, and took the fallback
sentence because nobody had measured what its store says about an account. §3.9a's sweep
supplied one on 2026-08-09 and it now names its own seam; the fallback is back to being
a path nothing currently takes.
- **The spend line's count is of sessions, not of generations.** A conversation whose read
dropped some generations for failing the `thinking + answer == output` self-check still
counts as one contributing session, and its partial sum is in the total. The drop is
named in that row's Diagnostics and nowhere on this surface, so a reader looking only at
the fleet line cannot tell a whole conversation from a partly-refused one. The
alternative — excluding the whole conversation — would discard generations that passed
their own check and undercount by an amount nothing on screen could name, which is the
worse of the two silences.
- **The spend line's window shrinks silently when a conversation is deleted.** "on disk"
is the only thing saying so. There is no honest alternative from a passive read: the
scan cannot know what was on disk yesterday without keeping its own history, and a cache
of previous scans would be telltale asserting a past it did not measure at the time.
- **One over-age block costs the whole header line its gauges.** The escalation is decided
per block, but the dress cascade is decided per LINE — one level has to fit every block
on it — so the eight columns `⚠ stale` adds to a single vendor can push the line down a
level that every vendor pays for. That is exactly what the `usage-stale-relay` frame
above shows: at 120 columns with three vendors the header was already flush against its
budget, so escalating Claude's reading dropped the bars from agy's and Codex's fresh
blocks too. The trade is deliberate and it is the cascade's existing grammar rather than
a new rule — the percentage beside each bar carries the reading, and a quietly stale
number is a worse failure than a missing bar — but it does mean the frame that
demonstrates the fix is also the frame that pays the most for it. A per-block dress
would fix it and is declined: blocks of different heights and vocabularies on one line
is a grid a reader has to reconstruct, and §7.2's whole argument is that they should
not have to.
- **The models census counts sessions, not turns.** A model that ran once six hours ago and
a model that has been running all morning are the same entry in the row. There is no
per-model activity anywhere the HUD reads, so weighting the list would be an invention —
but a reader who takes the order for importance will be wrong, which is part of why it is
alphabetical rather than ranked.
### 7.18 The scan keeps up: what a poll actually costs, measured (2026-08-09)
The footer's `⚠ last scan Ns ago` (`view.go` `staleAfter = 3s`) was on permanently on the
owner's machine. That notice was **correct** — the scan really was taking longer than three
seconds against a 1 s `pollInterval` — which is the worst version of the problem: a true
signal that fires constantly stops being read, and the one field the HUD has for saying
"do not trust what you are looking at" becomes wallpaper. Raising `staleAfter` was
considered and rejected outright; it hides the measurement instead of fixing what it
measures.
**Measure first.** Profiled against the live corpus on 2026-08-09 (Windows 11, i7-7700K,
1,404 sessions across six vendors — 967 claude, 346 codex, 65 agy, 53 grok, 7 cursor,
1 gemini). Nothing from that corpus is in this repository; only the timings are.
| vendor | refs | discover | read (all refs, gate 8) | per-read p50 / p95 |
|---|---|---|---|---|
| claude | 967 | 29 ms | **2.612 s** | 12.8 ms / 63.7 ms |
| codex | 346 | 27 ms | 138 ms | 1.0 ms / 13.2 ms |
| agy | 65 | 1 ms | 42 ms | 2.5 ms / 14.4 ms |
| grok | 53 | 5 ms | 11 ms | 1.0 ms / 3.3 ms |
| gemini | 1 | 0 ms | 3 ms | 3.2 ms |
| cursor | 7 | 0 ms | 2 ms | 1.0 ms |
Whole `Scan`, warm: **1.84 s / 1.91 s / 3.37 s** over three consecutive runs.
Four of the five things worth suspecting were already fine, and saying so is the point of
recording this:
- **The vendors are already concurrent.** `Scan` fans out one goroutine per adapter and
`readAll` fans out per ref behind a semaphore of 8. Serialization was not the finding.
- **The reads are already bounded.** Head 64 KB + tail 128 KB held at this corpus size —
164 MB read against 693 MB on disk. Cursor's SQLite store was already snapshot-cached on
a two-stat check, and grok's `updates.jsonl` never showed up in the profile at all.
- **The scan is already off the UI goroutine** (`scanCmd`), so a slow scan lagged the
display; it never froze input.
- **The 8-hour idle filter is genuinely expensive** — 1,235 of 1,404 sessions were read and
then hidden — but see the rejected optimization below.
The finding was **JSON parsing, and specifically re-parsing work that had not changed**.
Broken down over claude's 967 files: open 73 ms, stat 14 ms, head read 185 ms, tail read
364 ms, subagent stat pass 321 ms, and **`json.Unmarshal` 2.65 s** — 46,727 records, every
second, of which on a typical tick approximately none had been written since the last one.
A pure `os.Stat` pass over the same 967 files costs 50 ms.
**The fix** is a per-transcript parse cache in `internal/adapter/claudecode`, keyed on the
file's `(size, mtime)` — the same shape `internal/adapter/cursor` already uses for its
store snapshot. The honesty constraints decided its boundaries, not convenience:
- **`last_activity` is not cached.** The §6 Q8 ruling makes it `max(mtime, newest record
timestamp)` and *both* inputs move, so only the record-timestamp half — a pure function
of the bytes the parse read — is stored. The mtime half is re-stat'ed and the max
re-folded on every read. This is the field the display's whole staleness story hangs
off; freezing it would have been the exact self-defeating fix.
- **The sub-agent count is not cached.** It is a function of `now` as much as of the disk
(§7.13's recency horizon): a fan-out expires with no file changing. It runs on every
read, cache hit or not. It was always a stat pass and stays one.
- **Diagnostics and degradation replay verbatim.** A hit reports the same torn records and
the same drift verdict a fresh read would. `drift.Watch.Fold` only reads its watch, which
is what makes replaying a stored one sound.
- **Absence stays absence.** The stat happens *before* the cache lookup, so a deleted
transcript is `ErrSessionGone` and never a replay; and a field the parse did not source
is stored as empty and rebuilt as nil.
- Entries are pruned in `Discover` against the live set, so a HUD left running for a day
does not accumulate one per session that ever existed.
The residual risk is every mtime cache's: a rewrite landing on byte-identical size *and*
identical mtime is invisible. NTFS timestamps are 100 ns, the vendor appends rather than
rewrites, and the cursor adapter already accepts this trade.
**Rejected: skipping the read for sessions the idle filter will hide.** It is the biggest
apparent win on the table — 1,235 of 1,404 rows — and it is a lie. Q8 exists precisely
because NTFS defers mtime while a writer holds the file (observed lags of ~100 s hot,
~20 min closing), so a stat-only recency prefilter would hide the hot sessions the HUD
exists to watch. The scan also cannot know the filter: `a` toggles `ShowAll` and the rows
must already be there. The cache buys the same speed with none of that.
**Result**, `BenchmarkScan` in `internal/hud` over a synthesized 1,400-session corpus
(generated in the test, never committed; five runs each, same machine):
| | before | after |
|---|---|---|
| warm scan (steady state) | 798 ms median (533–931) | **82 ms median (64–174)** |
| cold scan (first, after launch) | 896 ms median (708–1156) | 994 ms median (655–1200) — unchanged within noise |
On the live corpus the warm whole-`Scan` went from 1.84–3.37 s to **181–204 ms**. The cold
scan is untouched by design: the first frame must genuinely read everything, and that is
what the spinner is for.
The benchmark's corpus is Claude-shaped only, which is a deliberate narrowing recorded here
so nobody reads it as a whole-fleet figure: claude was 967 of 1,404 sessions and 2.6 s of
the 2.9 s of read time, and it is the only adapter carrying this cache. The other five
together were under a fifth of the budget. It runs at full scale in CI (~30 MB of temp
files) and drops to 50 sessions under `-short`.
Codex's 138 ms is the next-largest item and is now co-dominant with everything else put
together. It is left alone: the scan is an order of magnitude inside its budget, and the
same cache would need its own correctness argument against its own read path.
### 7.19 `w`: the week page — the slow windows, one line per vendor (2026-08-09)
The owner's question, verbatim: "one view of the weekly usage for these models — it would
help with scoping work." §7.17 already holds every reading that view needs and spends a
block per vendor to say it, with the census, the spend lines and the five-hour windows in
between. A scoping glance wants none of those; it wants the slow pools, one line each. So
`w` opens a LENS over §7.17's data — the same `quotaVendors`/`usageBlocks` assembly, so the
two surfaces cannot disagree about an account — rendered as a tight table.
**Which windows, and why the rule is honest.** The page shows every window the vendor
itself names weekly, plus the vendor's longest. Neither leg infers a duration, which is the
constraint that shaped this section (§4a.1; quotaAgeWarn's ruling that a length parsed out
of "5h" or "seven_day" is a guess wearing a fact's clothes):
- the `-weekly` suffix is vendor vocabulary read verbatim — agy names its buckets
("3p-weekly", "gemini-weekly" observed, §3.8) and quotacache carries the names as ids
unchanged. Reading the vendor's own suffix is reading, not translating.
- the LAST window rides `model.QuotaWindow`'s ordering contract — "display order, shortest
first" — so it is the vendor's longest pool by structure rather than by arithmetic.
Claude's slice ends on `seven_day`, Codex's on `secondary`. No id is parsed for a length,
and each row renders the vendor's own label beside the reading, so the page never states
a duration the vendor did not.
One edge is deliberate: a vendor whose only surviving window is short — Claude relayed
after quotacache dropped an expired `seven_day` — shows that window under its own label.
It is the longest reading telltale holds, and the label says how long it is.
**What is kept off.** SPEND does not appear, and not because it is unimportant: a spend
total's accumulation window is "sessions on disk, this scan" (§7.16) — not a week, not any
calendar span — and rendering it under a page titled "this week" would claim a window the
number does not have. The u page renders spend correctly, one key away. A vendor with no
reading keeps §7.17's absence sentence rather than an em dash, because a dash would say
"no reading now" about vendors that are structurally unreadable (§4a.1's three kinds of
nothing). Sorting by remaining headroom was considered and rejected: readings, absences
and two-window vendors do not order on one axis, and a page that reshuffles when a
percentage moves spends §7.1 rule 4's churn budget to encode nothing the percentages do
not already say.
**The relayed age rides every row.** This page has no vendor headings to carry §7.15's
age, so it rides each row as a suffix, and past quotaAgeWarn it escalates in the header's
own grammar — the word (`stale`), the glyph, then the hue as the second signal. The REASON
sentence does not travel here; §7.17 carries the argument, this page carries the alarm:
```
claude 7d ██▉───────────────── 15% ↻ 4d07h · ⚠ stale 19h ago
```
**The frame, generated by the build** (`internal/hud/testdata/golden/week.txt`; the
week-stale variant is the golden quoted above). Both agy weekly pools under one vendor
name, the second on a continuation row; the four five-hour buckets in the fixture render
nowhere; 3p-weekly's measured 0% draws a full empty track while the three absence vendors
draw sentences — zero and absent, still different states on the page built for a glance:
```
telltale │ 7 sessions │ codex 1 gemini 1 agy 3 cursor 1 grok 1
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
this week each vendor's longest window, and every window it names weekly ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
claude 7d █▏────────────────── 6% ↻ 5d00h · 2h ago
codex 7d ███████████████───── 79% ↻ 22h48m
gemini no quota reaches disk anywhere telltale can read
agy 3p-weekly ──────────────────── 0% ↻ 6d23h
gemini-weekly ███████▎──────────── 38% ↻ 6d23h
cursor no quota anywhere · its store holds experiment values, not usage
grok no quota anywhere · no window, no ordinal, no reset time on its disk
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
esc close ↑/↓ scroll
```
**Known limitations, named:**
- **"This week" is the page's name, not each row's claim.** Codex's `secondary` window is
selected for being the longest, not for being seven days; if the vendor ever reports a
different length there, the label beside the reading says so and the page needs no code
change. The title states the question the page answers, the labels state the facts.
- **A new agy bucket that is weekly but not named `-weekly` would miss the suffix leg.**
It would still appear if it is the slice's last window; otherwise it waits for the
registry to learn the vendor's new word, which is the same posture every verbatim-
vocabulary surface here takes (§7.15's convert rule).
### 7.20 `--hide`: the standing hide list (2026-08-10)
The owner's request, near-verbatim: hide gemini and cursor, because only the CLI vendors
are in use these days. The `v` filter cannot answer it — a filter narrows one launch to
one vendor, and this is the opposite ask: every launch, every vendor EXCEPT two. So the
HUD takes a hide list: `--hide gemini,cursor`, with the env var `TELLTALE_HUD_HIDE` as the
flag's default. The env var is the standing per-machine preference (TELLTALE_ASCII's
precedent) and the flag always wins, including `--hide ""` to see everything for one
launch without unsetting the variable.
**Where the hide is applied.** To the snapshot, as each scan lands — not in Render. The
grid, the vendor lines, the fleet quota strip, the `u` page and the `w` page all read the
same snapshot, so stripping it once is what keeps five surfaces from ever disagreeing
about who is hidden. The header census comes from the same slices, so its counts match the
rows for free.
**How this survives the honesty rule.** A monitor that silently hides rows is a liar
(§7.1), and a hidden vendor has no backstop at all: it leaves the "N of M" census
entirely, and no keypress re-reads the choice. So the footer states `hidden gemini cursor`
for the whole run, and the notice outranks the filter and the query in the drop order —
only the two ⚠ facts sit above it. The `v` cycle skips hidden vendors, because a filter
that can only ever select an empty grid is a dead stop on a one-key cycle; a `--vendor`
naming a hidden vendor at startup is refused loudly rather than opened onto a
contradiction.
**What this is not.** Not an uninstall: the adapters stay registered, and the vocabulary
(`agy`/`antigravity`, `cursor`/`composer`) is parseFilter's own, so the two flags cannot
disagree about what a vendor is called. The list is deduplicated and sorted at parse time
so the footer's wording is stable no matter how it was typed. `--hide all` is refused: a
HUD told to hide every vendor is a request to not run it.
### 7.21 The event sink — every hook, one durable stream (2026-08-11)
**What it is.** `telltale events` is a loopback-only HTTP server plus a durable log. An
emitter POSTs one hook event to `/events`; the sink appends it to a JSONL day file under
`~/.telltale/events/`, and rebroadcasts it to every WebSocket client on `/stream`. Two read
endpoints serve a future viewer: `GET /events/recent?limit=N` (newest first) and
`GET /events/filter-options` (the DISTINCT of the three tag axes: `source_app`,
`session_id`, `hook_event_type`). The design is a re-implementation from a published
observability shape (an emitter, one events table, a live stream); no code was taken from
the reference, which carries no license.
**Why it exists.** The fleet runs four vendors and each one's hooks fire into that vendor's
own log, in that vendor's own shape. The sink is the one place a hook event from any vendor
can land in one shape: any process that can pipe JSON is a source. `tools/emit-event.py` is
the reference emitter (stdlib-only Python, no dependency to install): it reads the hook payload
on stdin, promotes the fields a reader filters on (`tool_name`, `tool_use_id`, `error`,
`agent_id`, `agent_type`, `stop_hook_active`), stamps epoch-millisecond time, and POSTs.
Its hard rules are the hook contract: a quarter-second connect probe before the POST, a
5 second timeout on the POST itself, no retry, and exit 0 on every path — a sink that is
down costs the agent the probe and one stderr line, never a failed turn. There is no
summarization pass anywhere: the payload travels and is stored verbatim.
**Trap 3 — "5 second timeout" never bounded the down path, and `localhost` was the wrong
host (measured 2026-08-15, Windows 11, during a fleet-wide slow-hook audit).** A refused
connect is not a timeout: Winsock retries it internally for ~2 seconds per address family
before surfacing WinError 10061, so the original emitter — one POST, no probe — cost
~4.4s of blocked agent time on EVERY hook event while the sink was down (transcripts
recorded p50 4.3–4.5s per PostToolUse across Read/Bash/Edit/Write). The `localhost`
default doubled the damage and taxed the up path too: the sink binds `127.0.0.1` only,
and Windows resolves `localhost` to `::1` first, so every event paid the full retry
cycle against `::1` before trying the address the sink actually listens on. Hence the
probe (a 0.25s bounded connect answers "is anyone listening"; measured: down sink now
costs ~0.26s plus interpreter start) and hence the default URL naming `127.0.0.1`.
**Distribution is one edit per repo.** The wiring pattern is a hook command of the shape
`python3 /tools/emit-event.py --source-app ` — `--source-app` is the only
per-repo change. Claude Code payloads carry `hook_event_name` and `session_id`, so no other
flag is needed there; a wrapper for a vendor whose payload lacks the name passes
`--event-type`.
**Why the store is JSONL, not SQLite.** The reference keeps an `events` table with indexes
on the three tag axes and the timestamp. This repo takes no dependency for a storage path
(decisions/001), and `internal/sqlite` is a byte-level READER with no write path — so the
contract the indexes serve is met with stdlib parts: one JSONL file per UTC day, the
retention window held in memory, distinct-value sets kept beside it. The two queries the
endpoints need — last N by arrival, DISTINCT of three columns — are exactly what that shape
answers. The WebSocket is hand-rolled for the same reason (`internal/eventsink/ws.go`):
the sink speaks one direction of one frame type, and that is a page of checked stdlib code,
not a dependency.
**Retention, which the reference does not have.** `--retain ` (default 30). The sweep
runs at startup and then hourly: memory drops events past the window, and a day file is
deleted only when its whole day is past it. The file-name pattern is affirmative
(`YYYY-MM-DD.jsonl`), so the sweep can never delete a file it did not write.
**The boundary this moves, said out loud.** This is the first content-bearing store under
`~/.telltale/` — rows carry the hook payload verbatim, not numbers-and-keys. Three facts
keep it inside the read/write contract: it is its own foreground mode the operator starts
(the `otel` precedent — the gauges gain no write), it binds loopback only and refuses any
other host at startup, and nothing in the gauges reads or renders these files. CLAUDE.md's
boundary section names the exception.
**Dark by design, and the v1 gate.** v1 is gate-held (§1) and this subsystem touches no
gate surface: no council code, no HUD or statusline render, no new seat. The sink runs dark
— events accrue and stream, and nothing in telltale displays them yet. A viewer is a later
call site, not a re-plumb; §7.16's held-display precedent is the model.
**Verified live, 2026-08-11, Windows 11.** A fake PreToolUse payload piped into
`tools/emit-event.py` against a running `telltale events`: the POST returned the stored
row, the row landed in `~/.telltale/events/.jsonl`, `GET /events/recent` returned it,
`GET /events/filter-options` listed its three axes, and a WebSocket client connected to
`/stream` received `{type:"initial"}` on connect and `{type:"event"}` on the insert. The
sink-down path was exercised the same day: with no server listening, the emitter printed
one stderr line and exited 0.
**Retention verified 2026-08-11, same box.** The startup sweep deleted a synthetic
`2026-06-30.jsonl` staged beside the live day file and logged `retention sweep deleted 1
day files`, and the same run left `2026-08-11.jsonl` byte-identical by SHA-256 — so the
sweep drops a day past the 30-day window and keeps a day inside it.
**Live drive, 2026-08-11/12.** The runs above used a piped fake payload. A real
vendor-invoked firing is now recorded. On the Windows reference box (Claude Code
2.1.226→2.1.228), an interactive Claude Code session (v2.1.228) in the repo directory
ran one Bash tool call. The project-local PostToolUse hook posted to the sink, and the
sink stored row id 4 with `source_app` `"telltale"`, the real session id, and the
verbatim PostToolUse payload (`tool_input`, `tool_response`, `transcript_path`, `cwd`).
The sink had run since about 11:58 local; the row landed 2026-08-12T02:38Z (timestamp
1786502282870). The drive found two traps, and both are measurements, not readings of a
document.
**Trap 1 — headless print mode does not run PostToolUse hooks.** In `claude -p` print
mode, PostToolUse hooks NEVER ran: not from the project's `.claude/settings.local.json`
with workspace trust accepted, not passed through `--settings`, with and without a
`matcher` key. The measurement is a breadcrumb test. Two hooks wrote a file by two
different mechanisms (`cmd /c echo` to a file, and `python -c` writing a file), and zero
files appeared. So the hooks were not invoked at all — this is not an invoked-and-failed
hook. **Do not generalize this to the council gate's stream-json surface.** §9.8's probe
measured PreToolUse firing over `--input-format stream-json` the same night. The two
runtimes differ; measure each one.
**Trap 2 — `uv run` in a hook command exits before the script runs.** A hook command of
the form `uv run /tools/emit-event.py` exits 2 with the error ``No environment file
found at: `.env` `` on any machine where `UV_ENV_FILE` is set globally. A PostToolUse hook failure is
silent in normal use, so the hook looks wired and stores nothing. The working shape is a
direct interpreter invocation. `tools/emit-event.py` is stdlib-only and needs no `uv`.
The `telltale events` usage text now recommends the direct form.
**Amended 2026-08-16 — a taken 4519, measured and then named.** §7.16a's amendment of the
same day fixed this shape for `telltale otel grok`, and left the identical residue here.
**Measured 2026-08-16**, Windows 11, `main` at `1995b34`, with a throwaway listener holding
127.0.0.1:4519: `telltale events` printed one line on stderr and exited 1.
```
telltale events: listen tcp 127.0.0.1:4519: bind: Only one usage of each socket address (protocol/network address/port) is normally permitted.
```
The failure was already loud and already correctly coded — nothing pretended to collect and
nothing hung. What the line did not carry is what to do next, and there are three parts to
that: who probably holds the port, that `--addr` moves this side (the flag already existed
and was already documented), and that moving this side ALONE stores nothing.
**The likely holder is a different one than the collector's, and that is the whole reason
this message is not a copy of §7.16a's.** 4318 is OTLP/HTTP's registered port, so a
collision there is very probably a rival receiver — Jaeger, the OpenTelemetry Collector, a
vendor agent. 4519 is telltale's own and nothing else in this fleet claims it, so the only
holder telltale can defend naming is **a `telltale events` sink the operator already
started**. The message says that, and it says the cheap thing first: check for the running
sink before moving the port, because if one is already listening the emitters already reach
it and a second sink buys nothing. On a port the operator chose it names no holder at all —
telltale cannot know who took 4600, and it says so rather than guessing.
**What else must move, and why dropping that half is worse here than for the collector.**
The emitters are the other side: `tools/emit-event.py` defaults to
`http://127.0.0.1:4519/events` and takes `--server-url`, so a moved sink needs
`--server-url http:///events` added to the hook command in **every** repo's
`.claude/settings.json` — a second per-repo edit beside `--source-app`. A sink moved alone
listens forever and stores nothing. §7.16a's version of that reads like "grok spent
nothing"; this one is quieter still, because the emitter's hook contract (measured
2026-08-11, and hardened by Trap 3 on 2026-08-15) is a 0.25s probe, one stderr line and
**exit 0 on every path**. A half-moved sink therefore produces no failure anywhere: the
fleet simply looks like one where no hook ever fired. That is why the redirect is in the
error text and in the flag help, not left to this document.
**The redirect is measured here, not merely prescribed.** §7.16a had to mark its
`OTEL_EXPORTER_OTLP_ENDPOINT` line unverified, because moving grok's own exporter needs a
fresh instrumented capture. This side owns both halves, so there was no excuse to leave it
at a prescription. **Measured 2026-08-16**, Windows 11, same build: `telltale events --addr
127.0.0.1:4520` bound and logged `listening on 127.0.0.1:4520`; a synthetic PreToolUse
payload piped into `tools/emit-event.py --source-app tt-moved-port --server-url
http://127.0.0.1:4520/events` exited 0 with no stderr; the sink logged `stored #1
tt-moved-port/sess-moved-live PreToolUse`; `GET /events/recent` returned the row with the
payload verbatim; and `2026-08-16.jsonl` appeared in the store. The run used a redirected
home directory, so it wrote to a temporary store and not to the operator's own.
The suggested port is the failed port plus one, and telltale does not scan for a free one: a
port free at the scan is not free at the bind, and a suggestion that looked verified would
be the dishonest one (§4a.1). The loopback bind stays absolute across the flag —
`TestAMovedPortIsStillLoopbackOnly` pins that, and it matters more here than for the
collector, because these rows carry hook payloads verbatim rather than four token counts.
`TestAHeldPortSaysWhoProbablyHasItAndWhatElseToMove` and
`TestTheDefaultPortCollisionNamesASinkAlreadyRunning` pin both branches of the message, and
`TestAMovedPortBindsAndStoresAnEvent` pins that the way out works.
**`internal/bindaddr`, added by the same change.** The busy-port detection, the loopback
test and the plus-one suggestion now live in one package that both §7.16a's collector and
this sink call; the two messages stay in their own packages, because what a collision MEANS
is per-mode and only the mechanism is shared. The extraction is not tidiness. The detection
carries a measured Windows fact that the portable-looking version gets wrong —
`errors.Is(err, syscall.EADDRINUSE)` is FALSE there for a real collision, because the bind
returns errno 10048 (`WSAEADDRINUSE`) while Windows builds define `syscall.EADDRINUSE` as
one of Go's synthetic `APPLICATION_ERROR` constants (536870914, measured go 1.26). Windows
is the primary target (ADR-002), so a second copy of that check is a second place for a
later reader to simplify it back to one arm and break the only platform CI runs.
`TestARealCollisionIsDetectedOnThisPlatform` provokes a real collision rather than
constructing an error value, so whichever arm a platform needs is the arm its suite
exercises.
**Amended 2026-08-16 — "it binds loopback only" was not containment.** Three facts were
offered above as what keeps a verbatim content store inside the read/write contract, and the
second of them was the load-bearing one. It did not hold against a browser. Measured the same
day (§7.24): a page on another origin posted a forged event into this sink, and — the worse
half — opened `ws://127.0.0.1:4519/stream` and was handed the `initial` snapshot, every
retained hook payload verbatim. A WebSocket handshake is exempt from CORS, so no content-type
rule or preflight was ever going to reach that path; the refusal had to move into the
handler. Every endpoint now refuses a request carrying `Origin`, the stream included and
**before** the upgrade, because a page that reaches `onopen` has already been handed the
snapshot. `/events` additionally requires `Content-Type: application/json`, which is what
`tools/emit-event.py` was measured sending. The containment sentence above should now be read
as three facts plus a fourth: no gauge reads these files, the operator starts the mode, the
bind is loopback, **and a web page is not a sender.**
**Amended 2026-08-17 — the sink gets its first reader, and it is its own mode.** "Dark by
design" above named a viewer as a later call site. This is that call site.
`telltale events view` lists what the sink stored, filters it by the three tag axes or by
day, and follows the store live. `internal/eventview` holds it.
**It is its own foreground mode, and that is the whole reason it may exist.** The four
facts the paragraph above just finished restating are what contains a verbatim content
store, and the fourth of them is that no gauge reads these files. A reader wired into the
HUD would have spent that one: hook payloads would land on a surface that redraws on every
tick, in a process the operator did not start for this. A separate mode spends none of the
four. `telltale snapshot` (§7.22) and `telltale otel grok` (§7.16a) set the precedent, and
`TestNoGaugeReadsTheEventStore` asserts it over the transitive import graphs of
`internal/hud`, `internal/statusline` and `internal/snapshot` rather than leaving it to a
reviewer's memory.
**It reads the day FILES, not the sink's own endpoints, on three grounds.** The sink serves
`GET /events/recent`, `GET /events/filter-options` and `/stream`, and this reader uses none
of them.
- **Trust: §7.24 already settled it, and it settled it the other way round from the
intuition.** That section measured a plain file write planting the same row a POST plants
with MORE control over it, and stated the boundary once: a program running as a principal
the store's ACL admits is trusted by these listeners exactly as far as the filesystem
trusts it. So the HTTP path grants this reader nothing the file path does not. The
endpoints are not closed to it either — `internal/localonly` refuses a request carrying
`Origin`, and a local program sends none — which is the point: a viewer that connected
would be served, and would gain nothing by it.
- **Availability, which is what actually decides it.** The sink is a foreground mode the
operator starts. Its endpoints answer only while that process is alive, and only over the
window that process loaded at startup. The day files outlive it. A reader that needed a
running sink would be dark in exactly the case it is reached for: after the fact.
- **A checkable boundary rather than a promised one.** With no network call in the mode at
all, "it makes no network call" is an import-graph fact. `TestTheViewerOpensNoSocket`
asserts it against `go list`, and says in its own comment what it does not cover.
**What it costs, and what the cost buys back.** Follow mode POLLS the day files on an
interval instead of receiving a push, so the interval is the honest latency bound and
`--interval` names it rather than the banner claiming "live". The trade is not one-sided:
the sink's `broadcast` drops a subscriber whose buffer fills, by design, while a file tail
cannot miss an event that way, because the file is the durable record and the tail only
moves forward through it. One `Tailer` serves both the startup listing and the follow loop,
so an event stored between the two can neither be missed nor printed twice — the seam a
separate priming read would have had to choose a failure mode for.
**Keys on the row, content behind a flag.** A row carries the arrival id, the stamp the
emitter sent, the three tag axes, and the promoted fields (`tool_name`, `tool_use_id`,
`agent_type`, `agent_id`, `stop_hook_active`). The payload prints only under `--payload`.
So does the promoted `error`, and that split is the one judgement call worth recording: the
other promoted fields are keys, while an error message is free text the hook was handed,
which makes it content in the same sense the payload is. The row still prints the WORD
`error`, because whether a row has one is what decides if the reader asks for the body.
Note that `session_id` is a key HERE and content in §7.22: the snapshot renders no session
id at all, because a gauge rollup has no use for one, while the three tag axes are this
subsystem's whole filter surface and the sink already serves them at
`/events/filter-options`. Two different contracts, each stated where it binds.
**Plain text and no colour, following `doctor` rather than the TUI.** This output is read
piped into a file and pasted into an issue by someone asking why a hook stored nothing, so
every distinction it makes is carried by a word — which satisfies the colour rule by having
no first signal that is not one. `--ascii` and `NO_COLOR` are therefore not flags on this
mode. Nothing is truncated either: the widest column is a 36-character session id, and that
is the field a reader carries into `--session`, into a vendor's own log, into an issue. A
clipped one is a session nobody can correlate.
**Zero and absent stay different here too** (§4a.1). A row with no timestamp renders the
word `absent`, never `1970-01-01` — epoch zero through a date formatter produces a date
that reads like a measurement. A `stop_hook_active` of `false` renders as a value and a
missing one renders as nothing at all. Both are pinned:
`TestAnAbsentTimestampIsNotNineteenSeventy` and `TestAMeasuredFalseIsNotAnAbsentField`. The
partial-read rule holds as well: an unparseable line costs that line, is counted, and is
reported on screen, so a store that is 40% unreadable cannot look like a quiet fleet. A
half-written line is not unreadable and is not counted as such — `internal/jsonl` holds it
back until its newline arrives, which matters because the sink appends while this reader is
reading.
`--day` selects a FILE, which is the day the SINK recorded the row, not the stamp the
emitter sent. The two differ across UTC midnight and whenever a sender's clock is off. The
file name is the fact this reader can check; the stamp is the sender's claim, and it is in
its own column to be compared against.
**What it deliberately does not do.** It writes nothing, and does not create the store
directory — a reader that created it would make "the sink has never run here" unanswerable
on the next run. `TestTheViewerWritesNothing` hashes the store before and after. It renders
no gauge, adds no seat, and touches no council code, so it stays clear of the v1 gate (§1)
for the same reason the sink itself did. And it does not summarize, group or count anything
about the payloads: the sink stores them verbatim and this prints them verbatim or not at
all.
**Verified live, 2026-08-17, Windows 11**, built from the branch, with a redirected
`USERPROFILE` so nothing touched the operator's own store. `telltale events` bound
127.0.0.1:4530; a synthesized PostToolUse payload piped into
`tools/emit-event.py --source-app tt-live-probe --server-url http://127.0.0.1:4530/events`
was logged as `stored #1 tt-live-probe/11111111-cccc-4ddd-8eee-000000000009 PostToolUse`.
The sink process was then **stopped**, and `telltale events view --payload` still listed
that row with its payload byte-identical to what the emitter sent. That last step is the
availability argument, measured rather than asserted: with the sink gone, every endpoint
this reader could have used was gone with it. Follow mode was driven the same day against a
synthesized store: it printed the retained tail oldest-first, and a row appended to the day
file appeared within one 500ms interval.
### 7.22 `telltale snapshot` — the read mode whose reader is a program (2026-08-11)
**What it is.** `telltale snapshot` runs one scan and prints the fleet's current gauge
state as a single JSON document on stdout, then exits 0. It is the same scan the HUD
runs — `hud.Scan` plus the account relay — reshaped by `internal/snapshot` instead of
rendered into a frame. Three flags: `--vendor` (one vendor only, the HUD's vocabulary),
`--compact` (one line instead of indented), `--timeout` (default 10s).
**Why it exists.** The fleet's own agents are now a reader of this data, and neither
existing read surface serves them. The statusline answers one vendor in one line of styled
text. The HUD is a full-screen TUI that runs until you quit it, and its output is a frame
of box-drawing characters an agent would have to scrape. Both are built for eyes. An agent
that wants to know whether anything is close to its context window, what the fleet has
spent, or which vendor stopped reading, had no answer that was not a screen-scrape — so it
either did not ask, or it read `~/.telltale/` directly and coupled itself to a cache
format that is nobody's API.
**A separate mode, for `doctor`'s reason.** What it prints goes somewhere else. The HUD
owns the alternate screen; this writes to a pipe and returns. Nothing on this path enters
the TUI, renders a gauge, or touches council — which is what keeps it clear of the v1
gate (§1). It is additive: no existing surface changes.
**The schema, and the four rules that shape it.** The document is
`{schema_version, generated_at, scan_error, fleet, vendors[]}`.
- **Zero and absent stay different** (§4a.1). A measured zero is the number `0`. An
absent value is `null`. No sentinel numbers, and **no `omitempty` anywhere**: an
optional key is always present, carrying `null`. That last part is the sharper half —
a key that vanishes when its value is absent makes "no reading" and "this schema moved
under me" the same observation for the consumer, which is the zero-vs-absent collapse
one level up. `internal/snapshot/testdata/golden/zero-vs-absent.json` pins it beside
the HUD's golden of the same name, and `TestZeroIsANumberAndAbsentIsNull` asserts the
two states differ in JSON *type*, not merely in bytes.
- **A derived value says so.** Each vendor block carries `estimated`, the sorted list of
`model.Field` names whose value here an adapter computed rather than read. It is the
JSON form of the render layer's `~`.
- **"Can't know" is not "absent now".** Each vendor block also carries `unsupported`, the
fields that vendor exposes nothing for, ever. A `null` on a field named there is a
capability statement; a `null` anywhere else is this moment's reading. The HUD spends a
whole column-drop rule on this distinction; JSON gets it for two lists. Both lists cover
the SESSION-sourced fields only, and the first live run is what settled that: an
adapter's quota capability describes what a session exposes, while the `quota` array
comes from the account relay, so listing quota in both put `agy` in the document with
two relayed windows and the word `quota` under `unsupported` — two true statements that
read as one contradiction.
- **Definitive empty states.** A list with nothing in it is `[]`, never `null`. A reader
must never handle two spellings of "nothing".
**Pre-computed aggregates, because the alternative is every consumer doing the same
arithmetic.** `fleet` carries the session count, the liveness census (`live`, `idle`,
`stale`, and `unknown` as its own count rather than folded into `stale`, which is an
age claim those rows cannot support), the vendor census by status, the highest context
percentage anywhere, and the total cost. `context_pct_max` is a max and not a mean: the
fleet question is "is anything close to its window", which an average over idle sessions
hides.
**Two things are deliberately absent.** There are **no per-session rows**. Partly that is
the rollup being the product — an agent wants one answer per vendor, not a list to fold —
and partly it is the read/write boundary: a session's honest identity is its name and its
workspace path, and this surface renders numbers and keys, never content.
`TestNoSessionContentReachesTheDocument` plants markers in every content-bearing field of
a session and requires that none survives, including the session id.
And there are **no token counts**. The relay is wired and the HUD reads it, but the
DISPLAY is held by the owner (§7.16's amendment, applied to grok in §7.16a). A JSON field
is a rendering; adding one here would end that hold as a side effect of a different
feature. `TestSpendIsNotRendered` pins the omission so the day the hold lifts is a
decision, not a drift.
**Quota comes from the account relay and never from a row** (§7.15, §7.1's sixth rule).
Hanging a window off a session would assert a per-session limit no vendor publishes.
`quota_read_at` travels with it, because a quota figure without the age of its reading is
a number the consumer cannot judge.
**It writes nothing.** The gauges' contract holds on this path with one item spare: it
reads vendor stores and the quota relay, calls no network, reads no credential, and does
not even write the quota relay — it renders no quota of its own to relay.
**Verified live, 2026-08-11, Windows 11.** `telltale snapshot` against the reference box's
real stores returned a document carrying all six adapters, 1423 sessions with a liveness
census that summed to that count, `agy`'s two relayed quota windows with their reading
time, and `estimated: ["subagents"]` on claude against `estimated: []` on cursor. The
zero-vs-absent pair appeared in that real document without being staged: `agy`'s
`3p-weekly` window carried `"used_pct": 0` — a measured zero — beside five vendors whose
`quota_read_at` was `null`.
`--compact` returned the same document on one line and `--vendor codex` returned that
vendor alone. `--json`, a positional argument, `--vendor chatgpt` and `--timeout 0` each
printed a corrective error, no document, and exit 1.
That run is also what found the quota-capability contradiction described above; the
contradiction was fixed and the run repeated.
**The stated reader has now consumed it, 2026-08-12.** At 2026-08-12T01:59Z an agent
session ran `telltale snapshot --compact` (binary built at main `01770ec`) and answered
real fleet questions from the parsed JSON: 6 vendors watching, 1,443 sessions, 1 live,
`context_pct_max` 75.8 on codex, and `agy` quota with 4 windows of which `gemini-weekly`
carried 11.9 `used_pct`. Zero-vs-absent held in the document the agent read: `cost` was
`null` everywhere, and a `used_pct` of 0 was the number 0.
**Amended 2026-08-16: the contract is published, and CI is its first measured consumer.**
Everything above was asserted by this package's own tests. A schema that only its author
validates against is a schema nobody has tested, so the contract now lives in a file a
consumer can fetch, and the gate reads that file rather than a second copy of its rules.
- **`docs/snapshot.schema.json`** is the document's JSON Schema, draft 2020-12. It
documents what the shipped binary emits, measured against the real output of a live run
and against the four goldens. Where the code and an intention differ, the code wins.
- **The `test` job validates the BUILT binary's `snapshot --compact` output against it**,
then validates the four golden documents too. The binary's own document runs on a
machine with no vendor CLIs, so the goldens are what cover the shapes that machine
cannot produce: a vendor the operating system refused with its message, a drifted
store, a scan error, and the zero-vs-absent pair. The step also re-asserts the
"writes nothing" half, as a before/after of `~/.telltale`.
It runs after the quota relay smokes, so the runner's document carries relayed quota.
Measured 2026-08-16 against that exact sequence: `agy` arrives with two windows and a
`quota_read_at`, one of them a `used_pct` of 0 — a measured zero on a bare runner,
unstaged — while `claude` arrives with `[]` and a null read time, because full.json's
reset stamps are absolute and now in the past. The step asserts that at least one
vendor still carries a relayed window, so the day that stops being true is a red build
and not a quietly narrower gate.
- **The gate is proved non-vacuous on every run.** Three mutations break the real
document one way the contract forbids, and the validator must reject all three: a
dropped optional key (the `omitempty` regression), the string `"n/a"` where a nullable
number belongs (the sentinel regression), and a `schema_version` bump nobody wrote a
schema for. A gate that has never failed is a gate nobody has measured.
- **The validator is pinned Python, not a Go module.** `go.mod` carries no direct
dependency outside the TUI stack, and this document records that choice four times over
— the SQLite reader (§3.2), the zstd reader, the OTLP listener (§7.16a) and the event
emitter (§7.21) each refused a library. A schema module used only by a test would still
enter the shipped module's graph. Hand-written Go assertions were the third option and
are not a schema gate: they restate the contract in a second place, and two statements
of one contract drift.
Run the gate the way CI runs it:
```
python -m pip install jsonschema==4.23.0
go build -o telltale.exe ./cmd/telltale
./telltale.exe snapshot --compact | python tools/validate-snapshot.py -
python tools/validate-snapshot.py internal/snapshot/testdata/golden/*.json
python tools/validate-snapshot.py --mutate drop-key
```
**What the schema deliberately does not promise.** Each of these is a limit of the format
rather than an omission, and a consumer that reads the schema as a stronger claim will be
wrong.
- **It does not promise which kind of `null` you got.** The schema guarantees that a
measured zero is the number `0` and an absent value is `null`, because the two have
different JSON types. It cannot tell a reader that the `0` in front of it was measured.
`estimated` and `unsupported` are what carry that, and `TestZeroIsANumberAndAbsentIsNull`
is what keeps the emitter honest about it.
- **It does not bound any value.** `context_pct_max` has no maximum and `cost_usd_total`
has no minimum, because nothing in `internal/snapshot` enforces one. A bound the code
does not hold is a claim the schema cannot back (§4a.1).
- **It does not close the object.** Every object sets `additionalProperties: true`. That
is the schema agreeing with `SchemaVersion`'s own rule: a field added at the end is not
a break, because every reader parses by name. `additionalProperties: false` would make
the schema call a break what the emitter does not, and would redden a consumer's
pipeline on a release that added a field. `required` is what carries the weight instead
— a renamed or dropped key still fails, which is the regression that matters.
- **It does not enumerate the vendors.** A seventh adapter adds a value, not a field.
`status` IS enumerated, because those four words are the whole of `hud.VendorStatus`
and a fifth would be a new state a reader must handle.
- **It does not promise the order of `vendors`.** The entries arrive sorted by vendor id
today. Nothing asserts that, so the schema claims nothing about it.
**Amended 2026-08-16: a human-visible consumer ships beside the machine one.** CI validating
the document proves the shape holds. It does not show anyone what the document is for, and a
contract whose only consumer is its own gate is a contract nobody has used. `tools/fleet-prompt.ps1`
is that second consumer: one PowerShell function, one `snapshot --compact` call, one parse,
one line of prompt text.
- **What it renders.** The vendor count that is watching, the session and live counts, any
vendor whose status is not `watching` named with that status, the highest context
percentage with the vendor holding it, the busiest relayed quota window, and the words
`scan degraded` when `scan_error` is not null.
- **What it demonstrates, which is the reason it exists.** The two rules the schema states
survive the trip into a caller. A null drops its whole segment rather than printing 0 or a
dash, which in PowerShell means `if ($null -eq $v)` and never `if (-not $v)` — `-not 0` is
true, so the idiomatic spelling is exactly the one that erases a measured zero. A figure
whose vendor lists `context_pct` in `estimated` keeps the render layer's `~`. An unknown
`schema_version` returns an empty string, because a prompt segment that guesses at a
contract it does not know is worse than one that is absent for a release.
- **Windows PowerShell 5.1, not 7.** That is what a Windows 11 box has before anyone installs
anything, and this is the primary platform (ADR-002). No ternary, no null-coalescing, no
`ConvertFrom-Json -Depth` and no `-AsHashtable`. Output is ASCII and holds no ANSI escapes;
colour is the caller's, and the line has to read on a console that has none.
- **`scan_error` renders as two words and never as its own text.** The message can be long
enough to break a prompt, and it can name a path. A prompt segment is the wrong surface for
a diagnostic string.
**Driven 2026-08-16, Windows 11, PowerShell 5.1.26100.9168**, against the real stores at branch
`snapshot-example` off main `08c300f`. The live document (1554 sessions, six vendors watching,
`context_pct_max` 75.8 on codex whose `estimated` holds `context_pct`, and agy's four relayed
windows) rendered:
```
tt 6 watching | 1554 sessions, 3 live | ctx ~75.8% codex | quota 12.2% agy/gemini-weekly
```
The four goldens drove the shapes a healthy fleet cannot produce, through the same `-FromFile`
path. `watching.json` produced `attn 1: codex unreadable` beside `ctx ~61.3% claude`;
`drifted.json` produced `attn 1: claude drifted` and `scan degraded`; `empty-fleet.json`
produced counts and nothing else, no `ctx` segment at all, because every reading in it is null.
`zero-vs-absent.json` is the pair that matters and it held: its `context_pct_max` of 0 rendered
`ctx ~0%`, a printed zero, against empty-fleet's null rendering as no segment whatsoever. A
`schema_version` of 2 returned the empty string.
The live document is quoted in the pull request unredacted, and no redaction was needed: the
surface renders numbers and keys, so the real output carries vendor ids, counts, percentages and
timestamps and no session name or workspace path. That is the "never content" claim above,
observed rather than asserted.
### 7.23 The drop-file relay — a row telltale did not measure, and says so (2026-08-16)
§4 promises a documented adapter interface, and §4a.7 works one example through. That
promise answers a contributor who will write Go. It does not answer the other question
the launch post raised: what happens to a tool telltale ships no adapter for, and is not
going to?
The owner ruled out a plugin runtime and lifted this feature's demand gate on 2026-08-16.
A runtime would execute a stranger's code inside a process whose entire contract is
"reads, never writes, no network, no credentials" — every guarantee in this document
would become a guarantee about someone else's plugin. The middle path is a **drop file**:
the vendor, or a script the user writes, puts one small JSON document under
`~/.telltale/dropfile/.json`, and telltale renders it as a fleet row. The reader
opens one file and can do nothing else. `docs/dropfile.md` is the spec; this section is
the reasoning.
**It adds no write exception.** The three sanctioned writes (§7.15, §7.16, and council's
`room.json`) are untouched. `internal/adapter/dropfile` creates nothing, writes nothing
and removes nothing; the directory is the operator's to fill, and a missing one reports
the vendor absent. The read/write boundary in `CLAUDE.md` needed no fourth bullet, which
is the strongest evidence this was the right shape.
#### The problem this format has and no other adapter has
Every other adapter reads a store its vendor wrote while doing its own work. Nobody had a
motive to write a flattering number into it, and each adapter's package doc names the live
corpus every field was grepped out of. A drop file is written FOR telltale, by whoever,
and every value in it is that writer's claim. **telltale measured that a file exists, when
it was last written, and what it says. It did not measure the session.**
So §4a.1's rule bites in a new place. The rule has always been about a *value* — measured,
inferred, or absent. Here the whole ROW has a provenance, and a drop-file row must never
be readable as a measured one.
#### Why the `~` estimate marker is the wrong mark
The obvious move is to mark every claimed value with the estimate marker and be done. It
is wrong twice over, and the second reason is the one that settles it.
It states a falsehood about mechanism. `~` means `CapDerived`: the adapter computed the
value from something that is not the value, and the snapshot's `estimated` array says so
in exactly those words. This adapter computes nothing — it reads the number the writer
wrote, verbatim. Marking these rows `~` would claim telltale did arithmetic it did not do.
And it collapses two provenances into one spelling. "telltale inferred this" and "somebody
asserted this" are different claims that a reader would discount differently, and after the
collapse no reader could tell which one they had. That is the failure §4a.1 exists to
prevent, imported through a new door — the same shape as the zero-versus-absent collapse,
one level up.
The rejected alternative is recorded rather than merely dismissed: a per-cell mark of any
kind, `~` or a dedicated glyph. Beyond the two objections above it repeats one fact in
every cell of a row where the fact is uniform, and §4a.2 already refused a per-field glyph
for degradation on the narrower ground that "we failed to read it" starts to read as data.
**Provenance is a property of the row, so it is marked once, on the row.**
#### The three marks, and none of them is a colour
1. **The vendor id is `self-reported` for every drop file.** The grid's identity column
reads `SR` and the header census reads `self-reported 2` in full. `SR` is the only tag
in `vendorTag` that is not a vendor abbreviation, deliberately: every other tag answers
"which tool did this row come from" because telltale measured that tool's store, and
here the only answer telltale can give is where the numbers came from.
2. **The claimed tool leads the row's label** — `windsurf: refactor the parser`. `SR` is
shared by every drop file and cannot separate windsurf from aider, so the row itself
names its claimant rather than hiding it in a pane nobody opens.
3. **`telltale snapshot` emits `"self_reported": true`** on the vendor entry, beside and
never inside `estimated`.
Both HUD marks are plain ASCII words, so `--ascii` and `NO_COLOR` change neither, and
`testdata/golden/self-reported-row.txt` — which renders through `PlainStyles` — is the
whole assertion rather than half of it. The drop-file vendor gets **no hue of its own**:
a hue is a vendor identity (§9.28) and these rows have none to give, so it takes the
identity-hue fallback as Gemini does.
#### Impersonation is unrepresentable, not merely rejected
The format has **no field for a vendor id**. A drop file cannot claim to be Claude Code,
because there is nowhere in the document to make the claim — the id is a constant the
adapter owns. Nor can it choose its own row: the session id comes from the FILE NAME, so
the operator's filesystem decides what a row is called and a document cannot rename itself
onto another one's row.
That is the same "the allowlist is the struct" mechanism `internal/cursorhook` uses against
a payload carrying reply text and an email address, and it is why the planted-credential
test can assert absence rather than sanitization: a key with no destination is dropped by
`encoding/json` before this package sees it.
#### One vendor id, not one per tool
Per-tool ids were considered and rejected. They would put a writer-chosen string in the
column that says what telltale measured — a file naming itself `claude` would draw a row
indistinguishable from a measured one — and `Capabilities()` is per-adapter, so a set of
tools sharing one adapter could not honestly declare different capability sets anyway.
With one id, `Capabilities` describes the FORMAT, which for this adapter is the only source
there is: `unsupported` names the fields the format cannot express, and a file that omits
`cost_usd` yields "absent now" rather than "can't know". Both statements are true, which is
what the merge costs and what it buys.
#### The one claim telltale can check, it checks
A writer that stops writing leaves a file that keeps asserting whatever it last said. A
file claiming `last_activity` of "now" would render live forever over a dead session, and
`--vendor self-reported` would become the way to pin a row to the top of the grid.
telltale cannot check a cost or a context percentage against anything. It CAN check
`last_activity`, because **a file cannot have activity newer than its own last write**, and
the mtime is telltale's own measurement rather than the writer's claim. So a claim ahead of
the mtime is replaced by the mtime, and the substitution goes in `Diagnostics`. The value
stays present and is NOT marked `Degraded` — a measured mtime is a better answer than
absence, and §4a.2 requires degraded fields to be absent.
Staleness proper reuses `internal/quotacache`'s rule and its constants rather than inventing
a boundary: 24 hours to expire, five minutes of future-skew tolerance, and the reading's age
travelling with it past five minutes. A file outside those bounds draws no row at all, which
is that package's own ruling that the honest display for "no reading" is absence.
#### Absence has two spellings on input and one on output
§7.22 emits every key with an explicit `null`, because a reader parsing a document must not
have to tell a missing key from a changed schema. An INPUT cannot hold its writer to that:
a key omitted and a key written `null` both have to mean absence, or the format fails
documents over fields the writer simply had no value for — which is §4a.5's partial-read
rule broken at the door.
So the semantic convention is snapshot's, exactly: **zero is a number, absence is nil, and
no sentinel number stands for either.** Only the syntax differs, two accepted spellings in
rather than one guaranteed spelling out. `TestZeroIsAMeasurementAndAbsentIsNil` walks all
three cases, and `testdata/golden/self-reported-row.txt` pins the render: a claimed 0%
draws a full empty track beside a claimed `$0.00`, and a row claiming neither draws the
absent marker in both cells.
#### The render, generated
```
telltale │ 3 sessions │ claude 1 self-reported 2
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
SESSION MODEL CONTEXT COST AGE
● CC │ telltale C:\src\code Opus 5 ███▊──────── 34% $1.20 │ 12s
● SR │ windsurf: refactor the parser C:\src\code gpt-5-codex ──────────── 0% $0.00 │ 40s
◐ SR │ aider C:\src\work — — │ 9m
──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
q quit / find enter detail u usage w week v vendor s sort a all ? keys
```
The header census says the word in full and the identity column echoes it per row. The
last two rows are the zero-versus-absent pair carried inside the claimed rows: windsurf
claims a measured 0% and $0.00 and draws a full empty track beside a printed zero, and
aider claims neither and draws the absent marker in both cells. This golden renders through
`PlainStyles`, so every mark visible here survives `--ascii` and `NO_COLOR`.
#### What the format cannot express, and why each door was never cut
- **Quota.** An account property (§7.1), sourced from the statusline relay (§7.15). A
session-shaped document has no account to speak for. The usage page says "no quota by
design" for this vendor rather than borrowing one of the six sentences that report a seam
which came up empty — this is a door never cut, not a seam that returned nothing.
- **Liveness.** §4a.4 allows a hint only from a positive vendor signal the HUD cannot see.
This is the field where a false claim is most tempting and least checkable, so there is no
field for it. Liveness is classified by the HUD from `last_activity`, bounded by the mtime
clamp above.
- **Token counts.** They feed §7.17's fleet spend sum. A claimed count summed beside
measured ones would yield a total carrying no mark at all.
#### No canary, no pin, no drift watch
`internal/adapter/pins` gets no row, and `internal/adapter/drift` is not wired in. Both
exist because an adapter reads a private, undocumented format that a vendor may move without
saying so, and the pin names the build the field map was surveyed at. **This format is
telltale's own and it is published**, so there is no vendor build to pin and no shape to
watch drift away from. `schema_version` does the job a canary does elsewhere, and it does it
better: a document whose contract number this adapter does not speak is skipped whole rather
than read leniently, because guessing that the field names still mean what they meant is
inventing every value at once.
#### Schema version
`self_reported` is a field ADDED to §7.22's vendor object, so `SchemaVersion` stays 1 under
that document's own rule: a field added is not a break, because every reader parses by name,
and every object sets `additionalProperties: true`. Nothing already in the document changed
meaning and no key left. `docs/snapshot.schema.json` carries the key in `properties` and in
`required` — the emitter always emits it — and `tools/validate-snapshot.py` passed against
the four goldens and against the built binary's live output.
### 7.24 Who may push to a loopback listener (2026-08-16)
**The question.** §7.16a and §7.21 each open a listening socket, and each treats the
loopback bind as the thing that contains it. Both write a store the product treats as
measured: `telltale otel grok` folds a push into `usage/grok.json`, which is the file the
HUD reads as grok's measured spend, and `telltale events` stores hook payloads verbatim.
Neither listener asked who was pushing. In a product whose whole claim is that a displayed
value came from measured vendor output, an unauthenticated input to that value is a real
seam, so it was measured rather than argued about.
**The measurement.** Windows 11, telltale built from `main` at `65a113c`, go 1.26, Chrome
151.0.0.0 headless. Every run used a redirected `USERPROFILE`, so nothing touched the
operator's own stores. The collector ran on 127.0.0.1:41318 and the sink on 127.0.0.1:41519;
the browser probe page was served from a separate origin, `http://127.0.0.1:41999`.
*A local program forges a measured total.* A stdlib Python script built an
`ExportLogsServiceRequest` carrying one `grok_code.api_request` record and POSTed it to
`/v1/logs`. The collector answered 200, logged `counted api request`, and wrote:
```
{"vendor":"grok","since":"2026-08-16T20:21:11.601084-04:00","written_at":"…",
"requests":1,"input_tokens":999000,"output_tokens":888000,
"cache_read_tokens":777000,"reasoning_tokens":666000}
```
Nothing distinguishes that from a real export. **The forger needs no secret**: the port, the
path and the record shape are all published here and in the package docs, and this repo is
public.
*A plain file write forges the same total, and more of it.* A hand-written `grok.json`
claiming 4242 requests and 111111111 input tokens, with a `since` six hours back, was
accepted whole by `readEntry` — the same function the gauge read path `ReadAll` uses. The
proof is that a later relayed request accumulated **onto** it (`requests` 4242 → 4243) and
kept the hand-written `since`. So the file write is not merely equal to the POST, it is
**stronger**: a POST may only add four non-negative counts to whatever window is open, while
the file writer picks the window's start, its request count and every total outright. The
sink's store measured the same way — a line appended straight to `.jsonl` was served by
`GET /events/recent` after the sink reloaded.
*A web page forges into both stores, with no local code at all.* This is the result that
decided the design. A page on `http://127.0.0.1:41999`, driven by a real headless Chrome,
posted a forged OTLP record into `usage/grok.json` and a forged hook event into the sink. It
works because neither handler read `Content-Type`: a `text/plain` body is one of the three
CORS-safelisted media types, so the request is a *simple* one the browser sends outright,
with no preflight to refuse. The response is unreadable to the page and that changes nothing
— the write has already happened.
*A web page reads the verbatim store.* Worse than the write, and it is the sink's, not the
collector's. A WebSocket handshake is exempt from CORS entirely, and `upgrade` never looked
at the request. The same page opened `ws://127.0.0.1:41519/stream` and was handed the
`initial` snapshot — the last hundred stored events, hook payloads and all. §7.21's
containment claim was "it binds loopback only"; against a browser that claim was doing no
work.
**The refuted argument, and the part of it worth keeping.** The tempting reading of the
first two results is that the seam does not matter: any local program can write
`usage/grok.json` directly, so a secret on the HTTP path buys nothing, and the honest fix is
a documented boundary. **The measurement refutes that as stated, and confirms half of it.**
The two paths do not have the same senders. A web page reaches the socket and cannot reach
the file — it is not a program on this machine at all. The file's reach is bounded by an ACL
(on the reference box `C:\Users\sanle\.telltale\usage` grants Full to the owner, SYSTEM and
Administrators, and Modify to two further profile-inherited principals); the socket's reach,
before this change, included every page the operator visited. So the boundary was not a
boundary yet. The half that survives is the local half: a program running as a principal the
ACL admits is trusted completely, and no HTTP-path secret would change that by one byte.
**What was built: make the socket's senders equal to the file's.** The fix is not
authentication and does not pretend to be. It removes the one class of sender the filesystem
already excludes, and then states the rest of the boundary plainly. `internal/localonly`
carries the check, extracted the way `internal/bindaddr` was for the same two modes.
- **Refuse any request carrying `Origin`.** This is the arm that generalizes and the only
arm that can cover the WebSocket handshake. Measured: every browser request carried
`Origin`, including the handshake — which carried **no** `Sec-Fetch-*` header at all,
which is exactly why the check reads `Origin` and not `Sec-Fetch-Site`. Neither real
sender carried it: `tools/emit-event.py` arrived as `Python-urllib/3.14`,
`Content-Type: application/json`, no `Origin`; the exporter-shaped request the same way
with `application/x-protobuf`. A page cannot suppress the header, because the user agent
attaches it rather than the script.
- **Require the media type the measured sender sends** — `application/x-protobuf` for the
collector (§7.16a's capture pins grok's own exporter to it), `application/json` for the
sink (`tools/emit-event.py`). Parameters are ignored, so `; charset=utf-8` passes.
Measured from the other side: with a non-simple `Content-Type` Chrome sent only an
`OPTIONS` preflight to each endpoint and **no POST followed**, because neither server
answers a preflight with CORS headers.
- **Both arms, because they fail in different directions.** `Origin` names the sender class
but rests on a header a future browser could stop sending on some path nobody has measured;
the media type rests on nothing about browsers at all, and turns any such page into one
that must preflight. Neither is load-bearing alone.
The sink applies the `Origin` arm to its **read** paths and its stream as well, not only the
POST: a page that cannot plant a row can still ask for the rows already there, and those rows
are content. A refusal is `403` — not `401`, because no credential would help, and a 4xx
rather than a 5xx so an OTLP exporter stops retrying instead of looping against a door that
will not open.
**No token, and the reason is the measurement, not the effort.** A shared secret was the
obvious shape and it was refused three times over. It closes nothing against the principal
that matters, because a program that can read a token file beside the store can write the
store directly — measured above, and pinned by
`TestTheFileWriterSetsWhatTheRelayCannot`. It would put a secret on disk next to the thing it
protects, for a reader that already has the disk. And it would gate the collector's only
working path on an **unmeasured** knob: §7.16a already had to mark
`OTEL_EXPORTER_OTLP_ENDPOINT` unverified because moving grok's Rust exporter needs a fresh
instrumented capture, and `OTEL_EXPORTER_OTLP_HEADERS` is the same instrument problem. A
wrong guess there makes the collector count nothing while looking healthy, which is §7.7's
worst failure. **OS-level peer verification was refused too**: loopback TCP peer identity on
Windows needs `GetExtendedTcpTable`, which is not stdlib (decisions/001, the same rule that
hand-rolled the OTLP and SQLite readers), and it does not even answer the browser case — the
peer there is `chrome.exe`, a perfectly legitimate local program. An executable allowlist is a
different and worse contract.
**The trust statement, stated once so it can be quoted.** A program running on this machine
as a principal `~/.telltale/`'s ACL admits is trusted by these listeners exactly as far as it
is trusted by the filesystem, because it can plant the same row either way and the file write
is the stronger of the two. That is a deliberate boundary, not an oversight, and
`internal/usagecache/trust_test.go` pins it so a later session adding a bearer token walks
past the reason it buys nothing. What is **not** trusted, as of this change, is a web page.
**Verified after the change, same box, same probes.** The measurement is only worth what the
re-run says, so the identical pages were driven at the rebuilt binary. The collector logged
`refusing a request from a web page (Origin: http://127.0.0.1:41999)` and counted nothing;
the sink logged the same refusal twice, once for the POST and once for the stream handshake;
and the WebSocket probe, which had reported `EXFILTRATED: {"type":"initial",…}` before the
change, now reports `BLOCKED (error)` and `CLOSED code=1006` with no snapshot delivered. The
other half held too: an exporter-shaped push was still counted
(`counted api request — in 42 · out 42 · reasoning 42 · cache read 42`) and
`tools/emit-event.py` still stored a row and exited 0.
**What this does not claim.** The capture is one machine, one day, one browser engine
(Chromium 151). Firefox and Safari were not driven; the `Origin` behaviour they rest on is
the Fetch and RFC 6455 requirement rather than a measurement here, and §3.4's discipline
applies before extending the claim. Nothing here defends against a local program, and it is
not meant to. The gauges' no-network rule and the loopback-only bind are untouched and remain
absolute: this change only narrows who may talk to a socket that was already loopback.
### 7.25 `telltale mcp` — the same document, in front of an agent (2026-08-18)
**What it is.** `telltale mcp` serves [§7.22](#s7-22)'s snapshot document over the Model
Context Protocol on stdio. An MCP client starts the process, speaks JSON-RPC down its stdin
and reads frames off its stdout, and calls one tool — `fleet_snapshot` — which runs one scan
and returns the document. One flag: `--timeout` (default 10s), which is the deadline for each
CALL rather than for the process. Nobody types this command; it goes in a client's config
once, e.g. `claude mcp add telltale -- \telltale.exe mcp`. That spelling is the CLI's
own — read off `claude mcp add --help` at Claude Code 2.1.233 on 2026-08-18, where
`claude mcp add my-server -- my-command --some-flag arg1` is the documented stdio form — and
not a shape assumed from memory. Running it is the operator's to do; see the closing paragraph.
**Why it exists, given that §7.22 already shipped.** The snapshot mode answered "a reader that
is a program" and it answered it well: CI validates it, and `tools/fleet-prompt.ps1` renders
it. Both of those readers are SCRIPTS somebody wrote on purpose. The reader this repo actually
has most of is an agent, and the honest difference is not that an agent cannot shell out —
several can — it is that a command nothing told it about is a command it does not run. A tool
its client LISTS is the version it reaches without being asked, with the argument names and
the value rules in front of it. That is the same gap §7.22 described one layer up: the data
was honest and reachable, and nothing was reaching it.
**It is a fourth READER, not a second document.** `internal/mcpserver` calls
`snapshot.Encode` on the document `snapshot.Build` produced from the scan `cmd/telltale` runs
for `telltale snapshot` — one scan path (`scanDocument`), one adapter roster (`allAdapters`),
one vendor vocabulary (`snapshotAdapters`), one serializer. The tool result's text content is
byte-for-byte what the CLI prints. That is the whole design, and
`TestTheToolResultIsTheSnapshotDocumentUnchanged` is what holds it: every honesty property
this surface claims is a property of those bytes — a measured zero is `0`, an absent reading
is `null`, no optional key is omitted, `estimated` names what an adapter computed,
`unsupported` names what a vendor can never source, `self_reported` names an entry whose
writer claimed it. A second serializer here would be a second statement of that contract, and
two statements of one contract drift — which is §7.22's own argument for a published schema
over hand-written assertions, applied to itself.
`structuredContent` carries the same document as JSON beside the text, for a client that
reads it that way. It is the identical value marshalled twice by one package, so the two can
never disagree. No `outputSchema` is declared beside it, deliberately: the document's schema
is published at `docs/snapshot.schema.json` and CI validates the shipped binary against that
file, and embedding a copy in the binary would be exactly the second statement just refused.
**One tool, not a family.** Every fleet question — what is close to its window, what has been
spent, which vendor stopped reading, how old is the quota reading — is already one parse of
one document. A second tool answering a subset would have to re-serialize part of it, and that
is where the zero-vs-absent rules get restated ([§4a.1](#s4a-1)).
**The tool DESCRIPTION carries the honesty rules, because the model reads it.** A caller that
does not know them reads a `null` as a zero and an estimate as a measurement — the collapse
this repo exists to prevent, moved one process outward into the agent. So the description says,
in the text the model sees before it decides to call: null is absent and never 0, `estimated`
means telltale computed it, `unsupported` means the vendor can never report it, and
`self_reported` means its writer claimed it. `TestTheToolListNamesTheOneTool` pins all four
words.
**Stdio only, and that is what makes [§7.24](#s7-24) not apply.** telltale's two other
machine-facing surfaces listen on loopback, and §7.24 exists because a loopback bind is not
containment on its own — a measured headless Chrome planted a usage row and read the whole
event store. This mode binds nothing. The client owns both pipes and starts the process, so
there is no third party to refuse and no `Origin` to check. `TestTheServerOpensNoSocket`
asserts the direct imports rather than the transitive graph, and says so: the graph already
reaches `net` through `internal/hud`'s TUI framework, and linking that code is not calling it
([ADR-002](#adr-002)'s distinction).
**It writes nothing**, with the gauges' contract one item spare for §7.22's reason: it renders
no quota of its own to relay. `TestTheServerWritesNothing` drives a whole session with the
home directory redirected and compares the tree before and after.
**The protocol surface is four methods and stops there.** `initialize`, `tools/list`,
`tools/call`, `ping`, plus the notifications a client sends and expects no answer to. Absent:
resources, prompts, completion, sampling, subscriptions, and any server-initiated request.
Their absence is *stated* in the capabilities object rather than discovered by a client that
tried one. Two shapes are refused with the reason rather than half-answered: a JSON-RPC batch
array (removed from MCP in 2025-06-18, and this server answers revisions on both sides of
that), and a message with an id and no method, which is a response to a request this server
never sent and would be a protocol error to answer.
**Version negotiation echoes what the client asked for, within a list.**
`supportedVersions` is `2024-11-05`, `2025-03-26`, `2025-06-18`, `2025-11-25`; the narrow
surface above is spelled identically across all four, so echoing is a true statement rather
than a compatibility guess. An unknown request gets `2025-11-25`, the latest supported, which
is the lifecycle's own rule. **`2026-07-28` is deliberately not on the list, and it is the
newest revision, so the omission is the interesting one**: that revision requires a server to
implement `server/discover`, and this one does not. Claiming the version would be claiming a
method a client is entitled to call — ADR-001's failure in protocol form, a capability
asserted rather than built. The same revision's versioning section says a client may invoke
methods inline instead, which is the path that works here.
**Two error channels, and the split is about who can fix it.** A bad `vendor` argument and a
scan that failed come back as a tool RESULT with `isError` set, because the model asked and
the model can correct. A tool name this server never listed, an unknown method and a
malformed line come back as JSON-RPC errors, because those are the client's plumbing. A failed
call carries no `structuredContent` at all — there is no measurement behind it, and an empty
document would be a fleet with nothing in it, which is a different claim.
**Hand-written JSON-RPC over `encoding/json`, not the official MCP Go SDK.** This follows the
module's standing position rather than inventing one: `go.mod` carries no direct dependency
outside the TUI stack, and this document records the same refusal for the SQLite reader
([§3.2](#s3-2)), the zstd reader, the OTLP listener ([§7.16a](#s7-16a)) and the event emitter
([§7.21](#s7-21)). Four methods and one tool is a smaller surface than the SDK's own API.
**Verified live, 2026-08-18, Windows 11, against the built binary** (branch `mcp-server` off
main `4b58258`, `go build -o telltale.exe ./cmd/telltale`), driven by a scripted stdio client
that writes request lines and reads response lines. Seven messages in, six responses out — the
seventh was the notification, which is answered by silence — exit 0, empty stderr, ~2.3s for
the whole session including one scan of the real stores:
- `initialize` at `protocolVersion: "2025-06-18"` echoed that version and returned
`capabilities: {"tools":{}}` with `serverInfo`; `notifications/initialized` produced no
response at all, which is the half a client would hang on.
- `tools/list` returned the one tool.
- `tools/call fleet_snapshot` returned a document of 8 vendors and 1,523 sessions, and
`structuredContent` parsed EQUAL to the text content. **The zero-vs-absent pair appeared in
it unstaged**: `agy`'s `3p-weekly` window carried `used_pct` as the integer `0` — a measured
zero — beside `gemini-weekly` at `0.2`, while `cost_usd_total` was `null` fleet-wide and
five vendors carried `quota_read_at: null`. `codex` carried `estimated: ["context_pct"]`
against `cursor`'s `[]`, and the drop-file entry carried `self_reported: true`.
- **That document was written to a file and validated against `docs/snapshot.schema.json`
with `tools/validate-snapshot.py`: `ok`.** The published contract holds on the new surface,
checked with the gate the CLI's own document is checked with rather than with a second
reading of it.
- `--vendor chatgpt` came back as a tool result with `isError: true` naming the accepted
words; `fleet_quota` came back as JSON-RPC `-32602` naming the tool that does exist;
`resources/list` came back as `-32601` naming the methods that do.
- **Writes nothing, measured on the real store**: `~/.telltale` held the identical 35 files,
sizes and modification times before and after a full session.
CI drives the same sequence against the built binary on every run, and pipes BOTH spellings of
the tool's document — the text content and `structuredContent` — through the same validator
(`.github/workflows/ci.yml`). **That gate was proved non-vacuous the way §7.22's schema gate
was**: one sentence was removed from the tool description, the binary rebuilt, and the gate
failed naming the missing word; the sentence was restored and it passed again.
**What is NOT verified, stated so nobody reads the run above as more than it is.** **No
third-party MCP client has connected to this server.** The drive was telltale's own scripted
client, which proves the framing, the document and the error channels, and says nothing about
how a shipped client negotiates a version, orders its requests, or renders a tool result.
Wiring one up writes an entry into the operator's own client configuration, which is his to
make; until he does, this section claims a correct server and not a working integration.
`STATE.md` carries the debt.
### 7.26 `telltale history` — what one vendor spent, day by day, from its own files (2026-08-29)
Every reader this product has answers about NOW. The statusline and the HUD render a scan;
§7.22's `snapshot` and §7.25's `mcp` serve one scan to a program; §9.42's `doctor` reports a
preflight. None of them can answer *where did last week go*, and the reason is structural
rather than an omission: the HUD's Claude read is a head+tail parse — 64 KiB in, 256 KiB back
(§3.1) — because §7.18 measured a whole-corpus walk at 164 MB and 46,727 records per second
against a 1 s poll. A history needs every record of every file. **Measured on the owner's
own corpus, 2026-08-29: 701 transcripts, 143,304 records, 10.4 s for a 30-day walk.** That is
four orders of magnitude off a poll tick, so this is a foreground mode that reads once and
returns, on `doctor`'s precedent, and never a page in the HUD.
#### It is SPEND-shaped, and §7.16 and §7.17 are the vocabulary
Nothing here is new grammar. §7.17's table of two claims puts this surface entirely in the
right-hand column, and the consequences are the ones already ruled:
- **No gauge, no percentage, no bar, no countdown, no ceiling.** There is no denominator
anywhere in a token count. `TestTheFrameBorrowsNoneOfQuotasVocabulary` is
`TestUsageSpendBorrowsNoneOfQuotasVocabulary` on this surface, and CI asserts the same
properties on the built binary's stdout.
- **A sum never prints without its window** (§7.16's accumulation ruling). Every count
carries its day; the report carries the span it walked and the zone it resolved days in.
- **Never a fleet total.** §7.17 already rejected one as "arithmetic telltale invented"; this
mode goes further and reports **one vendor per run**, so the arithmetic is not available to
make. The vendor's name is on the title, on the window line and on both refusal
paragraphs.
**One rule here is narrower than §7.17's and is new.** The four columns are not added
together either. Input, cache read, cache write and output are four separately billed
categories and telltale holds no price, so their sum would be a number that reads like a bill
and is not one — the §7.16 derived-`inputTokens` trap arriving through addition instead of
through a vendor's own arithmetic. `TestNoTotalIsRenderedAnywhere` computes what such a total
would render as and fails if it appears anywhere in the frame.
#### Absent is not zero, on two axes
§4a.1's rule, applied to a table instead of to a cell:
| | What it means | What it renders |
|---|---|---|
| a request whose usage block reported zeros | a measured zero — the request happened | `0` in all four columns, `1` request |
| a workspace with no request that day | it sent nothing | **no row** |
| a day inside the window with nothing at all | nothing was written that day | **no row** |
The failure this refuses is the natural implementation: pre-seeding a bucket for every
(day, workspace) pair so the table comes out rectangular. **A rectangular table here is a
table of claims** — every zero cell in it would assert a request nobody sent. What makes a
missing day readable instead is the window line: the report states the span it walked, and a
sentence under the table says in words that a day with no row carried no token-bearing
record. `zero-vs-absent.txt` is the golden, named after `internal/hud`'s for the same reason.
#### The derived value is the DAY, and it is disclosed in words rather than with a `~`
The vendor writes an instant. A calendar day is that instant resolved in a time zone, and the
zone is a choice telltale made — so the window line reads `days resolved in EDT UTC-04:00`,
and the OFFSET is on it because a zone abbreviation alone is ambiguous across regions. This
is deliberately not a `~`: that marker means an estimated VALUE (§4a.1), and a day bucket is
an exact reading under a stated convention. Marking it `~` would say the count might be
wrong, when what is conventional is which side of midnight it fell on.
Two related refusals, both measured rather than assumed:
- **A record with no readable timestamp is in no day.** It is counted and named in
diagnostics, never folded into today — that would move a measurement onto a day nothing
said it belonged to.
- **A record stamped ahead of the clock cannot be dated**, on every adapter's own
`futureSkew` rule. A skewed clock must not be able to invent a day's spend.
#### The survey: which vendors could support this, and which cannot
The mode covers **claude only**, and the six it does not cover are named on **every run**,
each with the reason. That block is not documentation politeness — it is the whole of what
stops a table headed with one vendor's name from being carried away as a fleet answer, and
it prints unconditionally for the reason `doctor` prints its three-state legend every time.
The survey is a **source read of this repository's adapters** at the revision it was written
on — each adapter's record struct, its package doc, and the live-corpus verdicts those docs
already carry. It is **not** a fresh measurement against a live vendor, and the difference is
stated because CLAUDE.md's measured-claims rule makes it load-bearing: the version pins these
verdicts rest on are the adapters' own `VerifiedAgainst` constants.
Two questions, in order. A vendor joins only when both answer yes, and **the second is the
one that surprised the survey**: a count with no timestamp is not a smaller history, it has
no day axis at all.
**Amended 2026-08-29, the same day: the source read was caught out on grok, exactly where
the caveat above says it can be.** grok's row said the record carried a total and no date.
A live re-measure at grok 1.0.5 read a real `turn_completed` record off disk and found a
full input/output/cache split beside the envelope's own `timestamp` ([§3.9a](#s3-9a)'s
2026-08-29 block). The reading of `internal/adapter/grok` was correct — that struct does
parse `totalTokens` alone. The error is that **a record struct is an allowlist, so what it
omits is a decision and not an absence**, and the verdict reported the omission as the
file's shape. The row below is corrected and grok stays uncovered, because nobody has built
or measured the coverage; it is no longer refused on fields. The general rule this buys is
worth more than the row: **before a vendor is built here, re-read its records, not its
struct.**
| vendor | counts on disk? | dated? | verdict |
|---|---|---|---|
| **claude** | **yes** — `message.usage` carries four raw counts (`input_tokens`, `cache_read_input_tokens`, `cache_creation_input_tokens`, `output_tokens`) on every assistant record | **yes** — RFC3339 `timestamp`, and `cwd` on the same record | **COVERED.** Four billed categories, dated, per request, per project. The only vendor with the cache split that makes four honest columns. Pinned at Claude Code 2.1.233 (§3.1) |
| codex | yes — a `token_count` event carries `info.last_token_usage` (this turn) beside a cumulative `info.total_token_usage` | yes — the rollout envelope's own `timestamp` | **the next slice.** Two things owed: which of the two a day may sum is a ruling nobody has made, and there is no cache split, so a codex block carries two columns where claude's carries four |
| gemini | partial — `tokens.input` is `promptTokenCount`, which [§3.7](#s3-7)'s adapter labels a context-occupancy proxy, and the cached subset is not separable from what it parses | yes | refused on UNITS. Summing an occupancy proxy per day counts one conversation's prefix once per turn, under a header that would read like uncached input |
| agy | yes, and the best-guarded in the fleet — `gen_metadata` carries uncached input and output per generation behind the `thinking + answer == output` identity §3.8 requires | **no** — the reverse-engineered field map carries no per-generation timestamp | refused for want of a DAY. Real numbers, no axis to put them on |
| cursor | **no** — `tokenCount.inputTokens`/`outputTokens` were 0 in 310 of 310 message rows; declared `CapNone` (§7.16) | n/a | nothing to read. This vendor's counts arrive by hook, not on disk |
| grok | **yes, and this row said "a total only" until 2026-08-29** — `turn_completed`'s `usage` carries `inputTokens`, `outputTokens`, `cachedReadTokens` and `cacheCreationTokens` beside `totalTokens`, measured at grok 1.0.5 and on disk since 1.0.0 ([§3.9a](#s3-9a)) | **yes** — the envelope's own `timestamp`, which `internal/adapter/grok`'s struct does not parse | **the old verdict described the adapter's STRUCT, not the record**, and both halves of it were wrong about the file. Not refused on fields any more; simply not built. One unit trap is owed first: `inputTokens` INCLUDES the cache read here and claude's `input_tokens` excludes it, so the four columns are not the same four |
| pi | yes — `message.usage.{input,output}` per assistant message, with a `cwd` | yes — record `timestamp` | datable, second after codex. No cache split, and it carries `usage.cost.total` per message — money, which this mode renders nowhere and would have to rule on |
`self-reported` is absent from the table and that is not an omission: §7.23's drop-file rows
are what a tool said about itself, with no session store behind them to walk. Giving it a
"not covered" verdict would imply a file this mode could learn to read.
**The block renders in fixed fleet order**, the same order §7.17's blocks and the header's
per-vendor counts walk. Ordering it by how close each vendor is to coverage reads better as a
roadmap and was **declined**: it would make one list in this product order vendors by a
property no other list orders them by, which is the reshuffle §7.17 spends a paragraph
refusing. The roadmap signal lives in the words — codex's verdict says it is the next slice.
#### The layout, and what each choice is for
The generated render (`internal/history/testdata/golden/ledger.txt`, at the default 100
columns):
```
telltale history — what claude spent, day by day, read from claude's own session files
read from C:\src\home\.claude\projects
window 7 local days, 2026-08-23 through 2026-08-29, days resolved in TST UTC-05:00
read 41 transcripts, 39,184 records
DAY WORKSPACE IN CACHE READ CACHE WRITE OUT REQUESTS SESSIONS
2026-08-24 C:\src\code\telltale 1,204 1,903,551 62,004 13,118 14 2
2026-08-27 C:\src\code\notes-api 96 0 4,102 812 3 1
C:\src\code\telltale 22,140 8,830,112 511,903 140,277 191 5
2026-08-29 ...rkspace-path\telltale 3 12 0 44 1 1
```
- **Plain text, no colour, no TUI** — `doctor`'s argument (§9.42), and it applies harder here:
a history is read in a pipe and pasted into a message. Every distinction the report draws is
a WORD, which satisfies §7.1 rule 2 by having no first signal that is not one. `--ascii` and
`NO_COLOR` have nothing to switch off, so neither is a flag: a flag that does nothing is a
promise that something was configurable.
- **`REQUESTS`, not `TURNS`.** One turn can produce several API requests, so "turns" would be
a count telltale did not take. The column is the number of records that carried a usage
block, which is exactly what was counted.
- **Counts are exact and grouped, never floored.** `theme.Tokens` floors to `1.9M` on the
gauge surfaces because a header line has no room for digits and rounding *up* would invent
tokens nobody was billed for (§7.16). This is a table with the room, so it rounds nothing at
all — strictly the more honest of the two, and affordable only here.
- **The day is drawn once per day**, not restated on a second workspace row: repeating the
date makes the eye read a second reading where there is one.
- **Rows are day-ASCENDING**, so today lands at the bottom, next to the prompt the reader is
looking at. Sorting by spend was refused for §7.17's reason — position is the navigation.
- **The workspace column is the only one allowed to give way**, and it truncates from the
LEFT: a path's identifying half is its tail, and the marker sits at the front where it says
"something was removed" before the reader has read the value. Below the floor the table
overruns the wrap column rather than letting numbers collide (`narrow.txt` is the golden at
60 columns).
- **A record whose own record named no `cwd` gets its own `(no cwd)` bucket**, never a
neighbour's. Attributing it to the last workspace seen in the file would be a guess, and a
guess in the project column is indistinguishable from a reading once it is on screen.
#### The read/write boundary
**This mode writes nothing at all.** It reads one vendor's store, calls no network, binds no
port, reads no credential, and relays no quota — it renders none, which is `snapshot`'s own
argument for holding the contract with one item spare (§7.22). It joins statusline, hud,
snapshot and mcp as a **reader**; `CLAUDE.md` names it in that list.
`internal/history/boundary_test.go` is the mechanical half, on
`internal/eventview/boundary_test.go`'s precedent: `go list` answers what this package
imports, and the gate fails if it ever reaches `quotacache`, `usagecache`, `eventsink`,
`eventview`, `council`, `net/*`, `os/exec` or a TUI module. The check is on DIRECT imports
and says so — this package imports `internal/adapter/claudecode` for `Discover`, and an
adapter's dependency graph is not this mode's write surface.
**Content cannot reach a rendered value**, by the technique `internal/cursorhook` uses against
a payload carrying a user's email beside four numbers: the record struct IS the allowlist, and
`encoding/json` drops every field with no destination. The only strings that survive a parse
here are the workspace path and the timestamp. Diagnostics carry counts and never bytes.
#### What it reuses, and the one thing it does not
Sessions are discovered by `internal/adapter/claudecode`'s own `Discover`, so a session this
mode counts is a session the HUD would draw and the two cannot come to disagree about what a
session is — including the two traps that function encodes (the glob is not recursive, and a
basename is validated as a UUID; recursing inflated the live session list 2.4× and
double-counted every token). Records are framed by `internal/jsonl`, so the U+2028 trap and
the 1,004,230-byte record are handled in the one tested place.
What it does NOT reuse is that adapter's head+tail parse. A ledger needs every record, so it
walks whole files through `jsonl.Scan`. Three record classes are refused and each is counted
in diagnostics rather than dropped silently: unparseable records, `` records
(Claude Code's own locally generated notices, which carry a zeroed usage block and would
otherwise look exactly like the measured zero above), and inline `isSidechain` records — 0 of
179,614 in the live corpus, so a non-zero there is a vendor change worth seeing rather than a
routine skip.
#### Verified against the built binary, not only the suite
The suite drives the host in-process and across a process boundary it starts itself, which is
two thirds of the claim. The last third is that `telltale.exe` really does this, so the shipped
binary was driven by a client that shares none of its code.
**2026-09-01, Windows 11 Pro 10.0.26200, `telltale.exe` built from this branch.** The host was
started as `telltale council host --pipe \\.\pipe\telltale-council-smoke --workspace
--vendor claude --read`, and a **PowerShell** `NamedPipeClientStream` — no Go, no telltale code
— opened the pipe and wrote one line. What came back, whole:
```
{"kind":"welcome","protocol":1,"host_pid":15900}
{"kind":"room","room":{"version":1,"workspace":"…\councilhost-smoke","turn":0,"posture":"read",
"seats":[{"vendor":"claude","binary":"…\claude.exe","phase":"idle","drivable":true}]}}
```
Four things are verified there and each was a separate way to be wrong. The descriptor admitted
a same-user client. `host_pid` matches the process that was started, so the client-side
`GetNamedPipeServerProcessId` check ran against the real server. The seat resolved a real binary
and drew `idle` — **no vendor was spawned**, because nothing was dispatched, which is the "never
start a vendor to see whether it answers" rule holding across the new process boundary. And when
the PowerShell client disposed its stream, **the host exited on its own**, which is detach being
unexposed, measured rather than asserted.
#### Known limitations, named
- **The window is complete or it says so.** A walk stopped by `--timeout` prints what it read
and marks the report incomplete, in its own paragraph: the ROWS stay true and the WINDOW
stops being, and a reader who misses that sentence would read a lower bound as a total.
- **A day is a local calendar day.** A session that crossed midnight in another zone lands
where this machine's zone puts it. The offset is on screen; nothing converts.
- **`REQUESTS` counts records, not API calls, if the vendor ever writes two records for one
call.** No such case is known at 2.1.233; it is named because the column's honesty rests on
the vendor's record-per-request shape rather than on anything telltale can check.
- **Nothing here is cached.** Every run re-walks, at the cost measured above. A cache would be
a ledger that can disagree with the files it came from, and the mode is not on a tick.
### 7.27 `telltale council ls` — the saved room, read and never opened (2026-09-01)
`telltale council` is the only way to see what the saved room holds, and it is an expensive
way. It enters the alternate screen, it detects every vendor, and after this section's sibling
(§9.52) it also rebuilds the seats. An operator who only wants to know *what is saved* must not
pay for a room to find out. `telltale council ls` answers that question and does nothing else.
The mode also has a job that outlives this question. §9.52's rebuild and the later host work
both need one place that reports what is on disk. A discovery surface built under a deadline,
beside the feature that needs it, is a surface that inherits that feature's shape. This one is
built first and alone.
#### It is the SIXTH reader, and it holds the same contract
CLAUDE.md's read/write boundary lists five readers: `statusline`, `hud`, `snapshot` (§7.22),
`mcp` (§7.25) and `history` (§7.26). This is the sixth, and it is closest to `history` — it
reads no scan at all. It reads exactly one file, `~/.telltale/council/room.json`, through
`LoadRoom`, the same loader the room itself uses.
- It **writes nothing**. Not the file it read, not a cache, not a lock.
- It **spawns nothing**. No vendor process starts. `exec.LookPath` is the deepest it reaches,
and that resolves a name against `PATH` rather than running a program.
- It **binds nothing**. No port, no pipe, stdout only.
- It **relays no quota**, so it holds the contract with the same one item spare that `snapshot`
and `history` hold it with, and for the identical reason: it renders no quota of its own,
so it has none to relay.
The reason this is stated at the same length the other five state it is that council is the
product's one ratified exception. A council sub-mode that read like a gauge but wrote like the
room would be the exception growing by accident. This one is a gauge.
#### The three states a seat can be in, and why two of them are not one
§4a.1's rule is the whole of the per-seat output. A saved room names a set of vendors, and each
one is in exactly one of three states:
| what is true | how it renders | why it is its own state |
|---|---|---|
| a session id is saved, and this machine can run the vendor | `saved` | the only row a rebuild can act on |
| a session id is saved, and this machine cannot run the vendor | `saved, not installed here` | the id is real and unreachable from this box. A room opened here will not rebuild it. |
| no session id is saved for this seat | `no thread saved` | a measured absence: the seat was in the roster and never answered, or its thread was cleared |
Collapsing rows two and three would tell an operator on a second machine that a conversation is
gone when the id is on disk and the vendor is missing. That is the same class of error as a
column of dashes for a vendor that could never fill it.
#### What it deliberately refuses to say
**It never claims a thread is alive.** Nothing this mode can read proves that a vendor still
holds a session. Only the vendor answers that, and only when a process asks it to resume. So
the word is `saved`, never `live` and never `resumable`. §9.52 states the same limit from the
room's side: a launched process is not a proven thread either.
**It never verifies an id against a vendor.** Verification means a spawn, a spawn means a
vendor process, and on some seats a resume attempt is billable. A read mode that quietly spends
money is not a read mode. The cost of the honest answer is one word: `saved`.
**It prints no content, because the file holds none.** `room.json` is session ids, a workspace,
a brief PATH and a handful of scalars (`resume.go`'s `SavedRoom` doc comment is the contract).
This mode cannot leak a conversation because it has no conversation to leak.
**It does not restore the posture it prints.** The saved posture is a record of what the room
stood in, and `reattach` already refuses to re-apply it. A reader that printed it as though it
were the posture of the next launch would undo that ruling in a listing.
#### Shape
Words and no colour, on `doctor`'s and `history`'s precedent. Every fact carries its own label,
so no column alignment has to survive a narrow terminal.
**It takes no flags at all, and that is a decision.** The other readers take `--root` to point
at a corpus of *vendor* stores. This mode reads telltale's own state, so `--root` here would
have to mean a different thing, and one flag with two meanings across two modes is worse than
no flag. It is also a two-word mode rather than a `--list` on the room, on the `hook cursor` /
`events view` / `otel grok` precedent: none of the room's flags apply to it, and a flag would
have to explain why it ignored every one of them.
A refused file prints the reason `LoadRoom` gives, in `LoadRoom`'s own words, and exits 0. A
damaged file on disk is a state to report, not an error to fail on — the same ruling the room
makes when it opens anyway and says why. No saved room at all prints one sentence naming the
command that makes one.
### 7.28 `telltale council host` — the room in a process of its own (2026-09-01)
**What this section rules.** Council runs the room and the screen in one process. This section
splits them. A HOST process owns the vendor processes, the pipes and the room state. A CLIENT
process connects to the host and renders. The two speak newline-delimited JSON over a Windows
named pipe with an explicit security descriptor.
**What this section does NOT ship, stated first so nobody reads a promise into it.** Detach is
not exposed. No key detaches the room, no command rejoins one, and nothing survives the client.
The client starts the host, drives it, and kills it on exit. `telltale council` is unchanged:
the daily command still runs the single-process room, and every golden file in
`internal/council` still describes it. This is the process boundary and its transport, built
and measured, with the feature that needs them deliberately withheld. The ownership inversion
is the risk; shipping it alone is how the risk gets reviewed alone.
#### Why a host must PARSE, and why "hold the processes" is not a smaller version of it
The tempting cheap host holds the child processes and lets the client keep reading them. It
does not work, and the reason is mechanical rather than aesthetic.
`pumpStdout` (`internal/council/runner/runner.go`) drains each child's stdout continuously.
Nothing else drains it. If nobody reads, the operating system's pipe buffer fills, and the
next write the vendor makes blocks. The vendor then stops mid-turn. So a host that only holds
processes is a room that silently stops working the moment the reader goes away, which is the
opposite of the property the split is for.
The second cheap version fails on the same fact from the other side. "Do not kill the seats"
leaves the pipe handles owned by the process that made them (`cmd.StdoutPipe()` in
`session.go`). The agents then survive as processes nobody can read and nobody can write. That
is worse than killing them, because it spends quota with no channel to see it on.
**So the host parses. That is what makes it a host and not a babysitter, and it is not
optional.**
#### No pseudo-console, and this is a property of council's seats
A terminal multiplexer hosts a terminal. tmux and Zellij emulate one, which is most of their
size. Council's seats are not terminals. They are line-oriented JSON processes on anonymous
pipes (`session.go`), started with `CREATE_NO_WINDOW` and `HideWindow` so that they have no
console at all (`proc_windows.go`). There is no screen to emulate and no size to track.
`CreatePseudoConsole` is therefore **out of scope by ruling, not by omission.** A later session
must not reach for it by reflex. Council would need it only to host an interactive vendor TUI,
and council drives no vendor that way.
Repaint at any width is already free. `Render` is pure over `State`, and `TestRenderIsPure`
holds it there. A client hands its own width to the same pure function. The hardest problem in
a Unix multiplexer — tell the server the new size, resize the pty, repaint — does not exist
here.
#### The transport: a named pipe, and [§7.24](#s7-24) is the reason
[§7.24](#s7-24) measured a loopback bind and found it was not containment. A headless Chrome,
on a page the operator merely visited, planted a forged row in `usage/grok.json`, planted an
event in the sink, and read the sink's whole verbatim store over `ws://127.0.0.1:41519/stream`.
The fix was to refuse any request carrying `Origin` and to require the measured sender's media
type.
**This socket is strictly worse than those two if it is reached the same way.** It carries
transcript content in both directions, and it accepts dispatch commands. The room writes by
default, and three of the four seats are batch CLIs with no channel to ask permission on. A
page that can post a turn into a hosted room can spend the operator's quota and edit their
working tree. §7.24's two-arm check would have to hold perfectly, on a surface worth far more
than four token counters.
**A named pipe removes the class instead of filtering it.** No URL scheme addresses
`\\.\pipe\...`. `fetch`, `XMLHttpRequest` and `WebSocket` cannot reach it. That makes
`internal/localonly`'s check unnecessary here rather than merely satisfied, and unnecessary is
the stronger of the two positions: the check §7.24 exists to perform has nothing left to
perform. **This is a successor to §7.24's ruling and not an exception to it.** §7.24 narrowed
who may talk to a socket that was already loopback. This section picks a transport that the
excluded sender cannot address at all.
Loopback TCP is therefore **refused**, and the refusal is recorded here so a later session
does not re-derive it. A file-based transport is refused too, on three counts: it cannot carry
a live stream without polling, it cannot answer "is the host alive" without a liveness
heuristic, and it reopens "who may write this file" with no bounding principle. A pipe's
answer to the last one is an ACL the operating system enforces.
**The dependency ruling: `golang.org/x/sys/windows` only, and no `go-winio`.** x/sys/windows is
already a direct dependency and carries `CreateNamedPipe`, `ConnectNamedPipe`,
`DisconnectNamedPipe`, `CreateFile`, `CancelIoEx`, `CreateEvent`, `GetOverlappedResult`,
`SecurityAttributes` and `SecurityDescriptorFromString`. That matches this repo's recorded habit
of a page of checked stdlib code over a dependency (`decisions/001`).
**AMENDED 2026-09-01, the same day, and the amendment is the interesting part.** This section
first ruled that overlapped I/O was the one thing that would justify a dependency and that the
design did not need it — a blocking `ConnectNamedPipe` on one goroutine, one client at a time,
because a second simultaneous client is refused anyway. **That was reasoned, and it was wrong.**
A synchronous pipe DEADLOCKED the room on its first streamed frame. Measured on Go 1.26.6,
Windows 11 Pro 10.0.26200: the host's reader was parked in `ReadFile` waiting for the client's
next command, the host's writer was parked in `WriteFile` holding a 277-byte room frame, and the
client was parked reading. **Windows serialises every operation on a synchronous handle**, so a
read that is waiting blocks a write on the same handle until it finishes.
The mistake was not about pipes; it was about the shape of this protocol. One client at a time
bounds CONCURRENCY and says nothing about DIRECTION, and this host is full duplex on one handle
by construction: it pushes room frames while it waits for commands. Half duplex would mean the
room could only draw when the operator typed, which is not a room.
So both ends open with `FILE_FLAG_OVERLAPPED` and go to `os.NewFile`, which associates them with
the Go runtime's completion port. The only hand-rolled overlapped work is `ConnectNamedPipe`,
about ten lines, and it pays for itself twice: `CancelIoEx` on that pending connect is how
`Close` wakes a waiting `Accept`, which retired a self-connect hack the synchronous handle had
forced. The dependency ruling is unchanged. **The general lesson worth keeping: "one client at a
time" is a statement about how many peers, never about how many directions.**
**The wire format is newline-delimited JSON, one frame per line.** Not for speed. The project
already has the parsers, the framing rule (§4's JSONL framing rule), and the habit across every
seat protocol (`runner/protocol.go`).
#### The security descriptor, measured rather than cited
**The default is a leak, and it must never be used.** With `lpSecurityAttributes == NULL`,
`CreateNamedPipe` gives the pipe a default descriptor. Microsoft's own page says that
descriptor grants read access to the Everyone group and to the anonymous account. This repo's
rule is that a claim about behaviour is measured and not read off documentation
([ADR-001](#adr-001)), so it was measured.
**MEASURED 2026-09-01, Windows 11 Pro 10.0.26200.** A pipe created with a NULL
`SecurityAttributes`, its DACL read straight back off the handle with `GetSecurityInfo`:
```
D:(A;;FA;;;SY)(A;;FA;;;BA)(A;;FA;;;S-1-5-21-...-1001)(A;;FR;;;WD)(A;;FR;;;AN)
```
`WD` is Everyone and `AN` is ANONYMOUS LOGON, each carrying `FR` — `FILE_GENERIC_READ`. The
documentation page is confirmed on this machine. What council's pipe carries instead, read back
the same way:
```
D:P(A;;FA;;;SY)(A;;FA;;;BA)(A;;FA;;;S-1-5-21-...-1001)
```
`GA` is mapped to `FA` at creation, which is why the applied string and the object's own string
are not byte-identical, and why a test must compare ENTRIES rather than the whole line.
**The trap this measurement caught, recorded because it will recur.** The first version of
`TestTheDefaultDescriptorIsTheLeakWeRefuse` looked for the raw SID `S-1-1-0` and came back
CLEAN — it would have reported the pipe safe while Everyone was on it. Windows renders a
well-known SID as its two-letter alias in an SDDL string, and which spelling you get is the
API's choice rather than the object's state. **Both spellings are checked.** A security
assertion that matches on one rendering of an identity is not an assertion.
That test is the NEGATIVE CONTROL for the test beside it. Without it, the positive test only
proves that a string was applied, and not that the string prevents anything.
**The descriptor this design applies:**
```
D:P(A;;GA;;;SY)(A;;GA;;;BA)(A;;GA;;;)
```
- `D:P` makes the DACL protected. No inherited entry is added.
- `SY` is LocalSystem. `BA` is the Administrators group. The third entry is the literal SID of
the account the host runs as, read with `Token.GetTokenUser()`.
- **Administrators are admitted deliberately.** An administrator can already open the host
process for debug, so a denial on the pipe would be theatre. Stating that is better than a
descriptor that looks stricter than it is. `gatehook.go` takes the same posture about `0600`
on Windows.
- **The literal SID, and not `OW` (CREATOR OWNER).** `OW` is a placeholder that an object
substitutes at creation, and whether a named pipe substitutes it the way a file does was not
measured. The literal SID needs no such answer.
`TestThePipeCarriesTheExplicitDescriptor` creates the real pipe, reads its DACL back with
`GetSecurityInfo`, and asserts three things: the three intended SIDs are present, `S-1-1-0`
(Everyone) is absent, and `S-1-5-7` (ANONYMOUS LOGON) is absent.
**Who can still connect: any process that runs as the same user.** That is exactly the boundary
§7.24 ratified and pinned. Quoted from it, because the sentence is the contract: *a program
running on this machine as a principal `~/.telltale/`'s ACL admits is trusted by these
listeners exactly as far as it is trusted by the filesystem.* The descriptor above makes the
pipe **equally** permissive as `~/.telltale/`, and never more.
**Another user on the machine: no, with this descriptor. Yes, read-only, with the default.**
That difference is written down here rather than left to a reader to work out.
#### Peer verification, in both directions
§7.24 had to REFUSE operating-system peer identity for loopback TCP. It needs
`GetExtendedTcpTable`, which is not stdlib, and it does not answer the browser case anyway,
because the peer there is a legitimate `chrome.exe`. **On a named pipe the check is available
and it does answer.** This is the strongest single argument for the transport, over and above
the browser argument.
- **Server side.** The host calls `GetNamedPipeClientProcessId`, opens the client process,
reads its token user, and refuses any client whose SID is not the host's own.
- **Client side.** The client calls `GetNamedPipeServerProcessId` and checks the same way. This
is the anti-squatting arm. The classic Windows attack is a lower-privilege account that
pre-creates a well-known pipe name, so that a later server or client attaches to theirs.
**Measured at `golang.org/x/sys` v0.47.0:** `GetNamedPipeClientProcessId` and
`GetNamedPipeServerProcessId` are both exported. `ImpersonateNamedPipeClient` is **not**. The
design uses the process-id route for both arms rather than adding a `syscall.NewLazyDLL`
binding, and the choice is an improvement rather than a fallback: impersonation changes the
calling thread's token and must be reverted on every path out, and a missed `RevertToSelf`
leaves a thread running as somebody else.
`FILE_FLAG_FIRST_PIPE_INSTANCE` is set on the create. A name another process already holds then
fails the create outright instead of adding an instance to a pipe somebody else owns.
#### Containment: a nested room job, and the property that must not be lost
`proc_windows.go` states the property in its own words: `JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE`
covers "the case we cannot code around" — if telltale dies, the handle closes and Windows reaps
the whole tree. **That property is preserved and moved one process outward. It is not
weakened.**
The per-seat job stays exactly as it is, so that seat eviction and turn cancellation still kill
exactly one tree. One ROOM job is added. It carries `KILL_ON_JOB_CLOSE`, the host assigns
ITSELF into it first, and then every seat lands in the hierarchy under it. The host is the only
holder of that handle.
| how the host ends | what happens to the seats |
|---|---|
| clean quit | the host kills them deliberately, then exits |
| panic or unhandled fault | the process dies, the handle closes, Windows reaps every seat |
| `taskkill /F`, Task Manager End Task | the same — the handle closes, every seat is reaped |
| the machine loses power | nothing survives anyway |
**The nested-job reap is measured, not read.** Microsoft's *Job Objects* page states that a
nested job with `KILL_ON_JOB_CLOSE` terminates its processes and its child jobs when the last
handle closes. `TestAHardKilledHostReapsEverySeat`
(`internal/councilhost/roomjob_windows_test.go`) runs it instead. The test binary re-executes
itself as a stand-in host, which builds the room job, assigns itself, starts a grandchild seat
behind a per-seat job, and reports the grandchild's pid. The test then calls
`TerminateProcess` on the stand-in host, which is what `taskkill /F` does, and asserts the
grandchild is gone. `TestTheRoomJobHoldsTheHostAndTheSeat` pins the structure with
`IsProcessInJob`, so that a pass cannot be earned by the per-seat job alone.
#### The crash blast radius, and the three things that bound it
**The baseline is identical to a telltale crash today.** Every seat dies, the conversation in
RAM is lost, and `room.json` survives. That is the current behaviour with a process boundary
moved, and it is not a regression.
What is new is that **the operator cannot see it happen**, because the host has no terminal.
Three mitigations, and all three are required.
1. **The client renders the death.** A broken pipe is reported as *the host exited, and the
seats went with it*. It must never render the same way as an ordinary disconnect. Two states
render two ways is [§4a.1](#s4a-1)'s whole discipline, applied to a process.
2. **`room.json` stays the floor.** The host writes it on the same schedule the single-process
room does. A hard-killed host therefore leaves the session ids behind, and the next
`telltale council` falls into the existing resume path. **This work is strictly additive: it
never removes the fallback.**
3. **No auto-start, ever.** No Windows service, no Run key, no scheduled task, and no
restart-on-crash supervisor. The operator starts the host and the operator ends it. A host
that resurrects itself is a host the operator cannot reason about, and it would make
`telltale doctor` dishonest.
#### The read/write boundary: what the host writes, and what it must never write
**Transcript content is never persisted. Not under `~/.telltale/`, not in a temporary file, not
compressed, not encrypted, and not "only the last N turns".** It lives in host memory and it
dies with the host.
`resume.go` already ruled this for the same data: every vendor stores its own history against
its own session id, so a second copy here would be a private conversation in a place the user
did not ask for. The rule does not change because the process holding the data changed. The one
precedent for verbatim storage is the event sink, and CLAUDE.md records that its containment
"is scope, not redaction" — its own foreground mode, started by the operator, read by no gauge.
A host is started by a client on the daily path, so that grant's own reasoning refuses to
extend here.
The host writes what council already writes: `council/room.json`, session ids and workspace,
never content. It adds one file, `council/host.json`, and the argument for it is that it is the
same class of file, in the same directory, holding the same class of value, for the same
purpose — the keys that let a later launch find what is already there. It holds `version`,
`pid`, `pipe`, `started_at`, `workspace`, `seats` and `turn`. Four of those are already in
`room.json`; the other three are process facts. `resume.go` already states the leak profile of
this exact shape — which directory was worked in, when, and a set of opaque ids, and not a word
anyone said — and a pid does not change that sentence.
**Liveness: the file says WHAT, the pipe says WHETHER.** A pid is reusable, and a stale
`host.json` is the normal case after a hard kill. So `host.json` is never read for liveness. A
host is running if the pipe opens, and that is not a heuristic.
#### The spawn guard, extended in the same change
`internal/council/main_test.go`'s `TestMain` makes the council package's spawn vars fail closed.
It exists because the opposite default was measured starting `codex exec --json -s
danger-full-access` from a plain `go test` run — a live agent turn, with full write access, on
the operator's own account. **CI can never catch that class**, because CI has no vendors
installed and nothing dispatches.
A host spawns from a DIFFERENT PROCESS and a different package, so it is outside that wrap.
Extending the guard is therefore part of this change and not a follow-up. Two new spawn paths
exist and both are guarded:
- **The host's own vendor spawn.** `internal/councilhost` has its own `TestMain`, which wraps
its spawn vars on the same rule: a binary this machine can resolve panics, and names the call
site and the full argv. The rule is copied rather than re-invented, because it is the same
question the operating system is about to ask.
- **The client's spawn of the host.** This one is sharper. It starts `telltale.exe council
host`, which resolves on any machine that built the binary, and that host then starts real
vendors — so an unguarded test would launch billed turns two processes away from the
assertion. It goes behind a var in package `council`, guarded in that package's `TestMain`
and stubbed in `countSpawns` with the restore added to the existing `t.Cleanup`.
#### Naming: "reattach" is already taken
`reattach` means *resume the vendor session ids saved in `room.json`*. It carries that meaning
in the code (`Reattachment` in `resume.go`), in the golden files
(`testdata/golden/reattached.txt`), and in the demo script. Overloading it would make the
room's own notice ambiguous at the exact moment the notice is trying to be honest about what
was restored.
**Ruling: `reattach` keeps its current meaning. The verb for a live host is `rejoin`.** Two
words, two facts, so a notice can say which one happened. Renaming the old one to `resume` is
cleaner and is not taken here: it is golden-file churn across at least four files, and it is
the owner's call.
`host` is the noun, over `server` and over `daemon`. `server` is the word this product's thesis
refuses out loud. `daemon` is a Unix word on a Windows-first product ([ADR-002](#adr-002)). A
host holds the room.
#### Known limitations, named
- **Host memory is unbounded.** A room accumulates turns for as long as it lives. A turn
ceiling is owed before detach ships, and the drop must be stated in the header rather than
applied silently, on the retention discipline `telltale events` already has (§7.21).
- **One client at a time, and the OPERATING SYSTEM refuses the second.** The pipe is created
with one instance, so a second open comes back `ERROR_PIPE_BUSY` and the client renders that
as "one client at a time". An earlier draft of this section said the refusal would name the
holder's pid. It does not: naming it would mean keeping a second instance open purely to
answer on, and a second accept path on a security-sensitive surface costs more than the pid is
worth. Multi-client attach is a tmux feature, it is not free, and it must not be acquired by
accident.
- **A stale host is deliberately not mitigated.** A host nobody returns to keeps running and
can keep spending. It must NOT self-terminate on idle: a detached room that dies on its own
is precisely the failure the operator cannot see. The answer is discovery, and discovery
ships before detach does.
- **Unix is not built here.** The equivalent is a Unix domain socket in a directory at mode
0700, and it is stdlib. It carries an asymmetry that must be measured into `PARITY.md` first:
`proc_unix.go` records that on macOS a process group does not bind lifetimes, so a `kill -9`
on a host LEAKS every seat there, while on Windows it does not. That is the reverse of the
usual direction. *Amended 2026-09-02: built, and the asymmetry is in `PARITY.md` —
[§7.30](#s7-30).*
- **A GATED room is refused, not hosted.** A gated seat blocks on a question, and carrying that
question and its answer over this wire is a card, a keystroke, and a reply written back down
the vendor's own stdin. None of that is built, so `councilhost.New` refuses
`PostureWriteGated` outright, and a gate arriving mid-turn puts a sentence on the seat's card.
A blocked seat and a slow seat must not render alike.
- **A CONVERSATIONAL seat is drawn as undrivable.** cursor-agent's ACP server cannot be handed a
turn by writing a line — its turn cannot be built until the vendor has answered a request of
the room's own ([§9.36](#s9-36)) — so the host names it and refuses it rather than dispatching
into a column that would never finish.
- **The unwatched-write ruling is still owed, and it is not owed YET.** Detach plus a write
posture plus `--auto` is a risk shape the docs have no ruling on: the room writes by default,
and three of the four seats cannot ask permission. It is not owed here because nobody is
unwatched — the client holds the room for the whole of its life. It must be ruled in the same
change that exposes detach, and never as a default that happened.
- **The room state on the wire is the host's own projection, not `council.State`.** `State`
carries pointers and rendered projections, and folding it onto a wire format is the work that
makes council's own `Model` the client's renderer. That is the next slice, and it is named
here so that nobody reads this rung as having paid for it.
### 7.29 detach, rejoin and kill — the room outlives the terminal (2026-09-01)
**What this section rules.** [§7.28](#s7-28) built a host and withheld the feature that needs
it. This section exposes detach: a client leaves and the host keeps the seats, a later client
rejoins the same live process, and `telltale council kill` ends it on purpose. It also carries
the unwatched-write ruling §7.28 said was owed with it and never before.
**What a reader must not read into it.** No key in the room's TUI detaches anything, and no
golden file moved. The reason is stated at the bottom under *the TUI has no host to leave*, and
it is a fact about what §7.28 built rather than a shortcut taken here.
#### The four verbs, and the one that was reserved
§9.52 reserved `rejoin` and spent nothing on it. This section spends it, and the three words
still name three different facts:
| word | what it is a fact about | when it is true |
|---|---|---|
| **reattach** | the FILE | the room read `room.json` and holds the saved ids |
| **rebuild** | the PROCESS | the room launched a NEW vendor process on a saved id |
| **rejoin** | the PROCESS THAT NEVER STOPPED | a client reached a host that was already running |
`kill` is the fourth and it is not one of that family: it is a verb about the host, not about a
thread. `stop` was refused because it understates what happens — this terminates agent
processes that hold live sessions and spend quota.
#### Detach is an explicit FRAME, and a bare disconnect still ends the room
The tempting version reads a closed pipe as a detach. It is refused, and the refusal is the
whole safety argument of this section.
**A client that died is not a client that left.** A crash, a `taskkill` on the terminal, a
power-off — every one of those closes the pipe exactly the way a deliberate detach does. A host
that could not tell them apart would keep a room running on an INFERENCE, and the room it kept
running would be one nobody chose to leave. That is the state this product refuses everywhere
else: §4a.1 rules that two facts render two ways, and here two facts must not even reach the
same code path.
So the wire grows one frame, `KindDetach`, and the rule is:
| what the client does | what the host does |
|---|---|
| sends `KindShutdown` | kills every seat and exits — unchanged from §7.28 |
| sends `KindDetach` | keeps every seat, re-arms the pipe, waits for the next client |
| closes the pipe with neither | kills every seat and exits — unchanged from §7.28 |
The third row is the one worth reading twice. **A bare disconnect keeps the meaning §7.28 gave
it**, so nothing about a crashed client changed when detach arrived. That is what makes this
change additive rather than a re-definition of the existing behaviour.
#### The host survives the client's PROCESS, and the mechanism was already there
Two properties have to hold, and only one of them is new code.
**The socket close is handled by the frame above.** That is the new part.
**The process exit needed nothing**, and this is the measured half. `spawn_windows.go` starts
the host with `CREATE_NEW_PROCESS_GROUP | CREATE_NO_WINDOW`, and both flags already do the work
a detach needs. `CREATE_NO_WINDOW` gives a console application **its own** console with no
window, rather than attaching it to the client's — so closing the operator's terminal sends
that terminal's `CTRL_CLOSE_EVENT` to the processes on ITS console, and the host is not one of
them. `CREATE_NEW_PROCESS_GROUP` keeps the client's ctrl+c out of the host for the same reason
one process group down. Windows does not end a child when a parent exits, so nothing else was
holding the host to its client except the shutdown-on-disconnect rule above.
**`DETACHED_PROCESS` is still NOT used, and §7.28's deferral is now a decision.** §7.28 named it
as what a host "wants" and refused to reach for an unmeasured flag. It stays unused, because the
property it would buy — no console at all — is one `CREATE_NO_WINDOW` already delivers for this
purpose, and the two flags are mutually exclusive. Swapping a measured flag for an unmeasured
one to buy a property already held would be the guess [ADR-001](#adr-001) refuses.
`TestADetachedHostOutlivesItsClientProcessAndStillReapsEverySeat` measures the claim rather than restating it: a real
host in a real process, a real client that detaches and then EXITS, and the host still answering
a second client afterwards.
#### The containment property is not weakened by detach, and that is measured too
§7.28's table says a hard-killed host reaps every seat, and
`TestAHardKilledHostReapsEverySeat` runs it. A detached host is the case that table was written
for and never covered: the host is now the only process left, so if detach cost the room job
anything, nothing at all would reap the seats.
**The same test carries both halves, and that is deliberate rather than tidy.** They are one
story — a host that outlived its client is exactly the host whose containment has nobody left to
check it — and splitting them would have meant two stand-in hosts measuring one process's life.
So the second half runs on the first half's host: a stand-in process runs a REAL `Host.Serve`
(real `NewRoomJob`, real `Listen`, real handshake) and starts a stand-in seat with **no per-seat
job of its own**, so the room job is the only thing that can reap it. The client process
detaches and exits, a second client rejoins from the test, and only then does the test call
`TerminateProcess` on the host the way `taskkill /F` does and assert the seat is gone.
**The seat is asserted ALIVE before the kill, and that assertion is the control.** A reap test
whose subject had already died would pass while measuring nothing, which is the failure mode a
containment test can least afford.
The no-per-seat-job detail is what makes it a measurement rather than a ceremony, and it is
carried over from §7.28's own test for the same reason: the per-seat job also carries
`KILL_ON_JOB_CLOSE` and its handle also dies with the host, so a seat wrapped in one would be
reaped by the mechanism that already existed.
#### Rejoin: the file says WHAT, the pipe says WHETHER, and the PID says WHO
§7.28 ruled that `host.json` is never read for liveness and left the probe unbuilt, because
nothing needed it and the obvious implementation was wrong. That ruling stands and the probe is
built here, on three readings that are three different questions:
1. **WHAT would I be rejoining?** `host.json`: pid, pipe name, start time, workspace, seats,
turn. Numbers and keys, and the file is unchanged from §7.28.
2. **WHETHER a host is running.** `WaitNamedPipe` with a zero timeout, which asks whether the
NAME exists and **does not connect**. That distinction is the whole reason discovery.go
refused to build this before: a probe that dials CONSUMES the host's single pipe instance,
and the host reads the close as its client leaving. Three answers, and all three are used —
`ERROR_FILE_NOT_FOUND` means no host, success means a host with nobody attached, and
`ERROR_SEM_TIMEOUT` means a host whose one client seat is already taken.
3. **WHO that pid is.** A pid is reusable, so the pid alone answers nothing. The probe opens the
process, reads its image name with `QueryFullProcessImageName`, and requires the same
executable name this binary runs under; then it reads the process creation time with
`GetProcessTimes` and requires it to be no LATER than the `started_at` the file claims. A
recycled pid fails the first check; a different telltale that took the number fails the
second.
**`WaitNamedPipe`, `QueryFullProcessImageName` and `GetProcessTimes`, measured at
`golang.org/x/sys` v0.47.0:** the last two are exported and are called directly. `WaitNamedPipe`
is **not** exported, so it is bound with `NewLazySystemDLL` — the same call `IsProcessInJob` in
`roomjob_windows.go` already makes, and the one `decisions/001` sanctioned for the hand-rolled
OTLP reader, the byte-level SQLite reader and the hand-rolled WebSocket. A page of checked code
rather than a dependency.
**All three readings must agree before a client rejoins.** A file with no pipe is a host that
died. A pipe with a pid that is not telltale is a name somebody else took, and `Dial`'s existing
server-side peer check is the second arm on that. Neither of those is an error to fail on; both
are states to render.
#### The four states this feature can leave an operator in, and none of them render alike
§4a.1's rule is the whole of this subsection. `rebuilt`, `survived` and `died` must never render
alike, and detach adds a fourth that must not look like any of them.
**You left** — printed by the client that detached:
```
detached. the host keeps the seats and the conversation, and it is pid %d.
`telltale council` rejoins it. `telltale council kill` ends it, and every seat with it.
```
**You came back and it was still there** — the `rejoin` case, and the only one of the four in
which a vendor process was never restarted:
```
rejoined the host that was already running — pid %d, started %s.
the seats kept working while you were away. nothing was rebuilt, and no session was resumed.
```
**You came back and it was gone** — the `died` case:
```
the host you left is gone, and the seats went with it.
it was pid %d, started %s, and nothing on screen could say when it ended.
the room's session ids are still in %s, so `telltale council` rebuilds those seats.
```
The third line points at §9.52's rebuild, and it uses §9.52's own word. A room that told an
operator their conversation was gone when the ids are on disk would be the same error `council
ls` refuses to make about a vendor that is missing from one machine.
**You asked to leave and the room refused** — the unwatched-write ruling, below.
The `rejoin` sentence carries the clause `nothing was rebuilt, and no session was resumed`
because that clause is the entire difference between this state and §9.52's `rebuilt`. Without
it the two are one sentence apart and an operator would read a rebuild as a survival, which
§9.52 calls the most expensive lie this surface can tell.
#### THE UNWATCHED-WRITE RULING (owed by §7.28, paid here)
**A room that writes to the workspace without asking does not detach. The host refuses it, and
it says why.**
The risk shape §7.28 named is detach plus a write posture plus `--auto`. On a hosted room those
three collapse into one condition, and the collapse is worth stating rather than leaving a
reader to derive:
- §7.28 already refuses `PostureWriteGated` outright — a gated seat blocks on a question this
host cannot carry. So a hosted room is never gated.
- A hosted room that is not read-only is therefore an **ungated** write room: every tool call
runs with nobody to ask. That is exactly what `--auto` means on the room's own surface
(`dispatch.go`'s `seatPosture`: write plus not-asking is `PostureWrite`).
So the condition is one word — the room's posture — and the refusal is keyed on it.
| posture | detach |
|---|---|
| read | allowed |
| gated write | allowed by this ruling, and **unreachable**: §7.28 refuses to host a gated room at all |
| write (ungated, which is `--auto`) | **REFUSED** |
The refusal sentence, verbatim, and it is one sentence on purpose:
```
this room writes to the workspace without asking, so it will not detach: telltale never leaves an agent working while nobody is watching.
```
The remedy is a second line rather than a longer sentence, because §9.17's tell is that a
refusal without a remedy is this room's stated defect and a run-on sentence is not a remedy:
```
the room is still here and still yours. open it with `telltale council --host --read` to get a room you can leave.
```
**Three things this ruling is NOT.**
It is **not** a claim that the write posture is unsafe. The room writes by default and that
ruling stands ([§7.28](#s7-28)'s parent, and the `--read` opt-out's own reasoning). What changes
is only whether the operator may walk away from it.
It is **not** a supervisor. Nothing watches the room for the operator, nothing re-approves
anything, and nothing self-terminates. The refusal is a refusal.
It is **not** the option the costing recommended. The scope ladder that produced this rung
recommended allowing the detach and reporting afterwards what happened while nobody watched —
turns dispatched, tool calls approved, files touched. **That was reasoned and it is overruled by
the owner** (2026-09-01). The report it proposes is a record of an act that already happened,
and the product's whole claim is that it does not act unwatched; a receipt is not consent given
in advance. The recommendation is recorded here so a later session does not re-derive it as new.
**Enforced in the HOST, never in the client.** The host is the process that would keep running,
so it is the process that must refuse. A check in the client alone would be a check a second
client could simply not make. `TestAWritingRoomRefusesToDetach` pins it against the host, and
`TestAReadRoomDetaches` is its positive control — without that pair, a refusal that refused
everything would pass.
#### `telltale council kill` — the fifth surface, and it is the room's own executioner
Sub-noun before the flag set, matching `ls` and `host` and matching `hook cursor` / `events
view` / `otel grok`. It takes no arguments, for §7.27's reason: none of the room's flags apply.
It reads `host.json`, runs the same three-part probe rejoin runs, and then calls
`TerminateProcess` on the pid. **That is deliberately the blunt instrument and not a shutdown
frame.** Three reasons:
1. It is the mechanism §7.28 already MEASURED. The room job carries `KILL_ON_JOB_CLOSE`, the
host holds the only handle, and `TerminateProcess` is what `taskkill /F` does — so `kill`
leans on the property `TestAHardKilledHostReapsEverySeat` proves rather than on a second
path that would need its own proof.
2. A shutdown frame needs the pipe, and the pipe may be held by a client. A `kill` that could
not end a room BECAUSE somebody was in it would be useless for the case it exists for.
3. The word says so. `kill` is honest about ending agent processes mid-turn.
It **refuses** rather than guessing when the probe disagrees with the file: a pid that is not a
telltale process is reported and not terminated, and a stale `host.json` is removed with a
sentence rather than acted on. Killing a pid the file names and the probe cannot confirm is the
one failure this command could make that nothing could undo.
#### `telltale council ls` gains a live-host section and stays the SIXTH READER
§7.27's contract is unchanged and is re-stated here because the temptation to break it is
exactly what a new section invites.
- It **still writes nothing** — including no cleanup of a stale `host.json`. A reader that
tidied would be a writer. **The ROOM removes that file instead**, on the died path, and the
asymmetry is the point: council is already the ratified writer of that directory, and
`host.json` is a file this same feature added, so a room that removes a record of a process it
has just proved is gone is tidying its own state. `TestCouncilLsLeavesAStaleHostFileAlone`
pins the reader's half.
- It **still binds nothing and connects to nothing.** The liveness probe asks whether a NAME
exists; it does not open the pipe. That is why the probe had to be built the way it was:
a dialling probe would have made `ls` capable of ending the room it was listing.
- It **still spawns nothing.**
- It **still relays no quota**, so it holds the boundary with the same one item spare.
What it prints is what the probe measured, and the three states stay apart in the same way the
seat states do: a live host, a `host.json` whose process is gone (named as stale, with the
remedy), and no file at all.
#### The TUI has no host to leave, so no key was added and no golden moved
**`telltale council` runs the single-process room, and it always has.** §7.28 built the host
beside it and wired no daily path to it, so there is no host for a key in the TUI to detach
from. A `d` in that room would have to either do nothing or lie, and a hint on the help panel
for a key that does nothing is the honest-gauge failure this project exists to prevent, spent on
its own surface.
So the way into a hosted room is **`telltale council --host`**, and it is an opt-in flag rather
than a change to the daily command:
| command | what happens |
|---|---|
| `telltale council` with no live host | the single-process TUI room, unchanged |
| `telltale council --host` | a hosted room, drawn by §7.28's plain client, which you can leave |
| `telltale council` with a live host | **rejoins it** — §7.28's plain client again, with the rejoin notice |
| `telltale council` with a dead host's file | prints the died notice, then opens the ordinary TUI room, which rebuilds from `room.json` (§9.52) |
| `telltale council` with a live host somebody else is in | **refused**, naming `telltale council kill` |
The last row is a refusal rather than a fall-through, and that is the load-bearing one. Falling
through would open a SECOND room over the same workspace on the same saved session ids, which is
two rooms rebuilding one conversation — worse than any refusal.
**The renderer is §7.28's plain-text `Render`, and that is stated rather than hidden.** §7.28
calls it "a legible proof that the wire carries a whole room — not a second council TUI", and
making council's own `Model` the client's renderer is still the next slice. So a rejoined room
looks different from the TUI room, and the client's banner says so in its own words rather than
letting the operator discover it.
**The client is line-oriented, deliberately.** A blank-prompt line dispatches a turn; `/detach`
leaves; `/quit` ends the room; `/interrupt` abandons the turn in flight. Keys would need the TUI,
the TUI needs `Model` on the wire, and that is the slice this rung is not.
#### What this rung deliberately does not do
**It persists no transcript.** Unchanged, and it is the rule this whole area is built around. A
rejoining client is handed the host's CURRENT projection over the wire, not a replay from disk.
The host holds the conversation in memory and it dies with the host. `resume.go` ruled this for
the same data, §7.28 restated it, and a second client arriving does not make a second copy any
more acceptable.
**It does not bound host memory.** §7.28 named this and said a turn ceiling was owed before
detach shipped. It is **not paid here**, and saying so is better than a ceiling picked without a
measurement: nothing has yet measured what a room accumulates per turn, so a number here would
be a guess presented as a limit. A detached room that runs for a week grows, and the honest
statement is that nobody has measured how fast.
**It adds no auto-start, no service and no supervisor.** §7.28's third mitigation is unchanged
and this rung is where it would have been tempting. The operator starts the host and the
operator ends it.
**It does not make a stale host self-terminate.** §7.28 refuses that outright — a detached room
that dies on its own is precisely the failure the operator cannot see — and detach is the rung
that makes the refusal cost something. What answers a stale host is discovery, and discovery
shipped first, on purpose.
**It is Windows-only, like the host it extends.** The Unix equivalent carries the asymmetry
`PARITY.md` already owes: on macOS a process group does not bind lifetimes, so a `kill -9` on a
host LEAKS every seat there. A detached host makes that asymmetry worse rather than equal, which
is a reason to measure it before building rather than to build and label. *Amended 2026-09-02:
no longer Windows-only — [§7.30](#s7-30) builds the Unix transport, measures the asymmetry on
Linux, and records it in `PARITY.md`.*
**No vendor has been dispatched to through a host, detached or not.** Every seat spawn in this
package's suite is stubbed, by design — the spawn guard exists to stop a test spending a real
turn — so neither CI nor a session can close that debt. It is the operator's to pay, and
`STATE.md` carries it.
### 7.30 the room outlives the terminal on macOS and Linux too (2026-09-02)
**What this section rules.** [§7.28](#s7-28) built the host on a Windows named pipe and a Job
Object and withheld the Unix side until one measured asymmetry reached `PARITY.md`. [§7.29](#s7-29)
exposed detach on the same terms. This section builds the Unix transport and the Unix containment,
so that `telltale council --host`, `/detach`, rejoin, `telltale council ls` and `telltale council
kill` do on macOS and Linux what they do on Windows — and it states, rather than labels, the one
thing they do differently. The owner's daily machine is a MacBook, so before this the crew tool's
central feature refused on the machine it was for, and `council ls` there printed `telltale
council --host opens a room in one` as the remedy for an absence: a true absence with a false
remedy. That sentence is unchanged and is now true on every platform; the fix was the transport,
not the words.
**What a reader must not read into it.** No frame was added to the protocol, no field to
`host.json`, and nothing about what reaches disk changed: the room's conversation still lives in
host memory and dies with the host. `protocol.go` is untouched. The seats' own containment
(`runner/proc_unix.go`, a process group per seat) is untouched too, and the reason is the next
heading.
#### The containment is the host's SESSION, and it is not a Job Object
The Windows room job has one property no Unix primitive has: `JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE`
binds every seat's **lifetime** to the host's last handle, so a host that dies by any route takes
its seats with it, and a child cannot leave the job by anything it does. `runner/proc_unix.go`
measured the Unix side on 2026-08-17: a process group NAMES a set of processes and does not bind
their lifetimes — SIGINT, SIGTERM, SIGHUP and SIGKILL on the parent, and the `Setpgid` child
survived all four.
So the Unix containment is built from the two facts that do hold, and `roomjob_unix.go` states
the difference in as many words:
- **Membership is inherited and cannot be lost by accident.** The client starts the host with
`Setsid`, so the host leads a session of its own with no controlling terminal; every seat, and
every shell a seat starts, is born carrying that session id, and it changes only when something
in the tree calls `setsid(2)` itself. **That is the one escape, and it is the measured
difference: a child that calls `setsid` leaves the session and nothing here can see it go. A
Job Object child cannot leave.** Nothing council dispatches is known to do this; a vendor that
daemonised its tool calls would be a `PARITY.md` finding.
- **Lifetime is not bound, so the host reaps.** `NewRoomJob` installs a SIGTERM/SIGINT handler
and ignores SIGHUP; `Serve` turns a caught signal into `Shutdown`, which kills every registered
seat through its per-seat group and exits on the ordinary path, removing `host.json`.
`telltale council kill` sends SIGTERM to the host, SIGTERM to every other process in the host's
session at the same time, waits a bounded grace for the host, and then sweeps the session with
SIGKILL — so a host that could not run its handler still has its seats ended by the command
that ends it. On a **dead** host's stale file, `kill` sweeps the dead session too and prints
how many processes it ended, because a dead leader's session id still names its orphans and
the kernel does not hand that pid out again while any of them carries it.
**Why the session and not the host's process group.** A room-level container must hold every
seat; a per-seat container must kill ONE seat's whole tree on an interrupt or an eviction without
touching the others. `runner/proc_unix.go` builds the second from a process group per seat, and
it is the containment the ordinary room measured and relies on. Seats that joined the host's own
group would give that up: `kill -TERM -pgid` on one seat would signal every seat. A session sits
one level above the process groups and holds all of them — the same nesting the Windows side has,
the room job over the per-seat jobs, in the platform's own hierarchy.
**What is left uncovered, stated once here and once in `PARITY.md`:** `kill -9` the host and walk
away, and the seats keep running — holding sessions and spending quota — until the next
`telltale council kill` sweeps the dead session. `TestASigkilledHostLeaksItsSeatAndKillSweepsIt`
pins BOTH halves: the leak, so the asymmetry cannot be forgotten, and the sweep that answers it.
#### The transport: a socket in the council directory, and two locks beside it
The socket is `~/.telltale/council/telltale-council-.sock`, in the directory council already
writes, and `Listen` refuses a directory whose mode admits another account (`chmod 700` is the
named remedy). The node is 0600, and both connect directions read the peer's uid and pid from the
kernel — `SO_PEERCRED` on Linux, `LOCAL_PEERCRED` plus `LOCAL_PEERPID` on macOS — so the program on
the other end is trusted exactly as far as the filesystem trusts it, which is §7.24's boundary
unchanged. A browser has no URL scheme for a filesystem path any more than for `\\.\pipe\...`, so
§7.28's transport ruling stands; the gate that enforces it (`TestTheHostBindsNoPort`) moved from
"this package does not import `net`" to "this package reaches `net` for Unix domain sockets, by
name, and for nothing else", and it walks every source file on every platform's build tags.
**One client at a time is enforced by the host here, not by the kernel.** A pipe with one
instance refuses a second open; a Unix socket accepts into a backlog whether or not anybody will
ever `accept`. So the listener runs its own accept loop for its whole life, hands a connection to
`Accept` only when `Accept` is waiting, and answers everything else with `KindRefused` at once,
naming the pid that holds the room. A second client is told "one client at a time" instead of
sitting in a backlog until the first one leaves.
**The probe cannot connect, and the socket alone cannot answer it.** `WaitNamedPipe` asks "is an
instance available" and opens nothing. A socket path answers only "does a node exist", and a node
outlives a hard-killed host; a probe that CONNECTED would be accepted the moment the host was
free, read as a client with no handshake, and would end the room it was listing. So two facts are
carried by `flock(2)` on two zero-byte files beside the socket, and a probe reads each by trying a
SHARED lock without waiting — it fails at once if the host holds the exclusive one:
- `.lock`, held for the listener's whole life: somebody is listening. A second `Listen`
fails on it with the sentence `FILE_FLAG_FIRST_PIPE_INSTANCE` earns, and a stale socket node
whose lock is free is unlinked safely.
- `.held`, held while a client is attached: the room is held. `Dial` reads it before
connecting, so a client never connects into a held room; the host's refusal is the second arm.
A flock is released when the process exits by any route, SIGKILL included, so a dead host never
reads as listening. That is the property a pid in a file cannot have, and it is why the file says
WHAT and the lock says WHETHER.
**flock, and not fcntl record locks — measured.** The first cut used `F_SETLK` byte ranges on one
file, two facts on one node, and every in-process test failed: fcntl locks belong to a PROCESS,
so a probe inside the host's own process read its own lock as free, and closing the probe's
descriptor released the listener's locks with it. flock belongs to an open file description, so a
second open in the same process conflicts exactly as another process would. The cost is a second
zero-byte node.
**The pid says WHO, in four readings.** `kill(pid, 0)` alone answers "is there a process I may
signal" and nothing else, and a pid is reusable. `verifyHostProcess` asks four things: the pid is
alive and this user's; it leads its own session (`getsid(pid) == pid`, which every host does);
its image name is this binary's (`/proc//exe` on Linux, `kern.proc.pid`'s `p_comm` on macOS,
sixteen bytes and compared as a prefix); and it started no later than `host.json` says
(`/proc//stat` field 22 against `btime`; `p_starttime` on macOS). A recycled pid fails on the
second, third or fourth, and an identity that cannot be read is "not a host", never "gone".
#### What is measured, and where
**Linux, in this repository's own CI.** The `race` job runs the whole suite on `ubuntu-latest`,
and every claim above is a test there: `TestOneClientDrivesAHostedRoomEndToEnd` and
`TestAHostAcrossAProcessBoundaryTakesAClient` (the transport, the peer check, the one-client
refusal by both arms, a bare disconnect ending the room);
`TestADetachedHostOutlivesItsClientProcessAndKillReapsEverySeat` (a client process detaches and
exits, the host survives, a second client rejoins and receives the projection, and `kill` reaps a
seat reachable only through the session); `TestAProbeDoesNotConsumeTheHostsSocket` (five probes,
then a connect, then `PipeBusy` with a client attached); `TestAStaleHostFileIsNeverMistakenForALiveHost`
(a pid that ran and exited, reported dead by `Probe` and left on disk, removed by `kill`);
`TestASigkilledHostLeaksItsSeatAndKillSweepsIt` and `TestTheIdentityCheckRefusesARecycledPID`
(the containment and the identity, on stand-in processes). The built binary was also driven by
hand on Linux on 2026-09-02 — `--host --read` with `/detach` on stdin, `ls`, a rejoin, `kill`, and
a `kill -9`'d host reported stale by `ls` and left alone — and the host's `ps` line read
`SID == PID`, `PPID 1`, no TTY.
**macOS is measured by the test suite, and not yet by hand (amended 2026-09-02).**
`peer_darwin.go` and `identity_darwin.go` compile under `GOOS=darwin` for both architectures, and
every call they make exists in x/sys v0.47.0. The `darwin` CI job's first run on Apple Silicon
(the crew PR, run 33657422294) ran this package's `_unix_test.go` suites in-process on the
runner: host, client, detach, rejoin, one-client refusal, kill and the stale-file path all
passed there, which exercises `LOCAL_PEERPID` on every dial. What that first run also found: a
sandboxed home under `/var/folders///T/` put the socket path at 105 bytes, past
macOS's 104-byte `sun_path`, and the host refused while the client heard nothing for ten
seconds; `PipeName` now retreats to a short per-uid directory when the council directory's path
would not fit (pipe_unix.go). Still owed: the built binary driven by hand on a Mac through
`--host`, `/detach`, `ls`, rejoin and `kill`, and `p_comm`'s sixteen-byte truncation against a
binary named `telltale` under a real `kill` sweep; until then `PARITY.md`'s row says "suites
measured, hand cycle owed".
#### What this rung deliberately does not do
**It does not put the seats in the host's process group.** The heading above says why, and
`PARITY.md`'s row names the consequence: `council kill` reaps through the session, and a host
signalled by anything else reaps through its handler.
**It does not make the host survive its own SIGKILL.** Nothing on this platform can, and the
sentence that says so is in three places on purpose — `roomjob_unix.go`, `PARITY.md`, and the
test that pins the leak.
**It does not add a field to `host.json`.** The socket path goes where the pipe name went. The
two lock files and the socket node carry no content; `TestATurnIsNotPersistedAnywhere` asserts
the lock files are zero bytes and walks the directory they live in.
**It does not change the protocol or the client.** `Client`, `RunClient`, `Rejoin`, `StartHosted`,
`KillHostedRoom` and `ls`'s words are the same code on every platform; the platform split is
entirely below `Listen`, `Dial`, `ProbePipe`, `verifyHostProcess` and `killProcess`.
### 7.31 the hosted room draws with the room's own columns, and you can leave it from there (2026-09-02)
**What this section rules.** [§7.28](#s7-28) built the host and drew its room with a plain
text client. [§7.29](#s7-29) exposed detach on that client and said, in its own words, that
the TUI had no host to leave. This section makes council's own `Model` and `Render` the
renderer of a hosted room. A room that outlives the terminal now draws the five columns, the
posture badge on each column, the tool trace, the turn page, the panes ([§9.51](#s9-51)) and
the help panel, the same way the single-process room draws them. It also gains `/detach`
inside the TUI. The plain client stays as the degraded path, and the paragraph below says
when it is used.
**What a reader must not read into it.** No golden file of the single-process room moved.
Every hosted frame is pinned by a new golden under `testdata/golden/hosted-*.txt`. No
transcript content reaches disk: `room.json` and `host.json` stay numbers and keys. No new
spawn path exists. The client still reaches the host through `startHostedRoom` and
`joinHostedRoom`, which the spawn guard already wraps.
#### Design point 1: what crosses the wire
Three shapes were possible.
1. **`council.State` on the wire.** The host would own the canonical `State` and the client
would only render it. Refused. `councilhost` cannot import `council`. `boundary_test.go`
pins that direction, and this section is the reason the direction exists. `State` also
carries facts that belong to the client's machine and to the reader: the posture claim
for each vendor, the focus, the scroll, the pane sizes, the help page, the draft. A host
that owned those would own the reader's eyes.
2. **Events on the wire.** The host would forward every `runner.Event`, and the client would
fold them. Refused. A rejoining client needs the whole current room, and the host holds
no event log. A replay from the host would be a second copy of the conversation, which
`resume.go` refuses for this data.
3. **The host's projection, widened.** The host's `Room` is already the fold. It already
travels whole on every frame, so a rejoin already receives the whole current room as its
first frame. This section widens `Room` to carry every fact `Render` reads about a seat,
and package `council` builds a `State` from it with one pure function.
**Ruling: shape 3.** One process owns the conversation, and that process is the host. One
process owns the view, and that process is the client. The function that joins them is
`stateFromRoom` in `internal/council/hosted.go`. It is pure over its two arguments, the
`Room` and the previous `State`. It copies the conversation from the seat and keeps the view
from the previous column with the same vendor: scroll, follow, last focus, the quota reading
the client read from its own relay. `TestHostedStateIsPureOverTheRoom` pins the purity.
What the seat carries now, and did not before:
| field | why `Render` needs it |
|---|---|
| `Turn`, `Prompt`, `Quoted` | the column echo, the band, the turn page, `PageTurns` |
| `History` | the transcript, the turn page, the act ledger |
| `Acts` as `{Text, Status, Detail}` | the trace's outcome marks, which a flat string lost |
| `Started`, `Ended`, `Elapsed` | the header clock, the finished column's figure, the inbox |
| `CostUSD`, `CostSession` | the cost cell and its `session` word |
| `Settling` | `done · exiting` on a batch seat, [§9.33](#s9-33)'s linger made visible |
| `Skipped`, `NoteDetail` | the skip line and the two-line card |
`Acts` was a list of rendered strings. It is structured now, because a string had already
folded the outcome into words and the TUI draws the outcome as a mark. The plain client
renders the same structure through `actLine`, so its output did not change.
**Redaction stays a client choke point.** The host does not redact. The client passes every
body, act and note through `Redact` and `sanitize` in `stateFromRoom`, on the way into
`State`. That is the same rule the single-process room applies in `applyEvents`, applied to a
whole string instead of a stream.
**The frame cost is named.** A frame now carries every seat's history. The 50 ms coalescing
tick bounds the rate, and `maxFrame` bounds the size at 8 MiB. Nothing has measured a room
against that ceiling. The host's memory ceiling is still owed ([§7.28](#s7-28)), and this
section makes the frame grow with it.
#### Design point 2: the crew is the model
[§9.54](#s9-54) made the seats busy one at a time. The host still refused a second dispatch
while any seat was running. That was the committee, and this section removes it from the
host.
- `KindDispatch` carries `Seats`. The client resolves the route (`@codex`, `-@claude`,
`@all`, the default) against its own `State`, and sends the explicit vendor list. An empty
list still means every drivable seat, which is what the plain client sends.
- The host refuses per seat. A busy seat is named on the room's notice and the idle seats
in the same list still go. The refusal reads the seat's phase and its `Settling` flag,
under the room's lock, in the same call that marks the seats it accepts. The old
`watchTurn` poll is gone.
- A drivable seat the dispatch did not name is marked `not addressed in turn N`, with
`Skipped` set, exactly as `sendTurn` marks it.
- `KindInterrupt` carries `Seats`. The client's `ctrl+c` interrupts the focused seat when
it is busy, every seat when the focused seat is idle, and ends the room when nothing is
busy. That is `viewKey`'s own rule, and the footer's cancel cell names which.
- The header's `turn N → route · K in flight`, the inbox strip and the seat numbers all
compute from `State`, so they work over a hosted room without a second implementation.
#### Design point 3: keys, and what a key must not do
**`/detach` is a room verb, not a key.** A single keystroke that walks away from five agents
is a keystroke that can be pressed by accident. The verb is bare-only, on `/read`'s rule: a
sentence that opens with the word is prose.
**In a single-process room `/detach` refuses.** The notice says `this room has no host to
leave. open one with telltale council --host`. A verb that did nothing would be the
honest-gauge failure spent on the room's own surface. The verb sits outside `roomVerbs`,
because the refusal line that lists that table fits its narrowest room with no cell to spare
(`refuseUnknownCommand`); the hosted help panel teaches it instead.
**`q` and `ctrl+c` still END the room.** They send `KindShutdown`, the host kills every seat,
and the closing line says so. Nothing about leaving happens on a key.
**The footer did not change.** Every hint on the mode line is the same in a hosted room. The
two places that say `hosted` are the header, which gains the word `hosted` after the posture
word, and the composer's border label, which gains `hosted pid N`. Both are words, so they
survive `--ascii` and `NO_COLOR`. The help panel's `/cd` row becomes a `/detach` row in a
hosted room, and its `ctrl+c / q` row says that `q` ends every seat. The panel keeps its 16
rows.
**What is refused inside a hosted room, with one notice each.** `/cd`, `/seat`, `/unseat`,
`/read`, `/write`, `/arena`, `/flow`, `/hand`, `/adopt`, `/retry`, `/trace`, `c`, `u`, `x`,
`o`, `a`, `s` and `ctrl+r`. Each of them changes a seat's process, tree, thread or gate, and
the client holds none of those. The rebuttal (`ctrl+r`) is owed: a quoting turn hands each
seat a different prompt, and the dispatch frame carries one.
#### Design point 4: four states, four renders
[§4a.1](#s4a-1) rules that `rebuilt`, `survived`, `died` and `refused` never render alike.
The TUI keeps every sentence [§7.29](#s7-29) wrote, and adds none of its own:
| state | where it renders | the sentence |
|---|---|---|
| detached | printed after the alternate screen closes | `RenderDetached` |
| rejoined | the room's notice line, on the first frame | `RejoinedNotice`, a one-line form of `RenderRejoined` that keeps the clause `nothing was rebuilt, and no session was resumed` |
| the host died while you were away | printed before the ordinary room opens, unchanged | `RenderHostDied` |
| the host died while you were in it | the TUI quits and prints it | `RenderHostExit` |
| refused | the room's notice line | `UnwatchedWriteRefusal`, verbatim, from the host's own `KindRefused` |
The refusal is the host's sentence and never the client's. The client sends `KindDetach` in
a write room too, because [§7.29](#s7-29) enforces the ruling in the host, and the client
draws the reason the host returns. `TestTheFourNoticesNeverRenderAlike` walks the one-line
form beside the others.
#### Design point 5: the live seat is scoped out
`--live` with `--host` is refused before anything opens, with one sentence. A live seat is a
pseudoconsole child, and in a hosted room that child would have to live in the host and its
cell grid would have to cross the wire on every repaint. That is a second wire format and a
second spawn guard, and it is owed. `STATE.md` carries it.
#### Design point 6: the plain client stays
`councilhost.Render` and `RunClient` are unchanged and still used. The TUI needs a terminal.
When `stdin` is not a character device the hosted room falls back to the plain client, so a
scripted drive (`/detach` on a pipe, which is how the Linux measurement in [§7.30](#s7-30)
ran) still works. `plainClient` in `hostcmd.go` is the one place that decides, and it is a var
so a test can force either. The notice tests stay in `councilhost`, because the sentences
live there.
#### Flags a hosted room refuses, and why
`--brief`, `--record`, `--trace` and `--live` are refused with `--host`, one sentence each,
before the room opens. The host takes no brief file, holds no recorder and no trace sink, and
spawns no pseudoconsole. A flag that was accepted and did nothing would be a promise the room
could not keep. `--fresh` is honoured. `--auto` and `--shared-tree` are accepted, because a
hosted room already runs ungated and already runs every seat in the workspace.
#### The host drives the crew's live shapes through their measured adapters
[§9.57](#s9-57) moved codex and grok to request/response live shapes in the registry, and
the host marks a conversational seat undrivable. From that merge until this section a hosted
room could drive only claude and agy, and nothing recorded it. The host now takes the batch
adapter `vendors.LiveFallback` names for such a seat, the way the room retreats to it on a
refused handshake (`fallback.go`), and the seat's badge wears that adapter's measured claim.
The wire carries `FellBack` so the client can draw it. A conversational seat with no fallback,
which is cursor, is still refused in words. The agy seat is not conversational and keeps its
live stream shape, which §9.57 lists as unmeasured.
#### Known limitations, named
- **A hosted room starts every seat fresh.** The host is handed a roster and never a saved
session id, so no `Restored` card is drawn and no thread is resumed. This predates this
section. It is named here so the missing card is read as honest rather than as a defect.
- **The host does not write `room.json`.** [§7.28](#s7-28) said it would, on the room's own
schedule. It does not, and no code in `internal/councilhost` reaches `SaveRoom`. A hosted
room's session ids therefore never reach disk. `STATE.md` carries it.
- **No vendor has been dispatched through a hosted TUI room.** Every seat in both suites is
stubbed. The stubbed end-to-end in `internal/council/hosted_e2e_test.go` re-executes the
test binary as a real host with an empty roster and drives the `Model` over a real pipe.
The owner's second live brief on the built binary is the closer.
## 8. Roadmap (decided 2026-08-01; adoption track added 2026-08-02, ADR-005)
Rigor stays the floor; features and front-end craft are the priority axis from here.
Each item names its incumbent inspiration and the honest-gauge twist that makes it ours.
Sources rule unchanged: a segment ships only when this doc names its source.
Read the version numbers below against §1: v1 cuts on the snapshot gates, so an item
marked for v1 is not an item waiting on the gauges. They are done.
ADR-005 adds a second axis: external adoption is an explicit product goal, and
adoptability is a design input rather than a lagging indicator. That does not reorder the
feature track below — it adds the adoption track that runs beside it, and one of its
items lands *before* v1 is done, which is why this section is no longer titled
"after v1".
### Adoption track (ADR-005)
ADR-001's sequence stands unchanged — dogfood → eval + design doc → launch post — so
these items are ordered by that sequence, not by version number.
1. **Activation slice — runs in parallel with the dogfood window**, i.e. now, not after
v1. Four pieces, all packaging rather than capability: prebuilt binaries via
goreleaser; one-command install with **scoop/winget first**, per Windows-first
(ADR-002); a README hero visual; and a useful **zero-config first frame** — the
binary's first run has to show something true without being configured, because an
install that lands on an empty screen has spent its only attempt. macOS/Linux binaries
are cross-compiled; macOS ships labeled **"smoke-verified on macOS — Windows is the
continuously verified target"** and Linux keeps **"built, not verified"**: ADR-002's
"no macOS/Linux work until v1" is amended for *distribution only*, the
no-porting/no-verification-effort rule stands, and both labels are ADR-001's
flagged-limitation pattern applied to a platform instead of a segment. The macOS
label is point-in-time and SHA-bearing — the suite, build, statusline smokes, a
53-session live Claude Code read and the HUD all ran on macOS at `052a9d6`
(ADR-005, second amendment) — while CI still runs `windows-latest` only, and five
of the six adapters have still never met a live macOS corpus.
The README positioning line — *"one local HUD for every coding agent you use"* —
lands **with** this slice and deliberately not before it: a positioning claim that
arrives ahead of a one-command install is a promise the reader has no way to act on.
2. **The launch post tests one hypothesis: cross-harness visibility** — do multi-harness
power users want one honest local HUD across the agents they already run? That, and
only that, is what the launched product contains. The signal that answers it is
evidence someone actually *ran* telltale — a version-bearing bug report, a real-session
screenshot, a PR grounded in running it, package-manager feedback, or an unsolicited
statement of use. Engagement without that (a comment, a question, a hot take) answers
a different question, and is read as such.
**Amended 2026-08-15: the hypothesis is room-led, because the product is.** The wording
above was written 2026-08-02 and names the HUD, four days before the
council-is-the-product ruling (§1); the post that goes out leads with the room, so the
signal above is read as evidence about cross-harness visibility of the ROOM — one brief
answered by five vendor CLIs side by side, with the gauges as the infrastructure under
it. The signal test itself is unchanged, and so is item 3's exclusion.
3. **Needs-input / blocked / done state is the first post-validation feature** — the
attention-routing job, and the reason the product is positioned the way it is. It is
built where the vendor seams already support it: Claude Code hooks, Codex notify
events, agy's `agent_state` (observed live transitioning `tool_use` → `idle`, §3.8),
and Cursor Hooks (documented and versioned — §3.9, and the reason the Cursor adapter's
`status` field is deferred rather than mapped). It is judged on its own terms rather
than on the launch's: the launch explicitly does not claim this ground, so its result
is evidence about cross-harness visibility only, and neither validates nor falsifies a
capability the launched product did not contain.
4. **The agy disk-seam re-survey RAN the same day this track landed — verdict: OPEN**
(§3.8 re-survey block; prompted by ccusage issue #1402). Both claimed surfaces
verified against the local 1.1.9 corpus: the advertised transcript.jsonl is real
(ADR-004's watch item resolved — the first survey's "never written" was wrong at the
same version), and `gen_metadata` token counts decode with a self-checking
arithmetic identity. **The next adapter work item is therefore the agy HUD adapter
itself** — transcript-first (Name/Workspace/LastActivity/liveness scaffolding from
plain JSONL), `gen_metadata` only for Model + token counts, honoring §3.8's build
cautions (WAL sidecars, stale summary index, PII fields, assert-the-identity). agy
stops being a statusline-only vendor when that ships, and telltale ships the lane's
only Antigravity HUD adapter. Liveness and subagents stay structural-only until
observed live. — **BUILT the same day** (decisions/006,
`internal/adapter/antigravity` + `internal/sqlite`, §3.8 "Adapter built"): four
reported fields, zero new dependencies, the token identity asserted at read time and
holding 16/16 on the live corpus. agy is now a HUD vendor; the two deferrals stand.
5. **Cursor (Composer) is the fifth vendor and the SIXTH HUD lane** — surveyed and
**BUILT the same day** (§3.9, decisions/007, `internal/adapter/cursor`). It is the
first vendor to persist its own context percentage, and the first whose store holds
live credentials, so the adapter's most load-bearing property is its read allowlist.
Two watch items follow it and neither is blocked on effort: **Cursor Hooks**
(cursor.com/docs/hooks — a documented, versioned payload carrying `conversation_id`,
`model`, `workspace_roots`, `transcript_path`, and context numbers on `preCompact`) is
where item 3's needs-input signal should come from for this vendor, rather than
reverse-engineering `status` out of the store; and the `cursor-agent` CLI keeps a
separate store that is not installed on the survey machine and stays unverified and
out of scope until it is. — **The second watch item CLOSED 2026-08-29.**
`cursor-agent` is installed on this machine now, its store was surveyed, and
`internal/adapter/cursor` reads its per-session manifest. §3.9's 2026-08-29 addendum
carries the field map, the `schemaVersion` pin and the composition argument. The
Cursor Hooks watch item is untouched and still open.
#### Packaging decisions (settled 2026-08-08; §6.5 closed here)
Adoption item 1's first piece — prebuilt binaries and a one-command install — is
**built and not yet fired**. `.goreleaser.yaml` and `.github/workflows/release.yml`
exist; no tag does. §6.5 deferred distribution naming "to packaging time", and this is
it, so the rulings are here rather than in the open-questions list.
**Tag day is one command.** `git tag vX.Y.Z && git push origin vX.Y.Z`. A `v*` tag is
the release workflow's only trigger; merging to main releases nothing. The workflow
then runs the repo's own gate, builds, and stages a **draft** release. The runbook
lives in [packaging/README.md](../packaging/README.md) — this section is the argument,
that file is the procedure, and neither restates the other.
1. **Targets, and each one's label.** Four: `windows/amd64`, `darwin/amd64`,
`darwin/arm64`, `linux/amd64`, CGO off. The labels above are binding on the release
notes and are printed per download in the release body: Windows **continuously
verified**, `darwin/amd64` **smoke-verified on Intel macOS** (point-in-time,
`052a9d6` — the Mac that ran it is Intel, and the label says which arch was under
the smoke rather than "macOS" flat), `linux/amd64` **built, not verified**.
`darwin/arm64` is the one addition to the ADR-005 list and it takes the Linux label
verbatim: **built, not verified**. Shipping it was weighed against withholding it —
most Macs are Apple Silicon, so refusing to build it replaces a labelled binary with
no binary at all, which serves nobody and teaches nobody anything. The label is the
whole claim. `windows/arm64` and `linux/arm64` are not built: no verification story
and no known user, and an unlabelled binary nobody has run is the packaging form of
a rendered guess.
2. **Names: `telltale` everywhere.** §6.5 named `telltale-hud` as the fallback if the
bare name collided. It does not — checked at packaging time against the scoop
`Main` and `Extras` buckets and against `microsoft/winget-pkgs`, both clean — so the
fallback stays unused and the scoop app is `telltale`. winget needs a
publisher-qualified id and gets `sanlee-ys.telltale`, the GitHub-handle convention
`junegunn.fzf` and `ajeetdsouza.zoxide` already use. npm remains skipped, not
renamed: the bare name there IS taken (an unrelated option parser), and §6.5 already
ruled npm optional for a Go binary.
3. **The scoop bucket is in this repo, `bucket/`.** Rejected: a second repo,
`sanlee-ys/scoop-telltale`. The deciding cost is a credential rather than
convenience — goreleaser pushing a manifest into a *different* repo needs a
cross-repo PAT held as a release secret, while pushing into its own needs only the
workflow's built-in `GITHUB_TOKEN`, which the release already holds to upload
artifacts. A project whose stated posture is that it reads no credentials should not
mint a long-lived write token for one JSON file. `scoop bucket add` accepts any git
repo carrying a `bucket/` directory, so the user's command is the same length under
either choice, and the second repo buys nothing but a second thing to keep alive.
4. **The release is a DRAFT and stays one.** goreleaser's job ends at "artifacts and
notes are staged"; publishing is outward-facing and is a human action. The honest
consequence, recorded rather than glossed: the scoop manifest is committed in the
same run, so between that commit and the publish click, `scoop install telltale`
points at a URL that 404s. It fails cleanly and installs nothing — but the window is
real, and the answer is to publish promptly rather than to pretend it isn't there.
`skip_upload: auto` keeps snapshots and prereleases out of the bucket entirely.
5. **The release runs the repo's own gate, called and not copied.** `release.yml`'s
first job is `uses: ./.github/workflows/ci.yml` — vet, the suite, the build and the
three binary-level smokes, on `windows-latest`. A release that skips the gate is a
false green; a release running a hand-maintained second transcription of it is a
subtler one, and that is the only change `ci.yml` took (a `workflow_call` trigger).
goreleaser itself runs on `ubuntu-latest`: with CGO off the cross-compile is
host-independent, so the build host is a cost question, and what makes the Windows
artifact trustworthy is the Windows gate that already passed.
6. **Changelog: plain `git log`, ascending, ungrouped** — not goreleaser's default
Conventional-Commits grouping, and explicitly not `use: github-native`. This repo's
commit voice is a lowercase sentence describing the behavior change (CLAUDE.md), not
a `feat(x):` label, so CC grouping would file every commit under "Others" beneath
three empty headings; squash-merge already makes those subjects the PR titles, which
is the changelog anyone would write by hand. `github-native` is rejected for a
harder reason: it replaces the entire release body, which would delete the
platform-label table — the part of the release that has to be true.
7. **Not packaged, and why.** No Homebrew tap: macOS is smoke-verified on Intel only,
and a tap resolves just as happily on Apple Silicon, which would launder "built, not
verified" into "supported". No `.deb`/`.rpm`: a distro package is a support claim
Linux has not earned here, and the tarball carries the label a package would drop.
No npm, per (2). winget is packaged but **not automated** — submission is a pull
request against a Microsoft-owned repository, and a bot opening PRs on someone
else's repo every tag is a different thing from cutting a release. The manifest
draft (schema 1.12.0, `zip` + nested `portable`, `windows_amd64` only) and the
submission flow are in [packaging/winget/](../packaging/winget/).
**Amended 2026-09-02: the tap ships.** What the refusal above was waiting for
was a measurement on Apple Silicon, and `ci.yml`'s `darwin` job is one: it
builds the binary on `macos-latest` (arm64) on every commit and runs the suite,
`doctor`, the statusline smokes and `council ls` there, with the Windows job's
honesty assertions. The tap lives in `Formula/` of this repository, on (3)'s
credential argument for the scoop bucket, and goreleaser rewrites the formula at
each tag. It is a **formula and not a cask**: Homebrew fetches a formula's
archive with curl and writes no `com.apple.quarantine`, so the unsigned binary
runs as installed, while a cask arrives quarantined and goreleaser's cask
documentation answers that with an `xattr` post-install hook it labels a
Gatekeeper bypass. goreleaser v2.17.1 deprecates `brews` in favour of casks, and
`goreleaser check` now exits 2 naming that one key; the pin is exact, deprecated
keys are removed only at a major version, and `.goreleaser.yaml` carries the
measurement. Still not done, each with its reason: **homebrew-core**, whose
acceptance rules ask for a notable project (75 GitHub stars among the signals)
and nothing here measures that number; **winget**, per this item; **signing and
notarization**, per (8). packaging/README.md is the runbook.
**Amended 2026-09-11: winget carries `0.2.0`.** The manifest drafted above was
submitted by hand, as this item required, and merged into
`microsoft/winget-pkgs` on 2026-09-10 as #417671 (`sanlee-ys.telltale`,
the publisher-qualified id from (2)). Nothing about "not automated" moves:
`v0.3.0` was published 2026-09-09 and has no manifest there, because each
version is its own pull request against a Microsoft-owned repository and a
session does not open those. So the winget channel lags the release by
exactly the submissions the owner has made, and the README says which
version it delivers. No `winget install` has been run against the community
source; the first one is a PARITY.md entry, the same debt the tap carries.
8. **Not signed, and the decision is the owner's** (recorded 2026-08-16). No
artifact carries a signature. `.goreleaser.yaml` declares no `signs` block and
`release.yml` holds no signing secret, so the claim is checkable from the two
files rather than asserted. This is a **gap that is stated, not a gap that is
closed**, and it takes the same treatment as an unverified platform: the label
is the whole claim, and it now appears in `SECURITY.md` and in the README
install section. Windows ships without an Authenticode signature, and scoop and
winget deliver that same binary. The macOS archives are unsigned and not
notarized, so Gatekeeper refuses one that a browser marked with
`com.apple.quarantine` — **that path is unmeasured here**, because the macOS
smoke ran a binary built on the Mac itself rather than a downloaded archive,
and the sentence says so wherever it appears. `checksums.txt` stays the
verification this release can honestly offer: it proves the archive is the one
the workflow produced, and it proves nothing about who produced it.
**Why it is not built:** signing needs a code-signing certificate, or an Apple
Developer account and a notarization credential, held by the owner and stored
as long-lived release secrets. The credential is the deciding cost here, the
same way it decided the in-repo scoop bucket in (3) — except that argument
cannot be won by a config choice this time, because no arrangement of the
workflow produces a signature without an owner-held secret. So this is an owner
decision about spend and about identity, not a contributor task, and no
contributor should build the pipeline speculatively.
**Posture surface (added 2026-08-16).** `SECURITY.md` carries the private
reporting route, the trust model in the terms of the read/write boundary, and the
signing statement above. `.github/dependabot.yml` watches `gomod` and
`github-actions` weekly; it is also the watch on the TUI line, because
`ultraviolet` is pinned to a pseudo-version that never moves on its own.
#### The automatable remainder (added 2026-08-16)
The posture paragraph above shipped the two pieces that need no automation. This
subsection records the rest. It changes nothing in item 8: signing and
notarization stay owner decisions, and no contributor builds that pipeline.
**What now runs, and when.**
| Check | Fires on | Runner | Fails on |
|---|---|---|---|
| `ci.yml` | push to main, pull request, release | windows-latest, ubuntu-latest, macos-latest (Apple Silicon, added 2026-09-02) | vet, the suite, the build, the binary smokes, the schema gate, the install-script gate (added 2026-08-18) |
| `govulncheck.yml` | push to main, pull request, Monday 07:00 UTC | windows-latest | a reachable known vulnerability |
| `codeql.yml` | push to main, pull request, Monday 07:30 UTC | ubuntu-latest | a default-suite alert |
| `dependabot.yml` | weekly | none | nothing. It opens a pull request |
| SBOM, through syft | a `v*` tag only | ubuntu-latest | a syft failure |
| Provenance attestation | a `v*` tag only | ubuntu-latest | an attestation failure |
**govulncheck is a monitor, and it is deliberately not a step in the gate.** The
gate answers whether a change works, and the tree decides that answer. This scan
answers whether the shipped code is vulnerable today, and the Go vulnerability
database decides that answer. Two costs follow. The scan reads vuln.go.dev on
each run, and the gate makes no network call except the module download, so an
outage at that host inside the test job would fail a build that nothing broke. A
new standard-library CVE also fails an unchanged tree, and it fails every open
pull request at the same time, and no author can correct it inside their own
change. A gate that fails for a reason outside the change teaches contributors
to ignore it. The scan therefore runs beside the gate under its own name. It
still fails on a finding, because a monitor that reports green over a vulnerable
binary is the false green ADR-001 refuses.
**The scheduled run is the reason this exists.** A scan on a pull request cannot
find a vulnerability that becomes public after the merge. The pull request
trigger scans a new dependency before it lands, and it also proves the job in
the pull request that adds the job.
**What the first scan measured, 2026-08-16.** govulncheck v1.7.0, against a
database updated 2026-08-14, reported three reachable standard-library
vulnerabilities: GO-2026-6090 in `crypto/tls`, GO-2026-6089 in `net/http`, and
GO-2026-5972 in `encoding/asn1`. It reported five more that this code does not
call. The traces reach `internal/grokotel` and `internal/eventsink`, which are
the two packages that run a server. **No module that this repository requires
caused any finding.** All eight findings had one cause: the toolchain. go.mod
declared `go 1.26`, `actions/setup-go` resolved that to the 1.26.5 in the runner
image, and every fix version is go1.26.6. Run 31981243687 on main confirms that
CI built with go1.26.5, so this was a property of the shipped binary and not of
one workstation.
**The fix is a version bump, and go.mod is where the version lives.** The `go`
directive now reads `go 1.26.6`. That clears all eight findings, measured both
ways: the same scan under a 1.26.6 toolchain reports `No vulnerabilities found`,
and the local build then reports `go1.26.6` under `go version -m`. The directive
is the single source every job reads, so one line moved the race job, the gate,
the release and both new scans together. `check-latest: true` on setup-go was
the rejected alternative. It would float the toolchain to whatever patch exists
on run day, which is the same objection this document already made to a
goreleaser version range. A committed directive moves when a person moves it,
and the monitor is what asks for the move.
**The SBOM and the attestation take effect at the next tag, and at no earlier
point.** `release.yml` triggers on a `v*` tag only, so no merge to main produces
either one, and neither is added to a release that already exists. syft writes
one SPDX-JSON document for each archive. `actions/attest-build-provenance` then
attests the four archives, and a user verifies one with `gh attestation verify
--repo sanlee-ys/telltale`.
**Provenance is not a signature, and the difference is item 8's difference.** The
attestation proves that this repository's release workflow built this archive,
from a named commit, on a runner GitHub hosts. It does not prove who the owner
is, and it does not prove that the owner vouches for the content. It needs no
owner secret, because GitHub mints a short-lived OIDC token for each run. That
is exactly why it is buildable here while signing is not: item 8's blocker is a
long-lived credential, and this mechanism needs none. **Signing and notarization
stay owner-held and unbuilt**, on item 8's terms and for item 8's reasons.
**What this subsection did not verify.** The pull request that added these files
proved the two scans by running them, and the run log names each job. It could
not prove the release path, because a tag is that workflow's only trigger. The
evidence for the release path is a `goreleaser release --snapshot --clean`
rehearsal on the reference workstation, 2026-08-16, against the pinned
goreleaser v2.17.1. It built all four targets, wrote all four archives, reached
the SBOM stage, named the document
`telltale__windows_amd64.zip.sbom.json`, called `syft`, and stopped
with `exec: "syft": executable file not found`. That failure is the wanted one:
the stage is reached and configured, it fails loudly rather than passing in
silence, and the missing tool is the one thing `release.yml` installs and the
rehearsal could not. `goreleaser check` also passes against v2.17.1. The
attestation step ran nowhere. **The next tag is the first real proof of both**,
and item 8's existing advice applies: rehearse it with an `-rc` tag.
`.github/CODEOWNERS` names the owner for every path. SECURITY.md already tells a
reporter that one person maintains this project; that file states the same fact
where GitHub can act on it.
What this does **not** discharge: the README hero visual and the zero-config first
frame are the other two pieces of adoption item 1 and are untouched here, and the
positioning line still lands with the slice, not ahead of it.
*(Ledger, 2026-08-16: the zero-config first frame is DELIVERED — measured on a clean
profile first, then built narrow; §7.7's "measured 2026-08-15, then narrowed"
subsection is the record. The positioning line landed room-led with the identity
rewrite of 2026-08-15. Of adoption item 1, the hero visual alone remains, and the
recording chain below is what will produce it.)*
#### The recording chain: PowerSession-rs + agg (measured 2026-08-16)
The demo tape records with **PowerSession-rs 0.1.16 and agg 1.9.0**, both
installed from winget at user scope with no administrator rights
(`Watfaq.PowerSession`, `asciinema.agg`). VHS is rejected rather than deferred:
it cannot record on this Windows build (charmbracelet/vhs #631, dead since
2025-06) and it renders xterm.js, which is not the Windows Terminal this
product targets under ADR-002 — a tape of the wrong terminal is the packaging
form of a rendered guess. The chain was proven against the real binary in
ascending difficulty, and the hard case passed: `telltale hud` recorded its
alternate screen (`ESC[?1049h` and `ESC[?1049l` each captured once), its ANSI
palette colour (47 cyan, 19 green, 7 bright-black foreground codes) and its
restore. The restore was tested against real scrollback rather than an escape
count — a marker line, the TUI, a second marker line — and the rendered final
frame carries both markers and no TUI residue. Two limits are recorded with
it, because both can produce a tape that looks successful and is not.
`NO_COLOR` in the recording shell silently strips every hue, which is how the
first capture came back monochrome. And agg's `--rows` re-runs the byte stream
at a new size, so it recovers `telltale doctor`'s scrolled-off report (76 lines
from a 30-row cast) but does nothing for a TUI that drew to the size it read —
the HUD at `--rows 50` leaves twenty empty rows. The tape's geometry is
therefore chosen before the recording starts, not after.
[packaging/tape/README.md](../packaging/tape/README.md) is the runbook and
carries the full measurement; this paragraph is the decision. **No cast or GIF
enters this repository**: both captures were inspected, and they carry live
session names, workspace paths and the absolute path of every vendor binary on
the machine. The tape stays a personal artifact and the repository holds the
script that makes it. What remains is not a tooling item — the owner drives the
eight beats, because a scripted race would be an invented recording.
#### The one-paste Windows install (added 2026-08-18)
`packaging/install.ps1` is the third Windows route, and it is the only one that
needs nothing installed first:
```
irm https://raw.githubusercontent.com/sanlee-ys/telltale/main/packaging/install.ps1 | iex
```
**It exists because scoop is a prerequisite and winget is not submitted.** Item
1's "one-command install with scoop/winget first" shipped both of those, and
both assume the reader already has the package manager. A reader who has
neither had two choices before this: unpack an archive by hand, or build from
source. A competitor sweep on 2026-08-17 read the same gap the other way round:
every lane leader collapses README-read to first-run into one paste, and abtop's
README carries this PowerShell shape. That is a reading of their documents, not
a measurement of their installers, and it is cited as such.
**What it verifies, and what it refuses to claim.** The script downloads the
archive and `checksums.txt`, compares the SHA-256 **before** it unpacks
anything, and deletes the download on a mismatch. It then prints, in its own
output rather than only in a document nobody reads at install time, that the
binary carries no Authenticode signature and that the checksum proves what the
workflow built and not who built it. That sentence is item 8 restated at the
one moment the reader can act on it. The script signs nothing and prepares no
signing pipeline: item 8 stands unchanged.
**Three refusals are in the script rather than in a note.** A machine reporting
`PROCESSOR_ARCHITECTURE` other than `AMD64` is refused by name, because the
release builds no `windows/arm64` binary and installing the amd64 one there
would be the packaging form of a rendered guess. A tag with no published
release fails with the URL that 404'd, not with a bare status code. A
`checksums.txt` that names no entry for the archive stops the install rather
than skipping the check.
**Measured 2026-08-18, against the published `v0.2.0` release.** Windows 11,
two shells: PowerShell 7.6.5 and Windows PowerShell 5.1.26100.9168. Both
installed `telltale_0.2.0_windows_amd64.zip`, both computed
`7a2401aa…33772528`, and that value equals the digest GitHub reports for the
asset. The installed binary answers `telltale 0.2.0`, so the release ldflags
survive the route. The `irm | iex` shape was exercised as `Get-Content -Raw |
Invoke-Expression`, and the calling shell survived it: the script throws and
never calls `exit`, because `exit` inside a piped script ends the user's
session. The `PATH` branch was driven once with the real user variable and
restored byte for byte afterwards: the install directory reached the persisted
user `PATH` and the running shell's own `$env:Path`. That trial also measured
the one surprise in this route, and it is recorded rather than smoothed over:
the directory is APPENDED, so a `telltale.exe` already earlier on `PATH` — a
`go install` build, in the measured case — goes on winning. `Get-Command
telltale` names the one that runs. Prepending was the rejected alternative,
because a script that quietly outranks a binary the operator put there is doing
something the operator did not ask for. Two refusals ran end to end and
installed nothing: the arm64 refusal, and a `TELLTALE_VERSION=v0.1.0` run
against the tag that has no release.
**What had no end-to-end live trial was the mismatch refusal**, because driving
it needs a host that serves a corrupted archive. Its comparison was measured
live instead: the real `checksums.txt` was parsed, a byte was appended to the
real archive, and the two hashes differed. The branch that acts on that
comparison is three lines below it and was unexercised until the payment below.
**Paid 2026-08-19, on the MBP (Intel, macOS, PowerShell 7.6.4): two mismatch
trials and one control.** The missing instrument was the corrupt host, so the
trial built one. A `python3 -m http.server` on 127.0.0.1 served the published
`v0.2.0` `checksums.txt` beside a copy of the real archive with one byte
appended (`7a2401aa…33772528` became `8ea9111b…6eb47b8d3c`). The script under
test was the shipped file with one recorded change: the `$base` line pointed at
the local host. The OS and arch refusals read plain environment variables, so
`OS=Windows_NT` and `PROCESSOR_ARCHITECTURE=AMD64` let the real function run on
this machine; that shortcut is named below. Both trials refused, 2/2. The
thrown message named both hashes and said the download was deleted and nothing
was installed. No install directory was created, and no `telltale-install-*`
work directory survived. The message itself is the proof that the three lines
ran: it is the `throw` line's own text, and it carries the two hashes the lines
above it computed. The control reversed the one appended byte and changed
nothing else. The same server and the same patched script then installed
cleanly: `sha256 ok (7a2401aa…)`, and a placed `telltale.exe`. Stated honestly:
this trial redirected `$base` and satisfied the two host guards from the
environment, so it measured the mismatch branch and not the GitHub download
path or the Windows host — the 2026-08-18 block above already paid those two
on the real host. STATE.md no longer carries the gap.
**One footgun is recorded because it cost a parse, and the gate holds it.** The
file is ASCII only. Windows PowerShell 5.1 reads a BOM-less file as ANSI, so one
em dash inside a `throw` produced four parser errors under 5.1 and none under
PowerShell 7. The `irm | iex` path decodes UTF-8 correctly and would have hidden
this; the download-then-run path would not. `ci.yml` now parses the file under
Windows PowerShell 5.1 on every push, rejects any byte at or above 0x80, and
rejects an `exit` statement. It never executes the script, because executing it
would download a release on every push. The ASCII arm was measured non-vacuous
the way the schema gate's mutations are: one em dash appended to a copy, and the
gate reported three bytes and failed.
**No other channel is reshaped by this.** No Homebrew tap, no npm, no winget
automation. Items 2 and 7 rule each one, and a one-paste installer is not an
argument to revisit any of them. (The tap arrived on 2026-09-02 by item 7's own
amendment, on a measurement rather than on this installer.) macOS and Linux keep the measured `curl` and
`shasum` walk in the README, which is the same verification without a script.
#### The listing and launch cadence, recorded and not executed (added 2026-08-18)
This subsection records strategy that no contributor may execute. It exists
because this repository rejects unrecorded strategy, and because the pieces
below are owner actions on surfaces outside it.
1. **Directory listings.** `awesome-claude-code` and the neighbouring lists are
the lane's standing distribution channel, and an inclusion is a pull request
to somebody else's repository. That is the same class of act as the winget
submission in item 7, and it takes the same ruling: **a human action, never
automated, and never opened by a contributor session.** What lands in this
repository is the badge slot in `README.md` and this paragraph. The badge
goes in only after the listing merges.
2. **One Show HN per versioned feature, with the maintainer working the
thread.** Recorded as the cadence, with one binding limit: item 2 pins the
launch to ONE hypothesis, cross-harness visibility of the room, and a serial
cadence must not quietly widen that claim. A second post about a second
feature tests that feature's own question and is read as such. The first post
is chain link 3 in `STATE.md`, and it is sequenced behind links 1 and 2,
which are paid.
3. **Publish the run-evidence bar, the method, and the result.** Item 2 already
defines the signal that answers the launch hypothesis: a version-bearing bug
report, a real-session screenshot, a pull request grounded in running it,
package-manager feedback, or an unsolicited statement of use. Nobody in this
lane publishes that bar. Publishing it is the launch story only an
honest-gauge product can tell, and it costs nothing to tell, because the bar
is written down already.
**The threshold is an owner decision and is NOT taken here.** The candidate
sweep proposed "10 runs in 30 days". That number has no measurement behind
it and no ruling, so adopting it would be the invented figure ADR-001
refuses. What is settled is the KIND of evidence, quoted above. What the
owner names before the post: the count, the window, and what the result
means if it is missed. Whatever is published then cites measured evidence,
and it never cites a star count or an install count, because telltale
measures neither.
**`README.md` carries two slots for this work and no adoption content.** The
badge slot holds one badge, the CI result, which is GitHub rendering GitHub's
own run and therefore needs no third-party host and cannot go stale. A
directory-inclusion badge may join it after that listing merges. A star count,
a download count, an install count or a "used by" figure never may: telltale
measures none of them, and a third-party render of an unmeasured number is the
badge form of a rendered guess. The hero slot is for the animated capture,
which stays owner-driven under the recording chain above and under the
2026-08-17 per-frame review ruling. The still SVG hero and the positioning line
already landed; neither moves here.
Neither track discharges what verification already owes: §3.4's remaining passive-tail
items stay open (§3.7's first live Gemini pass ran and passed 2026-08-03), and
adoption work does not buy an exemption from them.
### v1.1 — the flagship trio (**BUILT**)
1. **Detail pane** (inspiration: abtop / CASS drill-ins) — **BUILT, §7.11.** Select a row,
get an expanded view: quota windows, extras (branch, CLI version, ctx tokens), session
id, and — crucially — the session's Diagnostics and degraded-field marks, which v1
carried with no surface. The honesty machinery becomes visible product. Shipped with
one thing the spec did not ask for and the design demanded: a `not sourced` line, which
makes §4a.1's "can't know" versus "absent now" legible for the first time.
2. **Burn-rate forecast** (inspiration: Claude-Code-Usage-Monitor / codeburn) — **BUILT,
§7.12.** The HUD samples the account window's `used_percentage` over its own runtime;
the slope is a telltale-measured value, rendered with the `~` derived marker AND its
sampling window (`~13:27 · 18m basis`). Never extrapolated from a guessed budget, and
below the minimum basis it renders nothing at all. Shipped with two refusals beyond the
spec's "minimum sample count/age": no projection past the window's own reset, and none
beyond 24 h, because the render is a wall clock with no date on it.
3. **Sub-agent chips** (inspiration: claude-hud's active-subagent display) — **BUILT,
§7.13.** Counts recently-written transcripts in a session's `subagents/` sidecar tree
(already discovered and excluded from rows in §3.1): a `⑂~2` chip on rows running
fan-outs. Corrected against the spec: the roadmap called it "pure sourced data" and it
is not — the files are counted exactly, but the recency boundary is an inference, so
the field is `CapDerived` and the chip carries the estimate marker.
Also in v1.1: **`/` type-to-filter** on name/path substring (CASS's kernel without the
embedding search) — **BUILT, §7.14.**
New schema field in v1.1: `subagents` (§4a.2), declared `CapDerived` by the Claude adapter
and `CapNone` by Codex. `model.Validate` gained a non-negative check for it; nothing else
in the schema moved.
### v1.2 — the Windows-native leap
- **`telltale notify`**: a third mode on the same binary, fed by Claude Code hook events
on stdin, raising a Windows toast when a session needs input or ends a long turn.
(agenttray had the idea; nobody has executed it well on Windows.) Read-only posture
holds: notify consumes hook payloads, sends nothing anywhere but the OS notification
API.
- **Statusline context-breakdown bar** (two-line statusline): stacked mini-bar of
`current_usage` components (input / cache-read / cache-creation / output) — fully
sourced from stdin.
### Later / unscheduled
- Themes + segment config file (ccstatusline's adoption driver).
- `telltale snap`: one-shot frame render to stdout (pipeable, screenshot-able; already
prototyped as a throwaway during v1 verification).
- ~~Gemini CLI adapter, once its seam is verified live (§4a.7 example becomes real)~~ —
**landed 2026-08-02** (`internal/adapter/gemini`, §3.7; the §4a.7 example is real,
with its original sketch kept as the postscript's evidence).
### Deliberately rejected
- Cross-device pairing/sync (codeburn): network egress breaks "telltale never writes".
- Plan-budget "% of plan" spend meters: the budget is a guess — the exact fabrication
this product exists to refuse.
- On-disk cost estimation via price tables: inventing dollars from token counts.
- A plugin runtime for third-party adapters: it would run a stranger's code inside the
process whose whole contract is "reads, never writes, no network, no credentials", so
every guarantee here would become a guarantee about someone else's plugin. The
drop-file relay ([§7.23](#s7-23)) is the middle path taken instead.
## 9. Council (ADR-008)
`telltale council` is the dispatch room: one brief typed once, routed to Claude's control
plane by default or to the explicitly `@mentioned` seats, replies streaming side by side.
`@all` convenes every seated vendor when independent answers are the point. It exists
because the alternative is four terminals and a clipboard.
[council.md](council.md) is the room's user-facing guide — the badge vocabulary, the
routing grammar, the reading keys, every flag. This section is the record underneath it:
what was measured on each vendor, what each finding cost, and what is still unverified.
It is the one subcommand that is not a gauge, and the boundary is worth stating precisely
rather than hand-waving. §7.8's invariant — no keybinding may mutate vendor state or send
anything to a running agent — is **unchanged and unweakened**; it is a rule about the
observation surfaces, and council is not one of them. Nothing in the HUD reaches council.
The only way in is typing the subcommand. What moved is the *scope* of the sentence in
`README.md`, from "telltale never writes" to "the gauges never write", because the old
phrasing had become false the moment ADR-008 was accepted and an accepted decision the
docs contradict is worse than either option on its own.
### 9.1 What v1 seats, and what it does not
Five columns: **Claude Code**, **Codex**, **Antigravity**, **Cursor**, **Grok** (the
fifth seat arrived 2026-08-09, §9.39 — this line said "Four columns" for six days after
it landed, which is its own small lesson about headline counts in a doc whose sections
are the record).
Cursor was originally written up here as deliberately absent, on the grounds that
`cursor-agent` was not installed. **That was not a judgement call, it was untrue** — the
binary was at `%LOCALAPPDATA%\cursor-agent\cursor-agent.cmd` and had been for a month, and a
test pinned the false claim in place so the day the world changed underneath it, it went on
passing. It is seated now (ADR-008, fifth amendment). What remains true from the original
paragraph: the `cursor` binary on PATH is the editor launcher, and council never drives it.
The seat is driven on Windows too — the `.cmd` shim turned out to be a nine-line wrapper
around a bundled `node.exe`, which detection resolves directly, so no prompt ever meets a
shell and §9.3's refusal has nothing left to fire on (ADR-008, tenth amendment). Its posture
badge stays `ro:requested`, now contradicted-by-capture rather than merely unobserved: under
`--mode plan` the agent was seen dispatching shell tool calls, and `--sandbox enabled` kills
the turn outright on Windows, so it is not passed there.
This is the inverse of the HUD, where Cursor *is* a built-in adapter (ADR-007) because its seam
is on disk rather than behind a CLI.
#### The vendors this room does not seat, and the class of evidence behind each (2026-08-15)
This list existed only in the owner's private notes, which meant it bound nothing and was one
lost file away from being re-derived vendor by vendor. It is here so a future session can tell
**"we looked and said no"** from **"nobody has looked"** — two states this project has already
ruled must never render alike (§4a.1), applied to a decision rather than to a value.
Every entry names the CLASS of its evidence, because they are not the same strength. An
unmeasured rejection stated as a measured one is the honest-gauge defect committed in prose,
and this doc's own §9.2 is the record of what that costs: the first draft of ADR-008 claimed
enforcement for four seats when one had a mechanism.
- **Warp** — **REJECTED, measured.** There is no blocking hook anywhere in the product: no
surface at which a tool call is held while something else decides. So the fleet's guard
requirement cannot be satisfied here, and that requirement is not council's own — council
gates the seats it spawns (§9.8), while agent-ops ADR-012 requires each vendor to be
guard-wired mechanically, per vendor. A vendor with nothing to wire cannot be brought inside
it, and routing work away from an unwired vendor is explicitly not how that gap closes.
- **Amazon Q** — **REJECTED on vendor status**, which is a weaker class than a measurement and
is named as one: the product is dead / end-of-life. There is nothing left here to measure.
- **Amp** — **REJECTED on platform.** WSL-only on Windows, and its threads live server-side.
Windows is this product's primary target (ADR-002), so a WSL-only binary is not a seat here;
and a vendor whose conversation state sits on someone else's server gives §9.4's native
resume nothing on this machine to resume from.
- **Goose** — **REJECTED on platform.** No Windows path at all. The same ADR-002 reason as
Amp, without the second problem.
- **Pi** (`earendil-works/pi`) — **the named instance of the re-host class below**, recorded
by name because its size makes it the one people ask about. A Pi seat re-hosts model
families already seated here. The HUD adapter is a different question with a different
answer: the format is surveyed (§3.9b), live-verified 2026-08-16, and
`internal/adapter/pi` observes sessions. Council still does not seat it.
- **Every BYO-harness re-host** — **REJECTED as a class**, rather than one vendor at a time.
Each of them re-hosts a model family that already has a seat here. §9.4's whole argument for
turn 1 being blind is that the answers are *independent*; a harness wrapped around a model
already in the room buys a column that agrees with an existing one for the reason a copy
agrees with its original. That is a surface, not an opinion, and the room is priced per seat
(§9.21).
- **Kimi Code 0.34.0** — **INSTALLED, WAITLISTED, UNVERIFIED. This is not a rejection**, and
it must not be read as one. The binary is on the reference box; its hook surface and its
session surface have never been observed, so there is no measurement to seat it on and none
to refuse it on either. One trap is recorded because it is the obvious wrong shortcut: the
legacy `kimi-cli` documentation is **VOID** for this binary — it describes a different
program — and any claim read out of it would be exactly the "read off `--help`" evidence
§9.2 refuses.
- **Qwen Code** — **RUNNER-UP: measured, and it fails on one specific thing.** It has a real
statusline hook, which is more than several seated vendors offer. What its payload does not
carry is any rate-limit window, so the quota relay (§7.15) would have nothing to write and
the header could speak for this vendor only by inventing the numbers it renders. Under
§4a.1 that is `CapNone` rather than a plausible fill, which leaves a seat whose gauge is
permanently blank. This is the one entry worth re-checking when a vendor payload changes;
the rejections above do not move until the vendor does.
### 9.2 Two claims the room refuses to leave implicit
Every column header carries its own **sandbox badge** and its own **streaming
granularity**, because vendors across the 4-vendor fleet differ on both and the first draft of ADR-008 got
this wrong in the direction that matters — it claimed "enforced read-only sandboxing" for
all seated vendors when only one had a mechanism named.
| | mechanism | badge |
|---|---|---|
| Claude Code | `--disallowedTools ` + `--strict-mcp-config` | `ro:tools` |
| Codex (macOS/Linux) | `-s read-only`, enforced by the OS sandbox | `ro:enforced` |
| Codex (Windows) | `-s read-only`, enforced since codex-cli 0.149.1 — before that, **no sandbox**: `-s danger-full-access` was the only mode that could spawn a process there (the 2026-08-29 amendment below) | `ro:enforced` |
| Antigravity | **no restriction flag at all**: `--mode plan --sandbox` were measured not to restrict writes, and were dropped once their only observed effect turned out to be a dead turn (§9.6b). Since the 2026-09-03 amendment council passes `--add-dir ` in every posture and `--mode accept-edits` in the write postures, because the cwd is no workspace to this vendor and a write inside a named workspace is auto-denied without the edit grant (§9.59) | `unsandboxed` |
There is no level that renders as an unqualified "read-only", and after the live spike there is
one that renders as the opposite. Antigravity was asked to write a file under both of its
read-only flags and wrote it — file confirmed on disk, reported permission mode and tool list
byte-identical to a run without the flags. That is refuted, not unverified, so it gets a fourth
level badged `unsandboxed`. Deliberately not `ro:none`: every other badge opens with `ro:`, a
reader scanning column headers takes in the prefix before the qualifier, and a vendor that
can edit your working tree must not read as read-only at a glance.
Council **stopped passing those two flags** on 2026-08-04 (ADR-008, seventeenth amendment).
The badge is unchanged and so is every word behind it — it never rested on the flags being
sent, it rested on the write having landed — and what moved is one clause of the detail, which
used to end *"the flags are still passed; they do not restrict it"* and could not go on saying
so. This is the second seat where a posture flag came off because it was doing nothing useful
and one measured harm; the Codex row above is the first, and the two were decided on the same
ledger. §9.6b carries the argument.
**Codex is the seat where the OS changes the answer, and until 2026-08-29 it wore that same
`unsandboxed` badge on Windows.** Re-probed 2026-08-04 against codex-cli 0.146.0: `-s read-only`
*and* `-s workspace-write` both failed every process spawn there with `CreateProcessAsUserW
failed: 5 (Access is denied.)`, including a control asked merely to list a directory. So
"read-only" on Windows was not a read/write distinction — it was a seat that could not read,
which is exactly how this surfaced: a live council turn answered a "thoughts on this repo"
brief with *"I could not inspect the repository."* Council passed `-s danger-full-access` on
Windows in **both** postures for as long as that held, because it was the only mode that ran,
and the badge told the truth about that rather than keeping a comfortable word. The containment
is the workspace, not the flag (ADR-008, third and twelfth amendments).
**Amended 2026-08-29 — the Windows sandbox re-measured at codex-cli 0.149.1, and the read
posture earned its badge back.** A chip re-probed the twelfth amendment's finding against the
build now installed, in throwaway directories, one turn per probe, files checked on disk rather
than read out of the reply:
- **`-s read-only` enforces.** The pwsh spawn still fails with the same `CreateProcessAsUserW
failed: 5` line — but only pwsh: the model retried through `C:\WINDOWS\system32\cmd.exe`,
which spawned *inside* the sandbox and obeyed it. A read turn listed the directory and read a
marker file, exit 0. A write turn (`cmd /c echo probe> wrote-ro2.txt`) came back
`Access is denied.`, exit 1, no file on disk.
- **`-s workspace-write` enforces and contains.** A write inside the workspace landed; a write
outside it (and outside the temp roots the mode allows by design) was denied with no file on
disk.
- **The `.git` carve-out holds on Windows and CANNOT be bought back.** The
`sandbox_workspace_write.writable_roots=["/.git"]` override that §9.6's macOS measurement
showed unlocking `.git` was passed in both the forward-slash spelling `gitWritableOverride`
emits and a backslash spelling — the `.git` write stayed denied in both, while the same
override named an ordinary outside directory and unlocked it. The deny outranks the override
at this build.
- **The resume override enforces.** A turn resumed with `-c sandbox_mode="read-only"` had its
shell write denied, no file on disk — the effect §9.7 recorded as owed and unobservable is
now observed.
So the postures split on Windows instead of collapsing: **read passes `-s read-only` and is
badged `ro:enforced`** — with the stated residual that the pwsh spawn failure is routed around
by the model's own retry, so a read turn can still fail to inspect when the model stops at the
spawn error (a liveness caveat, not a safety one; its cause is unmeasured) — while **write
keeps `-s danger-full-access`**, because a workspace-write seat that cannot write `.git` is the
exact defect the `.git` widening exists to prevent: it edits files all session and never lands
one. `vendors/codex.go` carries the capture on its constants; `TestCodexPostureIsPerOS` pins
both halves of the split, and `TestNoVendorClaimsUnverifiedEnforcement` now requires the
Windows claim to cite the build it was measured on. macOS and Linux are untouched.
**Amended 2026-09-01 — the installed build moved to codex-cli 0.151.0, and the sandbox claims
above were NOT re-measured on it.** A chip traced the seat's `failed (exit 1)` ([§9.58](#s9-58))
and re-ran the three seat argv shapes at 0.151.0 from a scratch directory. All three parsed and
produced `thread.started`, so the flags this section rests on are still accepted. The four
sandbox probes could not run: every turn that day died at the account's usage limit before a
tool was asked for. So the table above is pinned at 0.149.1 on a machine that runs 0.151.0.
[STATE.md](../STATE.md) carries the owed re-measurement.
`TestSandboxBadgesAreNeverBlanket` fails the build if a bare claim reappears, and asserts the
badges stay distinct — convergence on one string is how a per-vendor claim quietly becomes
a blanket one again.
What these badges *say* has not changed since; how loudly they say it has. §9.11 gives the
two that mean "this seat can change your files" weight and the warning hue, because the room
had been drawing `unsandboxed` at exactly the volume it drew `ro:tools` beside it. The words
still carry the whole distinction — that is why they break the `ro:` prefix — so nothing
here depends on colour; the weight only makes the word findable in a frame with four columns
of prose in it.
**The Claude row cost three attempts to get right, and the failure mode is worth recording.**
The original ADR claimed enforcement with no mechanism named. The first correction named
`--allowedTools "Read,Glob,Grep"`, which *sounds* exactly right and is not: it pre-approves
tools for permission prompts, it does not remove them from the session. Running the real
invocation and reading the `system/init` event's own `tools` array showed `Edit`, `Write` and
`Bash` still there. Every test in the package passed at that point, because every test asserted
the **flag** and none asserted the **effect** — which is this repo's own False Green failure,
committed inside the feature whose entire premise is refusing it.
What works is `--disallowedTools` plus `--strict-mcp-config`, and two parts of that are easy to
miss. Deny **PowerShell**, not just Bash — denying only Bash leaves a working shell on the
platform this product targets. And drop MCP servers, because without `--strict-mcp-config` the
session inherits whatever the user has connected; the verification run surfaced Gmail write
tools in a session with every built-in write tool denied, and no fixed deny list can name those
in advance.
The residual limitation is stated in the badge's own detail text rather than hidden: a deny list
cannot cover a tool that does not exist yet. The claim is *these named tools are absent,
verified*, not *this session cannot write*. The general rule this leaves behind: **a flag's name
is not evidence of its effect**, and the check that matters is what the session reports about
itself afterwards.
**Amended 2026-08-17 — the seat's own `capabilities` array, measured and deliberately not
read.** The rule above ends on *what the session reports about itself afterwards*, and Claude
Code's `system/init` carries a field that looks like exactly that: a `capabilities` array. It is
not the same thing. A tool list reports what the session **holds**, which is a fact about the
session. A capability token reports what the vendor **can do**, which is a claim about behavior.
`claude.go`'s Interrupt precedent already refused to rest on one, and this is the case that
precedent was written for.
Measured at **Claude Code 2.1.233** on Windows, from two live headless runs in a throwaway
directory: a plain `-p --output-format stream-json --verbose` invocation, and the read posture's
exact argv. Both init frames carried the same array, in the same order:
```
["interrupt_receipt_v1","interrupt_cancel_queued_v1","msg_lifecycle_v1"]
```
Three findings, and none of them supports a gate. The array does not move with our flags, so it
says nothing about the posture council asked for. The tokens do not date a build —
`interrupt_cancel_queued_v1` is already in the 2.1.226 bundle, so a version check resting on
their presence passes on three installed versions at once. And the one token council might
plausibly want, `interrupt_receipt_v1`, guards a capability a live run already verified, so
reading it can only weaken a claim that already stands.
`streamLine.Capabilities` therefore parses the field, and nothing branches on it.
`TestInitCapabilitiesAreParsedAndGateNothing` pins the three states apart — absent, `[]`, and the
measured array — and pins that every one of them still produces the same `KindSession` event.
Modelling a field nothing reads needs a reason (§7.16b). The reason here is that a modelled field
is a checkable record of a measurement, and a comment is not.
**`telltale doctor` cannot carry this line, and its charter is the reason.** The package doc is
explicit that doctor does not start a turn, because a turn costs real quota (ADR-008 §6).
`capabilities` arrives only on a `system/init`, and a `system/init` arrives only when a turn
starts. So no cheap local read of this field exists, and the check doctor could offer must spend
exactly what doctor exists not to spend. Recorded here instead of built.
Granularity is the same discipline applied to streaming, and the spike made the answer worse
than the guess. Claude streams token-level deltas, verified live. The other two were
provisionally labelled `events`, on the reasoning that a coarse stream is still a stream;
neither streams at all. Codex emits one `item.completed` per complete agent message and has no
message-delta feature even under development. Antigravity delivers an entire response as a
single `text_delta` — a one-word reply left its column blank for 73 seconds and then painted at
once. Both are `GranFinalOnly`. A vendor that emits nothing until it finishes renders `PhaseWaiting`
— a first-class phase, named as such in its column header — rather than an empty column that
looks like slow streaming. `TestWaitingIsNotStreaming` asserts the two never render alike. That
distinction was added on the theory that some vendor might not stream; it turns out to describe
two thirds of the room. This is §4a.1's rule (a dropped column and an em dash must not read the
same) applied to a surface where the ambiguity would otherwise be invisible.
The card's WORDING is not what carries it, and §9.14 is why that matters: the body used to
recite *"this vendor reports no incremental output, so nothing appears until the turn
finishes"* on every waiting turn, which is council's plumbing described in council's vocabulary
in the space a user came to read an answer. The distinction is carried by the header's own
phase word and the granularity badge beside it; the body says only `working — the reply
arrives whole.`
### 9.3 Execution: argv, never a shell
Prompts are arbitrary text — quotes, ampersands, whatever was typed — so no prompt is ever
interpolated into a command string. Specs are `{Binary, Args []string, StdinPrompt, Dir}`
through `exec.CommandContext`.
This is load-bearing on Windows specifically. `LookPath("codex")` resolves to `codex.cmd`,
an npm shim, and Go's `os/exec` runs `.cmd` and `.bat` through `cmd.exe`, whose argument
parsing cannot be safely quoted for arbitrary text. So `detect.go` classifies every
resolved path as `KindNative` or `KindShim`, Codex and Claude take their prompt on **stdin**
(Codex via its verified `-` sentinel), and a vendor that is *both* a shim and argv-only is
marked `AvailUnusable` and not driven at all. The refusal is the feature; the card tells the
user which env override fixes it.
### 9.4 Multi-turn is native resume, not transcript re-send
Turn 1 is blind: no vendor sees another's answer, which is what makes opinions across the 4-vendor fleet
independent rather than anchored. Later turns ride each vendor's own session-resume
(`claude --resume`, `codex exec resume`, `agy --conversation`). Re-sending the transcript
would grow input quadratically against metered quotas and flatten native turn structure into
quoted prose; resume sends only the new turn and makes the blind-round guarantee
*structural*, since each session holds only its own history. Cross-agent rebuttal is an
explicit opt-in toggle that quotes the previous turn's finals as labelled untrusted material.
### 9.5 Layout and testing
Same contract as §7.9, for the same reason: `Render` is pure over `State`, tests construct
state by hand, goldens live in `internal/council/testdata/golden/*.txt` and render with
`PlainStyles()`. Three columns at ≥96 cells; below that, or when a column would fall under
24 cells, the tier drops to a tab bar rather than shredding prose into unreadable ribbons.
Width is measured with `lipgloss.Width`, never `len()`.
Two of the frame's rows are no longer constants, and the ordering that makes that safe is
worth stating: the **tier is settled before any row is budgeted**. The tab bar costs a row,
and the fallback from columns to tabs is a width test — budgeting first and dropping the
tier afterwards worked only while the footer was a fixed three rows, and a taller composer
would have overflowed the terminal by exactly the tab bar. `resolveLayoutIn` therefore
finishes deciding the tier, then spends rows: header, footer chrome, tab bar if any,
collapsed-seat notice if any, then the composer up to its ceiling, and **the composer
yields before the body does** — at the minimum height a six-row draft would leave the
columns nothing, and a room you can type in but not read is not the trade anyone asked
for. A tab bar holding a single tab is not drawn at all: it selects nothing and repeats
the column header underneath it, which stopped being a rarity the moment dead seats began
folding away (§9.9).
One trap worth recording, because the golden tests could not have caught it: `padRight`
truncates rune by rune, so on text that already carries ANSI escapes it cuts through an
escape sequence and counts escape bytes as content. Goldens render with the identity style
set by design, so they are blind to it. Anywhere a line is assembled from differently-styled
pieces — the tab bar, the help body — padding goes through `fit`, which is ANSI-aware.
`TestFitIsANSIAware` is the regression guard.
### 9.6 Invocation traps, one per vendor
Each adapter hit a failure that is silent rather than loud, which is the kind worth writing down.
- **Claude**: `--allowedTools` pre-approves, it does not restrict (§9.2). Enforcement is
`--disallowedTools` + `--strict-mcp-config`.
- **Codex**: `codex exec` and `codex exec resume` **do not take the same flags**. `-s` and `--cd`
are rejected by `resume` with an argument-parsing error, and a parse error means *empty
stdout* — a naive resume would blank the column on every follow-up turn with no card able to
explain it. Resume carries the posture as `-c sandbox_mode=""`, derived from the same
function as the spawn path so the two cannot drift, and takes its workspace from `Spec.Dir`
alone. The session id is **positional**, not a flag value.
- **Antigravity**: `-p` is a **string flag whose value is the prompt**, not a boolean. Written in
the natural order, `agy -p --output-format stream-json ""` exits 0 and cheerfully
answers a question about the flag it just swallowed. `-p` must be last, brief immediately
after, every other flag before it. agy also rejects a prompt on stdin, so its brief goes in
argv and is bounded by the ~32K Windows command-line limit — a real ceiling on a long brief,
with no workaround short of upstream support.
- **Cursor**: the prompt is a variadic positional and needs a bare `--` in front of it, or a brief
that happens to open with `-` is read as an unknown option and the turn dies. And the stream
sends every passage twice — deltas, then the whole message again — which §9.6c covers, because
it is a parsing trap rather than an invocation one and it took two captures to state correctly.
The shared shape: all four failures produce a *plausible* result rather than an error. That is
why each one is pinned by a test asserting the argv this repo actually builds.
### 9.6a The activity trace carries outcomes — and says when it cannot
The trace answers *what did this agent do*. Until this landed it could not answer *did it
work*, which made it the same half-built gauge §4a.1 exists to forbid: `⚙ Bash: go test ./...`
renders identically whether the suite passed or the build never compiled. The results were not
missing, either — they were arriving in the same stream the commands came from and being
dropped on the floor. A room that discards knowledge it has is the mirror image of one that
invents knowledge it lacks, and both are the same failure.
Four statuses, and **Unknown is the one that earns the type**. Pending renders as the bare
entry, OK as `✓`, Failed as `✗` plus the vendor's own first line about why, Unknown as `?`.
ASCII gets `+`, `x`, `?` — chosen around everything already spoken for, since `*` is the
activity prefix, `>` the ellipsis, `]` focus and `#` the HUD's gauge fill. Every distinction is
a glyph before it is a colour, so all four survive `--ascii` and a monochrome terminal.
**Where the outcome comes from, per vendor, and how strong the claim is.**
| | signal | verified |
|---|---|---|
| Claude Code | `user` messages carrying `tool_result` blocks with `tool_use_id` + `is_error` | **live**, 2026-08-04, Claude Code 2.1.220 |
| Codex | `item.started` → pending; `item.completed` with `exit_code` / `status` | captured fixtures; `exit_code` and `status:"failed"` observed, `status:"completed"` **never** |
| Antigravity | `step_update` ACTIVE → pending, DONE → **Unknown** | live capture shows DONE carries no success signal at all |
Three things the Claude probe settled that a docs-first parser would have got wrong. Field
**order** differs between captured lines, so nothing may be read from position. `is_error` is
**absent** on some successes — a `Read` result carried only `tool_use_id`, `type` and `content`
— so absence is success rather than unknown; Claude Code marks failure and stays quiet about
the rest. And the results came back **out of order**, the second call's failure landing ahead
of the first call's success, on the very first probe. That last one is why correlation is by
id and never by arrival order: a trace zipped by position would have blamed the wrong command
on its first real run.
Antigravity is the case the Unknown status exists for. Its steps flip ACTIVE then DONE, and
every captured DONE line carries `duration_seconds`, sometimes a `tool_info` with the call's
parameters, and nothing whatsoever about whether the step achieved anything. agy reports
success or failure exactly once per turn, in the final `result` event, and that verdict is
about the *turn*. So a finished agy step renders `?` — not `✓`, and the code comment says why.
Reusing the success mark would be council inventing a result on a vendor's behalf, which is the
`--allowedTools` mistake (§9.2) wearing different clothes.
Codex carries the same discipline in a smaller way. `exit_code` is a **pointer**, because codex
spells "still running" as `"exit_code":null` and a plain int would flatten that to 0 — the
spelling of success, and the most expensive confusion available on that field. And an item that
completes with neither an exit code nor `status:"failed"` resolves **Unknown**, not OK: no
captured line has ever carried `status:"completed"`, and guessing the success spelling from the
observed failure one would be a success claim built on a string nobody has seen. That
deliberately weak mapping tightens the moment a live run shows the spelling.
`TestActOutcomesRenderDistinctly` fails the build if any two statuses ever render alike;
`TestOverlappingToolCallsResolveToTheRightEntries` replays the real out-of-order probe.
**A failed entry is a card now, and it was the one card §9.11 missed.** That pass gave every
card in a column one grammar — a title with its body hanging under it — and cited the trace as
somewhere the room *already* did it, on the strength of the failure detail's indent. The entry
itself never got it, and a live room showed why that matters: `run_command: pwsh -Command
"Get-ChildItem"` does not fit 37 cells, so the command wrapped to a continuation starting hard
against the column edge, reading as a second nameless entry with the outcome mark stranded on
it. It now hangs under its own `⚙`, which costs no rows and makes one call look like one call.
Two things then had to change with it, and both are the kind of detail that only shows up
against a real capture:
- **The reason indents FOUR, not two.** Once the command hangs at two, a reason at two lands in
the same column as the tail of the command it explains — telling them apart by colour alone,
which this product does not do. Goldens render with `PlainStyles`, so that golden is exactly
the artifact that proves it.
- **The reason is flattened and bounded.** `sanitize` preserves newlines on purpose, because a
vendor's prose reply is prose; a tool failure's detail is not prose, and multi-line stderr
pushed through the wrapper arrived as ragged fragments at random widths. It is now collapsed
to one flowing line and capped at three rows with the room's own ellipsis, so a clipped reason
can never read as a complete one. The clip has an answer — `f` expands the column to the full
frame, where the same reason typically survives whole — and a refusal behind it: the trace
answers *what did this agent do and did it work*, not *show me the log*, and the turn-level
failure still arrives in the column's note carrying the vendor's own sentence.
### 9.6b The agy trace was showing its message-passing and hiding its work
Driven live, the Antigravity column's trace read `user_input ?`, `system_message ?`,
`checkpoint ?`, `unknown ?` — and, for every real thing the agent did, a bare `tool ?`. Three
separate defects wearing one symptom, all found by reading captured stdout (agy 1.1.10,
Windows, 2026-08-04) rather than the adapter.
**1. The plumbing is suppressed, and the line that decides what counts as plumbing is not
"noisy".** Hiding a vendor's ACTIONS would be a false gauge — a quiet column for an agent busy
editing the workspace, which is §4a.1's failure with the sign flipped. Hiding its PLUMBING is
noise reduction. So the suppression is an allowlist defended per kind against a captured line,
never a filter on what looks like chatter, and it lives in the adapter (`ParseEvent` returns
`false`) rather than in the view, because `Render` is pure over `State` and a step that is not
an action must never become one.
| kind | why it is plumbing, from the capture |
|---|---|
| `user_input` | step 0 of every turn, `DONE`, nothing else on the line — the brief council itself just sent, echoed back |
| `system_message` | same empty shape; agy placing its own message into the conversation |
| `checkpoint` | `duration_seconds` and a ~120-token usage block, nothing else — a thread bookmark, never the workspace |
| `error_message` | an empty marker on a failing turn: no message, no error field, no duration |
| `unknown` | one per turn at a fixed preamble slot (step 1, right after `user_input`), 0.0005s and 0.0045s across two turns, no tool name, no parameters |
`error_message` needed the most care, because dropping the only visible sign that a turn went
wrong is the opposite mistake. It is safe for a checked reason: both captured failing turns end
`result` with `status:"ERROR"` and `error:"Agent execution terminated due to error."`, and that
path already produces a `KindError` carrying the vendor's sentence. The turn-level failure IS
reported, with words. A rendered `error_message ?` is strictly *less* than that — an ominous
name with a shrug attached. The result path now prefers `result.error` over the composed status
line precisely so that argument keeps holding.
`unknown` had to be argued rather than listed, since suppressing a step whose type the adapter
merely does not RECOGNISE is the same class of mistake as inventing an outcome for it. The
capture says this is agy's own label and not our ignorance: fixed position, half a millisecond,
and no tool name — while every step in every capture that did something carried one. What
would reverse the decision is written as code, not as a promise: an `unknown` step that names a
tool is **not** suppressed and renders under that name, so if agy ever starts acting through
this label the trace shows it that same turn.
**2. agy's real tool names were on the wire the whole time and were not being read.** A tool
step carries `tool_name` at the top level *and* `tool_info.name` with `tool_info.parameters`
beside it. The adapter rendered `step_update.step_type` — the literal string `"tool"` — so every
call, whatever it was, produced one indistinguishable entry. This is ADR-008's tenth amendment
repeating itself in a second costume: Cursor's `tool_call.tool.case` lookup matched nothing
because the oneof arrives flattened to a key on the wire, and every Cursor trace entry read
`tool call`. Same cause both times — the fields the vendor sends were never compared against
the fields the parser reads — and the same fix, which is to **parse what arrives**. Observed
names: `list_dir`, `run_command`, `write_to_file`, `list_permissions`.
The entry now follows the grammar the other three adapters already use — `Glob: **/*.go`,
`Bash: go test ./...` — so `⚙ tool ?` becomes `⚙ list_dir: C:\Users\…\antigravity-cli\scratch ?`.
The argument rule is deliberately small, because agy's parameter keys are vendor-specific
(`DirectoryPath`, `CommandLine`, `TargetFile`) and an arbitrary object is not a trace line: only
string values are candidates, one such value renders (which is every captured shape), several
resolve to the lowest key name by byte order, none degrades to the bare tool name. Rule three is
not a claim about which key matters — it is a refusal to let Go's randomised map iteration reach
a rendered line or a golden, pinned by `TestAgyToolArgIsDeterministic`.
**3. A failed agy tool call used to render as permanently pending.** There is a fifth state,
`ERROR`, and the switch handled `ACTIVE` and `DONE` only, so the line matched nothing and the
entry its `ACTIVE` twin had opened stayed pending for the rest of the room's life — the trace
claiming a command was running after the vendor had given up on it. It carries its own reason in
`tool_info.error.message`, so it maps to Failed with the vendor's own first line, exactly as
§9.6a specifies. Failed and not Denied: `ActDenied` is council's first-hand record of its own
gate keystroke, and a refusal read off someone's stream is not that. The `DONE → Unknown` rule
above is unchanged and narrows in one direction only — agy does report per-step failure, and it
still reports no per-step success.
**The resume note was a misdiagnosis, and the fix is to the claim rather than to the
mechanism.** A seat whose restored thread failed its first turn used to say *"the saved thread
was refused — this seat's history is gone."* **Measured**, single trial, 2026-08-04: `agy
--conversation ` **does** resume in 1.1.10. The same `conversation_id` came back,
`step_index` **continued** (10 → 11) rather than restarting at 0, and `result.num_turns` was 2.
That demonstrably live thread's turn nevertheless ended `status:"ERROR"` /
`"Agent execution terminated due to error."`, and a separate attempt died before any thread was
involved at all — a bare `result` with an **empty** `conversation_id` and *"Eligibility check
failed: UNAVAILABLE (code 503): The service is currently unavailable."* So agy turns fail
transiently for reasons that have nothing to do with the conversation, and "the history is
gone" is a claim the evidence does not support.
The behaviour was deliberately untouched at the time: one failed turn still dropped the id, for
the reasons ADR-008's ninth amendment gives at length, and no new signal was invented to tell
the two cases apart because none had been observed. Only the sentence was narrowed, to the
three things known — the first turn on the restored thread failed, the seat has let the saved
thread go, and the next brief starts a new session with the brief re-applied.
**That is now out of date, and the evidence above is what dated it** (ADR-008, sixteenth
amendment). The paragraph declined to change the rule on the ground that a record is not a fix,
which was right — and left a measurement sitting beside a rule it contradicted. The rule's
default is unchanged: a restored id whose first turn fails is dropped, because a seat retrying
a genuinely dead id rebuilds the same doomed invocation on every turn for the life of the room.
What changed is that a failure which is **identifiably transient** is now treated exactly as a
cancellation already is — nothing was learned about the thread, so nothing is forfeited, and
the seat stays on probation so the next unclassified failure still costs it the id.
Two classes qualify, and both are positive evidence that the vendor never reached the
conversation. **Pre-flight**: the failures `failureNote` already classifies off captured stderr
— not signed in, an untrusted workspace, a sandbox the vendor's own config demands and its own
help refuses, a binary that vanished — each documented at its case as exiting before any model
call; and, one step earlier, a dispatch that never started a process at all. **Vendor-reported
outage**: the 503 quoted above, matched on agy's own sentence, with the capture's empty
`conversation_id` as the corroboration that it died before a thread was involved.
Everything else still drops, and the asymmetry is deliberate — a lost conversation costs one
conversation, a wedged seat costs every turn of the room. Claude and Codex get nothing
vendor-specific here: neither has a measured transient signal, only measured *dead-thread*
strings pointing the other way, and their behaviour is unchanged. agy's commonest failure
sentence — *"Agent execution terminated due to error."* — is deliberately **not** classified,
because it was captured on a demonstrably live thread and is also what a dead one would
plausibly produce, and a string on both sides of a distinction is evidence for neither.
The classification is produced where the evidence is (the runner's stderr classifier, the
adapters' result parsers) and travels on the event as a small enum. It is never re-derived by
matching the rendered note: that note is prose written for a narrow column, and keying a
mechanism off it would make every wording change a silent behaviour change. It lives on `Model`
and never on `State` — a decision input the renderer has no business reaching.
**The card that says this changed shape too.** One ⚠ plus a single sentence carrying an outcome
and a mechanism wraps to three lines of uniform weight in a 37-cell column, and three of those
side by side reads as a room on fire over a seat that will simply start a new session. It is
now §9.11's card grammar: a short title — *thread not restored — starting fresh* — with the
mechanics hanging under it, quieter, and **no warning mark**. That is the same fact
`reattachCard` already states calmly at idle when no thread came back, learned a turn later;
the ⚠ has to go on meaning *something went wrong* for the notes where something did. The words
carry the card in every glyph set, so `--ascii` loses nothing.
**A side measurement, separately labelled, with its confound stated.** Under `--mode plan
--sandbox`, agy's `run_command` was refused with *"granting access to C:\: Access is denied."*,
the agent gave up, and the whole turn died `status:"ERROR"` with an empty response. The control
run with both flags **dropped** ran a shell command and returned `status:"SUCCESS"`. ADR-008 and
the `baseArgs` comment previously said `--sandbox`'s effect on the shell "was NOT tested and is
not claimed"; this is the first evidence on it, and that comment no longer says so. **It is one
trial per arm with an uncontrolled difference: the two turns issued different command lines
(`pwsh -Command "Get-Location; Get-ChildItem"` versus `Get-ChildItem`), so it is not a clean
A/B**, and the refusal's mention of `C:\` may be about a drive root rather than about the flag.
What it does establish is the flag's observed cost — a dead turn, with nothing rendered. The
posture flags were **not** changed on the strength of it; that is a decision to make
deliberately and separately, and this was a record, not a fix.
**That decision is now made: the flags come off, in both postures** (ADR-008, seventeenth
amendment). The open question this section carried — *should council keep asking agy for
`--mode plan --sandbox` when both are measured to do nothing?* — is closed the way the evidence
points, and the ledger is one-sided rather than a close call:
- On the **write** side, the flags were measured restricting nothing. Asked to write a file
under both, agy wrote it; reported permission mode and tool list were byte-identical to a run
without them, and `write_to_file` was still in the list. Refuted, not unproven.
- On the **shell** side, the one and only effect either flag has ever been observed to have is
the paragraph above: a refused `run_command`, an agent that gave up, and a turn that ended
`status:"ERROR"` with an empty response. The user sees a blank column.
- So the flags bought **no restriction that was ever observed**, at the price of turns that die
with nothing rendered. That is not caution; it is the appearance of caution paid for in the
vendor's actual answers. The confound above is unresolved and does not need to be: it concerns
*why* the turn died, and the decision only needs *that* it did, set against a benefit measured
at zero.
**No honesty claim moves with them, and that separation is the point of having waited.** §9.13
deliberately changed what the room *says* about this posture and nothing about the posture,
because a documentation pass that quietly retunes a safety flag is exactly what this file exists
to prevent. This is the other half, made on its own, by the owner. The badge stays
`unsandboxed`; the detail loses the clause claiming the flags are passed, because they are not,
and a detail describing council's own behaviour inaccurately is the one class of false claim
this repo has no excuse for. The containment was never these flags — it is the workspace (§9.2,
ADR-008 third and twelfth amendments), and agent-ops ADR-012 rules the same way independently.
Deliberately **not** part of this: `--dangerously-skip-permissions`. Dropping a flag that
restricted nothing and adding one that approves everything are different acts, and the second
stays refused on both seats that offer it.
**Amended 2026-09-03: two flags come ON, and neither is a restriction.** `--add-dir `
in every posture and `--mode accept-edits` in the write postures. The first names the workspace,
because the cwd alone is none to this vendor. The second is the vendor's edit-only grant, and
without it print mode auto-denies every write inside a named workspace. [§9.59](#s9-59) carries
the measurement. `--dangerously-skip-permissions` stays refused.
### 9.6c The Cursor stream says everything twice, and the second time does not always look alike
cursor-agent under `--stream-partial-output` sends a model call's text deltas and then that
call's **complete message** as one more assistant event. Appending both renders the passage
twice, which is the whole of this defect in both of its appearances.
The first capture (2026-08-04, a turn asked to reply `PONG`) showed deltas `"P"`, `"ONG"` each
carrying `timestamp_ms` and the repeat `"PONG"` carrying none, so the adapter dropped the event
whose `timestamp_ms` was absent. That rule was derived from turns with no tool call in them, and
**every such turn is one model call** — so it was a rule about the end of a *turn* being used as
a rule about the end of a *message*.
A turn that runs a tool is several model calls, each ending in a repeat of its own segment, and
those mid-turn repeats carry `timestamp_ms` like any delta. The column rendered the segment, the
segment again, then the next one — `X X Y` — which is what the owner saw on a long Cursor reply
and what replaying the captured turn through the old parser reproduces exactly.
What separates them is **`model_call_id`**: present on every whole-message repeat that ends a
mid-turn model call, absent from every one of 108 captured deltas, and carrying the vendor's own
per-segment numbering (`…-0-x7su`, `…-1-15l2`) that also appears on the `tool_call` events
between the segments. The adapter now drops an assistant event when `model_call_id` is present
**or** `timestamp_ms` is absent — the second is kept because the *turn-final* repeat still
carries neither, so dropping it would trade this bug for the first one.
Presence rather than absence is the point, and it generalises past this vendor: a missing field
cannot distinguish "the vendor is telling me this is a complete message" from "the vendor stopped
sending that field". `internal/council/vendors/testdata/cursor-segmented-turn.jsonl` is the whole
turn, redacted, replayed by `TestCursorSegmentedTurnRendersEachPassageOnce`, which asserts the
streamed body equals the reply the vendor itself put in its `result` event. That `result` remains
the safety net if both fields ever go: the room uses it whenever a column streamed nothing, so
the failure mode is a column that fills at the end, never one that is wrong. ADR-008's twentieth
amendment carries the argument.
### 9.7 Status
The room opens, detects the four seats, renders both layouts and every degraded state, takes a
brief, and dispatches it. Claude streams incrementally; Codex and Antigravity render the waiting
card and fill at once. Quitting the room kills every child, including the persistent one —
and no longer strands the conversation: a bare `telltale council` reopens the one saved room
by default, `--fresh` starts over, and `/cd ` typed in the composer moves the room to
another workspace between turns, with the persistent Claude seat following by respawn on its
own session id (ADR-008, ninth and eleventh amendments). Multi-turn is one live process for
Claude (§9.8) and for Cursor (§9.36), and native resume for the two seats that are still batch
programs.
Cross-agent rebuttal (§9.4) and per-column scrollback are **built and shipped**, and the
scrollback now spans the whole conversation rather than one turn (§9.9): the room keeps a
per-column transcript, echoes the brief that produced each turn, composes in up to six rows,
and folds the seats it cannot drive out of the grid. Not built: per-vendor cancel — `ctrl+c`
still ends the whole turn.
Known gaps, stated rather than buried. Codex's non-shell write path is untested — asked to
create a file with its own patch tool it declined, but that was a model choice and says nothing
about enforcement. Neither vendor was observed producing a failure event on stdout, so the error
branches are modelled on exit code plus stderr rather than an observed schema. Antigravity's
`--print-timeout` is left at its 5-minute default, which is a hard ceiling on a long council
turn and a policy choice worth making deliberately later.
Unverified and scheduled as a live spike before the Codex and Antigravity columns ship: the
Codex `--json` event schema and delta granularity, whether `codex -s read-only` engages on
Windows, and Antigravity's stream-json schema, conversation-id location, stdin support and
`--sandbox` semantics. Those columns render honest *requested* badges until it says
otherwise.
**That spike ran, and the Windows sandbox question is closed the other way.** `-s read-only`
does not engage on Windows in any useful sense: it fails every process spawn, reads included,
and so does `-s workspace-write`. Codex on Windows is invoked `danger-full-access` and badged
`unsandboxed` — see the §9.2 table above and ADR-008's twelfth amendment. What remains open on
this seat is narrower: whether the `-c sandbox_mode=` override actually changes behaviour on
the *resume* path. The key is accepted; its effect has never been separately observed, and
until this change every mode failed identically so there was nothing to observe.
*(Both halves of that paragraph moved on 2026-08-29, at codex-cli 0.149.1: the Windows sandbox
now enforces and the read posture is `-s read-only` again, and the resume override's effect was
observed — a resumed read-only turn had its shell write denied. §9.2's 2026-08-29 amendment
carries the measurements. The paragraph above is kept as the record of what was true at
0.146.0.)*
One claim in this section is looser than its measurement and is flagged rather than quietly
corrected, because fixing it is a separate change to a separate surface. The Claude column's
granularity word is `tokens`. Measured over a 250-word reply the deltas are **~80 characters
each, about three a second** — genuinely incremental, and not tokens. Measured identically
under the persistent invocation and under a spawn-per-turn control, so it is a pre-existing
overstatement rather than something §9.8 introduced.
### 9.8 One live process, and the gate it makes possible
Every Claude turn used to be a fresh `claude -p --resume`. A one-word "gm" cost about 25
seconds and $0.23, nearly all of it session init, paid again on every turn. That was the
visible cost. The structural one is what forced the change: **a batch process cannot ask
permission.** Its stdin is written and closed before the first token arrives, so it has no
channel to ask on and none for an answer to come back on.
`--input-format stream-json` keeps one process alive taking one JSONL message per turn on an
open stdin. Verified live against Claude Code 2.1.220 rather than read from documentation: two
turns down one stdin came back under the same `session_id` with the same pid, `system/init` is
re-emitted at the *start of every turn* (a parser reading it as "a new session" would reset the
seat once per turn), and one `result` per turn is the only end-of-turn signal there is, because
there is no exit to infer one from.
Cancelling a turn now **interrupts** rather than kills — `{"subtype":"interrupt"}` on the
control channel, measured to end the turn and leave the process answering a further one.
Killing would also work, and would throw away the session init the room just paid for, so
cancelling one turn would quietly make the next one expensive.
**The reported cost changed meaning and the badge changed with it.** `total_cost_usd` is a
running total for the process: $0.1061493 → $0.1177296 across two turns while the per-turn
`usage` block stayed at 2 input tokens both times. That cell has meant "this turn" everywhere
else in the room, so it now reads `$0.1177 session`. Rendering a session total unlabelled would
be a false reading of a true number; subtracting to recover the turn would be council inventing
a figure, which §8 rejects.
#### The gate
With `--write`, the seat that can ask **does** ask. Every tool call raises an approval card in
its column — the tool and its argument line, formatted exactly as the activity trace formats it
— and the room enters a gate state: `y` approves, `n` denies, and the vendor is stopped until
one of them is pressed. Blocking was measured, not assumed: the answer was withheld for twenty
seconds and nothing else arrived on stdout in that window.
Three flags turn it on, none is optional, and **two of them do nothing alone**:
| flag | what it does | what happens without it |
|---|---|---|
| `--permission-prompt-tool stdio` | routes the request onto the stream | **absent from `--help`** and real; alone, the session runs in auto mode, no request is ever emitted and the file is written |
| `--permission-mode manual` | makes the call ask rather than assume | alone, there is nobody to ask, so the call short-circuits to *"you haven't granted it yet"* and the vendor gives up |
| `--setting-sources ""` | stops the user's own permission rules pre-approving the call | measured on a machine allowing `Bash(mkdir:*)`: `mkdir zzz` **ran ungated** and the directory was created |
**The third flag became a FALLBACK on 2026-08-12** and the table row is the record of why it
was ever needed. Council injects its own `PreToolUse` hook instead, which runs at step one and
beats an allow rule at step five, so the operator's settings stay loaded. The dated block at
the end of this section carries the measurements and the build.
The third is the honesty of the whole feature, and it is the one nobody would have thought to
test. Permission *allow rules* in settings files are consulted **before** the callback, so a
call they cover never reaches the gate at all. Without that flag, "nothing writes without your
keystroke" is simply false — and false quietly, on a machine whose owner wrote those rules
years ago for a different purpose.
**Amended 2026-08-11, at the end of this section.** That paragraph is still true and it is
still the record of 2026-08-04. It reads one step of the evaluation, not the step above it: an
`ask` rule is consulted BEFORE an `allow` rule, and it reaches the callback the allow rule
would have skipped. So the third flag is no longer the only way to be honest. The measurement
and the build it implies are below.
**One limit is stated on the badge rather than buried here.** Shell commands the CLI itself
classifies as read-only are approved without asking — `git status` was ungated under both
setting-source configurations, and so is `echo` — so the claim is about calls that *change*
things and is worded that way everywhere.
**RETIRED 2026-08-12, and kept because it is the record of a hole that existed for eight days.**
Everything in the next four paragraphs describes council copying the USER's hooks into the
ephemeral file. The seat no longer drops their settings, so there is nothing to copy and a copy
would run every one of their hooks twice. The ephemeral file survives, built the same way and
for the same reason, carrying council's own gate hook instead.
**The second limit was a hole, and it is now closed.** Dropping the setting sources also
dropped the user's own hooks and user-level commands from that seat. Half of that is the
feature working: the allow rules are what the gate replaces. The other half was collateral — a
`PreToolUse` hook is a screen the user built, nothing was replacing it, and the calls it
covered are disproportionately the ones the gate never sees. Measured: in the gated posture,
`echo ` raised no request and simply ran.
`--settings ` composes with `--setting-sources ""` — the sources stay dropped and the
named file is still read — so council copies the user's `hooks` section into an ephemeral file
of its own and points the gated seat at it. Two properties of that file are load-bearing:
- **It is built by naming one key, never by deleting others.** The same spike showed a
`permissions` block inside a `--settings` file re-admits the allow rules: an allowlisted
`mkdir` ran with no request and the directory landed on disk. An allowlist of exactly `hooks`
cannot rot as Claude Code adds settings keys; a denylist would.
- **The badge is derived from whether the file exists**, not from whether council tried. An
unreadable settings file, an empty hooks section and a temp directory that could not be
created all end in the same place, and the column says the guard is absent rather than
claiming one.
The file is absolute (a relative `--settings` path resolves against the *child's* working
directory, which is the workspace, and fails), 0600, removed on teardown, never logged and
never rendered — the same privacy discipline `--brief` carries, for the same reason: only a
boolean crosses onto `State`.
The honest residual: hooks fire as that file described them at spawn time. Editing the real
settings mid-session does not propagate until the next room.
**A denial is not a failure, and the difference took a fifth outcome value to keep.** The vendor
reports a refusal as an `is_error` tool_result carrying council's own refusal text back — so
read off the stream alone it is indistinguishable from a tool that broke, and the trace would
say the command *failed* when what happened is that it was *not allowed to run*. `ActDenied` is
recorded from the keystroke, before the echo arrives, and the echo cannot overwrite it. It
renders `✗ denied by you`: the words carry the distinction, colour only seconds it, and it is
the one line in the trace that is not a reading of a vendor's words.
**The gate is Claude-only, and that is a fact about the other CLIs.** `codex exec` and `agy -p`
are batch programs — read a prompt, answer, exit. Neither has a channel a question could arrive
on. Their columns keep `WRITES`; only the seat that asks carries `gated`. Giving all four the
same badge would be the blanket claim §9.2 exists to refuse, one level up.
**Amended by §9.36, and the amendment is narrower than it looks.** The Cursor seat is now a live
ACP process and it *can* be asked: `session/request_permission` blocks it until answered, measured
on both branches. It still does not carry `gated`, because it does not ask about EDITS — measured
twice, it wrote a file and raised nothing. So the last sentence above holds with one word changed:
only the seat that asks about **everything that changes anything** carries `gated`. Council answers
Cursor's requests all the same, because an unanswered one blocks the vendor forever.
`--write --auto` restores the old behaviour for the times nobody is watching: `acceptEdits`,
the `WRITES` badge, the user's settings left alone — and therefore no injected hooks file
either, since a room that loads those settings natively would otherwise run every hook twice.
Gating is the default because the room the user opened is the one they are looking at;
unattended is the exception and has to be typed.
#### The gate can keep the user's settings, measured 2026-08-11 — and the build is not authorised
**Superseded by the block below it, 2026-08-12, and kept whole.** Nothing changed in the
product on this date; this is the record of the probe and of the ruling it waited for. The
ruling came the next day — measure the two open unknowns, build only if both support it — and
both did.
**The claim.** Claude Code evaluates a tool call in six steps, in this order: hooks, deny
rules, ask rules, permission mode, allow rules, then the `canUseTool` callback. That is the
live documentation of 2026-08-11 (`code.claude.com/docs/en/agent-sdk/permissions`), not
memory. `ask` is step three and `allow` is step five, so an `ask` rule reaches the callback
that an `allow` rule would have skipped. The same docs say a `PreToolUse` hook may return
`permissionDecision` `"allow"`, `"deny"`, `"ask"` or `"defer"`, and that `ask` beats `allow`
when both apply. If either holds in the binary, the gate can keep `--setting-sources` and
gate anyway — and keeping them carries the user's own hooks, deny rules and user-level
commands back in natively, which is the whole of what this seat gives up today.
**The rig, because this repo does not read a claim off a doc.** A probe replicated this
seat's own argv — `baseArgs` plus `Session`, gated posture, `--model haiku` added to keep the
turns cheap — spawned it against a throwaway directory, wrote one turn on stdin as `Turn()`
builds it, and answered any `can_use_tool` request with `behavior: "deny"` as `Decide()`
builds it. `--setting-sources ""` was dropped on every arm except the one that reproduces
what council ships. **The decisive observable is the filesystem, never the stream**: the
brief asked for one command that creates a marker, and the arm is read by whether the marker
is on disk. Claude Code 2.1.226, Windows 11, `claude-haiku-4-5` on every turn, two trials
each.
| arm | what it changed | requests | marker created | trials |
|---|---|---|---|---|
| **A** adopter | zero-rule `CLAUDE_CONFIG_DIR`, no `--setting-sources` | — | — | **blocked** |
| **A2** | user settings live, `touch probe-marker` | 0 | **yes** | 2/2 |
| **A3** | user settings live, `install -d probe-marker` | 1 | no | 2/2 |
| **B1** | user settings live, `mkdir probe-marker` | 0 | **yes** | 2/2 |
| **B2** | B1 plus an injected `ask` rule for `Bash(mkdir:*)` | 1 | no | 2/2 |
| **C** | user settings live, injected `PreToolUse` hook returns `"ask"` | 1 | no | 2/2 |
| **C control** | same hook wiring, hook returns no decision | 0 | **yes** | 2/2 |
| **shipped** | council's argv today, `--setting-sources ""` | 1 | no | 2/2 |
**B1 says the 2026-08-04 finding still reproduces.** An allow rule covers `mkdir`, the call
ran, and the directory is on disk. Nothing here retires the original measurement.
**B2 is the decisive arm.** One rule was added — `{"permissions":{"ask":["Bash(mkdir:*)"]}}`
in a `--settings` file — over settings that already allow the same shape. The call raised a
request, the denial was honoured, and nothing was created. The request named its own cause:
`"decision_reason_type":"rule"`.
**C says a hook can do the same job, and says more on the way through.** A hooks-only
`--settings` file whose `PreToolUse` hook returns `permissionDecision: "ask"` gated the same
allow-covered call, and the request arrived carrying `"decision_reason_type":"hook"` with the
hook's own sentence in `decision_reason`. The hook wrote a breadcrumb on every trial, so its
run is provable off the stream. **The control matters as much as the arm**: C changed two
things at once, a `--settings` file and an "ask" behind it, so the same file was run again
with a hook that returns nothing. The call went ungated and the directory landed. The
decision causes the gate, not the file.
**Two facts fell out that were not the question.** `--settings` composes with the user's
settings rather than replacing them — every sources-live arm ran the user's own `SessionEnd`
hooks, including the arms passing `--settings`, and the shipped arm ran none. And a write
shape no rule covers already reaches the gate with sources live (A3), so today's flag is not
what makes the gate fire; it is what makes it fire *uniformly*.
**What build this implies, and it is the hook rather than the rule.** A2 is why. `touch`
creates a file and it ran ungated on both trials, so the user's rules cover more shapes than
anyone would enumerate, and an `ask` list built shape by shape leaks exactly the way an allow
list leaks. That is the same defect `hookset.go` (now `gatehook.go`) already refuses by naming one key instead of
deleting many. A matcherless `PreToolUse` hook has no list to leak: the documentation's own
advice for a check that must run on every tool call is a hook, for this reason. So the shape
to build is council injecting its own `PreToolUse` hook that answers `"ask"`, into the same
ephemeral `--settings` file it already writes, and **dropping `--setting-sources ""`** — which
returns the user's deny rules, their user-level commands and their hooks to the gated seat,
and retires the hooks copy in `hookset.go` (the file became `gatehook.go`) along with the badge that reports whether it
worked.
**What is NOT settled, and each item is a reason the build waits.**
- **The adopter arm did not run.** A temporary `CLAUDE_CONFIG_DIR` holds no credentials, so
the turn died at `"apiKeySource":"none"` and `Not logged in · Please run /login` before any
tool call. Copying a credential store into a probe directory is a redline, so the arm stays
unrun. The shipped arm is the nearest evidence for the same question: with no rules in
force at all, the call reached the prompt on both trials.
- **A matcherless hook was never measured.** This rig measured a `Bash` matcher. The claim
that one hook sees every tool call is documentation, and documentation is what this section
exists to distrust.
- **Composition with the user's own `PreToolUse` hooks is unmeasured.** The docs rank `deny`
over `defer` over `ask` over `allow` when several apply. Council would be adding a second
hook to a file the user also populates, and the credential guard is exactly the hook that
must not be weakened by the addition.
- **A hook is a process per tool call.** The seat that was re-founded to stop paying process
cost per turn (§9.33, §9.36) would take on a spawn per call. Nothing here timed it.
- **The badge's sentence would have to change.** Today it claims a guard because a hooks file
exists. Under this build the guard IS the gate, and "the user's hooks are carried" stops
being a separate claim — a badge that kept saying it would be reporting a file that no
longer does that job.
#### The two deciding measurements, and the build, 2026-08-12
The owner ruled: measure the two items above that decide the build, then build only if both
measurements support it. Both did, and the build is in. Claude Code **2.1.228** (the box moved
on from 2.1.226 between the two dates), Windows 11, `claude-haiku-4-5` on every turn, two
trials per arm, throwaway directories, the same rig as the block above — a probe replicating
this seat's own argv, one turn written on stdin as `Turn()` builds it, `can_use_tool` answered
as `Decide()` builds it, and **the decisive observable is the filesystem, never the stream**.
The adopter arm stays unrun for the same reason: copying a credential store into a probe
directory is a redline.
**(a) A matcherless hook fires for every tool shape, and the ask reaches the card.** The arm
kept the operator's settings live — no `--setting-sources ""` — and injected one `PreToolUse`
entry with no `matcher` field, returning `permissionDecision: "ask"`.
| arm | `mkdir gate-a` (an allow rule covers it) | `install -d gate-b` (no rule covers it) | `Write gate-c.txt` (not a shell command) | on disk | trials |
|---|---|---|---|---|---|
| **M** matcherless hook returns `"ask"` | request, `hook` | request, `hook` | request, `hook` | nothing | 2/2 |
| **M control** same file, hook returns no decision | **no request, directory created** | request | request | `gate-a` | 2/2 |
| **shipped binary** the file council now writes, `telltale hook gate` | request, `hook` | request, `hook` | request, `hook` | nothing | 2/2 |
Every request carried `"decision_reason_type":"hook"` and the hook's own sentence in
`decision_reason`, forwarded verbatim to the `can_use_tool` card. **The control is what makes
this a finding**: the same file, the same hook process running — its breadcrumbs prove it — and
only the decision removed. `mkdir` went ungated and the directory landed, which also
re-reproduces the 2026-08-04 bypass on 2.1.228. The decision causes the gate, not the file.
**The matcher forms were measured against each other** in one turn, four entries side by side
writing to four breadcrumb files: **matcherless, `"*"` and `""` each saw both the `Bash` call
and the `Write` call; `"Bash"` saw only the `Bash` call.** That one hook sees every tool call
was documentation until this turn. The absent field is what ships, of the three equivalent
forms, because it is the only one that cannot later be read as a pattern somebody should widen.
**A Windows trap, and it is the worst failure this feature has.** Claude Code hands the hook
command to **`/usr/bin/bash`** — Git Bash, on the platform this product primarily targets. The
first three arms measured nothing because bash ate every backslash of a native Windows path:
```
/usr/bin/bash: line 1: C:UserssanleAppDataLocalTempclaudeC--…askhook.exe: command not found
```
`exit_code: 127`, `outcome: "error"` — and **a hook that fails to run makes no decision, so
every call ran ungated while the badge went on claiming a gate**. It was found only because a
`SessionStart` hook was planted in the same file to prove the file was read at all, which costs
no model turn. Council quotes the command and swaps the separators;
`TestTheHookCommandSurvivesGitBash` pins both.
**The first fix for it was wrong, and the Linux CI job is what said so.** `filepath.ToSlash` is
a **no-op on Linux**, where a backslash is a legal filename character, so it made the
conversion depend on the host Go compiled for. The string is not read by the host — it is read
by bash, on every platform, where a backslash is the escape character and cannot survive as
itself. The swap is now unconditional, and the test feeds a Windows path on every runner.
**(b) The hook costs tens of milliseconds, and the operator's own settings cost more.** The
measure is the interval between the assistant's `tool_use` block landing on stdout and the
`can_use_tool` request landing — the window Claude Code evaluates permissions and runs hooks in.
| arm | per-call gap, all samples (ms) | median | trials |
|---|---|---|---|
| **shipped today** — `--setting-sources ""`, no hooks at all | 23.0, 12.3, 21.2, 10.5, 6.8 | **12.3** | 2 |
| operator's settings live, **no** council hook | 284.6, 280.5, 291.0, 234.6 | **282** | 2 |
| live + council's hook, a 3 MB probe binary | 358.9, 351.7, 249.5, 348.4, 298.4, 258.4 | **323** | 2 |
| live + council's hook, **the shipped 14 MB `telltale.exe`** | 523.6, 491.2, 461.2, 451.8, 439.7, 362.7 | **456** | 2 |
Read the rows against each other rather than against zero, because **most of the delta is not
the hook**. Loading the operator's settings at all costs ~270 ms per call — that is their own
`PreToolUse` hooks running, and it is the thing this build BUYS, not a price it adds. Council's
own hook adds **~41 ms** as a small binary and **~174 ms** as the shipped one. Measured
directly, outside the CLI, 20 spawns through the same Git Bash: the small binary is
**36.2 ms median**, `telltale.exe` is **54.4 ms median** — the 14 MB single binary links the TUI
framework on a path that runs once per tool call, which is ADR-002's statusline argument
arriving at a second door. Against a warm Claude turn of 6.4 s (`STATE.md`'s traced `@all`), three
gated calls add ~0.5 s. That is the owner's "tens of milliseconds is fine" band at the binary's
own cost and inside it at the process's; it is nowhere near the "a second per call" that fails.
**What shipped.** Council writes an ephemeral `--settings` file containing exactly one key,
`hooks`, holding one matcherless `PreToolUse` entry that runs `telltale hook gate` — a new mode
beside `telltale hook cursor`, which drains stdin and prints one decision object and nothing
else. `--setting-sources ""` is **dropped**, so the gated seat loads the operator's deny rules,
their user-level commands and their own hooks again. `--permission-mode manual` is **kept**:
the documentation makes it an alias for `default` on 2.1.200+, which would make it decoration,
but every arm of both probes carried it and nothing has measured the seat without it — the same
rule that kept `--permission-prompt-tool stdio` when it was absent from `--help`.
**The fallback is the old build, not a hole.** A room whose hook file cannot be written — no
temp directory, a binary that cannot locate itself — passes `--setting-sources ""` and gates the
2026-08-04 way. It gives up the operator's settings and the column says so. Weaker in what it
keeps, never weaker at the gate.
**One cost was not on the list of five, and it changed the room.** The hook asks about
EVERYTHING, which is the point — and `Read`, `Glob` and `git status` raised **no request at all**
under the old flag, because Claude Code approves what it classifies read-only before the
callback. Under the hook all three raise one. Shipping only the hook would have tripled the
cards, and this room already knows what that costs: the first session with the gate carded the
user thirty-four times, which is why `autoApproveRoutine` exists. So council answers them
itself — `autoApproveRoutine` for shell commands, and a new positive list of tool names that
change nothing (`Read`, `Glob`, `Grep`, `NotebookRead`) for the calls that are not shell
commands. Positive, so a tool Claude Code adds next month draws a card rather than being waved
through; `TodoWrite` is deliberately absent, because nothing here measured what it writes.
**The badge's sentence changed, as the fifth item predicted.** It no longer claims the
operator's permission rules are dropped, because they are not. The wired branch says their
settings stay loaded and that council's own hook asks first; the fallback branch says the
settings were dropped and why. Both branches still say nothing runs until you answer.
**Of the five unsettled items, two are settled, two are retired by the build, and one remains.**
The matcherless hook and the per-call cost are measured above. The badge sentence and the
`--setting-sources ""` question are decided by what shipped. **Composition with the operator's
own `PreToolUse` hooks is still unmeasured** — council now adds a second hook to a file the
operator also populates, the docs rank `deny` over `defer` over `ask` over `allow` when several
apply, and the credential guard is exactly the hook that must not be weakened by the addition.
The ranking makes a weakening unlikely (a `deny` from their hook beats council's `ask`), and
"unlikely by documentation" is the standard of evidence this section exists to distrust.
**The per-call cost tolerance is HELD — owner ruling, 2026-08-15.** The measurement above
stands: the shipped 14 MB binary costs ~54 ms per gated call because the one binary links
the TUI framework, and a sibling hook binary would save ~133 ms per gated call. The owner
ruled the saving does not buy an ADR-002 amendment: roughly half a second across three
gated calls on a 6.4 s turn sits inside the "tens of milliseconds is fine" band the build
was accepted under, and the split's real price is not the binary — it is the packaging
(the scoop manifest ships one exe), `hookCommand`'s self-location, room/hook version skew,
and a missing-sibling fallback whose only honest shape is `--setting-sources ""`, which
trades away the operator's settings to save milliseconds. The one-binary shape stands
unamended. Reopen this only with a measurement that changes the arithmetic: more gated
calls per turn than the ~3 assumed, or a vendor change that raises the per-call floor.
#### The operator's deny beside council's ask, measured 2026-08-15 — the last item closes
**The last unsettled item of the five is now measured, and the operator's guard does not come
back weaker.** The item said this: council adds a `PreToolUse` hook to a file the operator also
populates, the documentation ranks `deny` over `defer` over `ask` over `allow`, and the
credential guard is the hook that must not be weakened. The rank held live. The deny also does
more than win — it stops the call before council's card is drawn at all.
**The environment, and the two claims it makes into hypotheses.** Claude Code **2.1.228**
(`claude --version`, recorded before any arm), Windows 11, `--model haiku`, two trials per arm,
throwaway directories. The floor for this measurement is 2.1.221: that release fixed auto mode
overriding a hook's `ask`, and 2.1.222 fixed exit code 2 not blocking. Both are changelog
claims, so both are hypotheses here, and both are confirmed by the arms below. A measurement
taken before 2.1.221 would be void.
**The rig is the shape of the two blocks above.** A probe replicates this seat's own argv
(`baseArgs` plus `gateArgs` plus `--input-format stream-json` plus `--settings`), writes one
turn on stdin as `Turn()` builds it, answers `can_use_tool` as `Decide()` builds it, and
**reads each arm off the FILESYSTEM, never off the stream**. The brief asks for one command,
`mkdir deny-probe-marker`, which the operator's own allow rules already cover. No credential
store is copied anywhere: `CLAUDE_CONFIG_DIR` is untouched, the operator's real login is used,
and `~/.claude/settings.json` is never edited.
**The deny hook has the credential guard's SHAPE and none of its content.** It is a
`PreToolUse` entry with matcher `"*"` that writes a reason to stderr and exits 2 — the
mechanism the real guard uses (`Exit 0 = allow, exit 2 = block`). It is installed in the
THROWAWAY workspace's own `.claude/settings.json`, which is a real setting source and belongs
to nobody. It writes a breadcrumb on every run, so its execution is provable off the disk.
**A wiring probe costs no model turn, and it is what makes the arms readable.** The seat was
started and its stdin closed without a turn. A `SessionStart` breadcrumb from the workspace's
own settings arrived while `--settings` carried council's gate file, so the two surfaces
demonstrably compose. That is the same `SessionStart` trick that found the Git Bash trap above.
| arm | hooks in front of the call | callback answer | requests | marker on disk | trials |
|---|---|---|---|---|---|
| **(a) control** | operator `deny` alone | allow | 0 | no | 2/2 |
| **(b) the question** | operator `deny` + council's gate hook | **allow** | 0 | **no** | 2/2 |
| **(b) again**, `--include-hook-events` | same | allow | 0 | no | 2/2 |
| **(b control)** | operator hook exits 0 + council's gate hook | allow | 1 | **yes** | 2/2 |
| **(c) anchor** | council's gate hook alone | deny | 1 | no | 1/1 |
**The callback answers ALLOW in arm (b), and that is the whole design of the arm.** A denial
pressed at the card would leave nothing on disk whatever the hooks did, so it cannot tell a
holding deny from a broken one. Answering `allow` inverts the test: if council's `ask` had
displaced the operator's `deny`, the marker lands. It did not land, on either trial.
**The (b) control is what makes (b) a finding.** One thing changed — the operator's hook exits
0 instead of 2 — and the marker landed on both trials. So the pipeline can create it, council's
`ask` still fires when nothing denies, and arm (b)'s empty directory is the deny.
**Both hooks ran, concurrently, and the stream says so.** `--include-hook-events` is
observability only and is NOT part of the seat's argv; it was added to one arm to see who
answered:
```
hook_started PreToolUse:Bash
hook_started PreToolUse:Bash
hook_response PreToolUse:Bash exit 0 outcome success stdout {"hookSpecificOutput":{"hookEventName":"PreToolUse","permissionDecision":"ask","permissionDecisionReason":"telltale council gates every call in this room"}}
hook_response PreToolUse:Bash exit 2 outcome error stderr BLOCKED by the operator credential-guard-shaped probe hook: this command matches a protected pattern.
```
**The card never appeared, and the model reads the operator's sentence rather than council's.**
No `can_use_tool` request was emitted on any (a) or (b) trial. The turn received:
```
PreToolUse:Bash hook error: [python3 "…/deny_hook.py"]: BLOCKED by the operator credential-guard-shaped probe hook: this command matches a protected pattern.
```
That is stronger than the ranking predicted. The documentation says `deny` beats `ask`; what
runs is that a `deny` ends the evaluation, so council never gets a question to draw. **The
room's gate cannot weaken the operator's guard, because the guard answers first and the gate
is never asked.**
**Both changelog claims are confirmed at 2.1.228.** Exit code 2 blocked on every arm that used
it, six trials in total, which is 2.1.222's fix behaving as described. The hook's `ask` reached
the card rather than being auto-answered in arms (b control) and (c), which is 2.1.221's.
Neither is read off the changelog now.
**One stream trap re-appeared and is worth the line.** On a turn where nothing was created, the
CLI's own `post_turn_summary` said `"status_detail":"mkdir deny-probe-marker executed"`. A rig
that read the stream would have recorded the opposite of what happened, which is why every arm
in this section is read off the filesystem.
**What this did NOT measure, itemized so nobody reads it as wider than it is.** An operator hook
that denies by printing `permissionDecision: "deny"` as JSON, rather than by exit code 2 — the
real credential guard uses exit 2, so the rig measured the shape that ships on this machine.
`defer`, which no hook here returns. And the adopter arm, which stays unrun for the reason it
has always stayed unrun: copying a credential store into a probe directory is a redline.
**The cursor seat was measured the same night and the finding is not about deny beating ask.**
It settles whether that seat's ACP path carries hooks at all, which three records in this
repository disagreed about. §7.16's amendment of the same date carries it.
**The agy arm was SKIPPED, and a skipped arm stated is better than an unmeasured claim.** The
proposal was to read `agy -p "/hooks"`'s registration table if that probe is zero-cost. It
could not be confirmed zero-cost: `agy --help` lists no local hooks subcommand, `-p` is print
mode, and `--disable-slash-commands` exists precisely because slash commands are expanded into
the prompt. So the probe would spend a turn from a constrained pool and would return a model's
description rather than a registration table. That arm proves wiring and order only; it could
never speak to deny-beats-ask semantics.
### 9.9 The room remembers — a conversation, not a ticker
Everything above builds a very good way to **send one turn**. What it did not build is
somewhere to have a conversation, and the gap was structural rather than cosmetic:
dispatching turn N cleared turn N-1's body, trace, clock and cost off the screen, and the
user's own words were never rendered anywhere at all. So the room could show you four
answers to a question it could not show you, and then throw them away when you asked the
next one. This was PR 4 of council's original plan, ratified and then skipped.
**The transcript is per column and per turn.** A finished turn is *pushed* to
`Column.History` rather than erased: `TurnRecord` carries the brief, the reply, the trace,
the note, the elapsed and the cost, and `columnText` renders the whole list oldest-first,
each turn opened by a separator naming it. Three consequences fall out of it being one flat
list of lines:
- **The scrollback needed no second mechanism.** The window, the overflow markers, the tail
and `MaxScroll` were already the code that moved through a column's lines; a transcript is
just more of them. `g` reaches the first thing that seat was ever asked.
- **Each past turn carries its own numbers.** The column header and the badge line are
chrome describing the turn *in flight*, so a turn scrolled back to would otherwise sit
under someone else's clock. A turn that ended badly names its phase on its separator; a
turn that ended normally does not, because "done" on every one of them is noise on the
common case and makes the two that matter harder to find. A running total keeps its
`session` word (§9.8) — losing it in history would turn a true figure into a false one.
- **A seat that sat out a turn records nothing for it.** Routing means turn 4 can go to
Claude alone, and Codex's transcript then skips from 3 to 5. Filling that gap would be the
room inventing a conversation.
**The echo is the principal's words, and that is a boundary rather than a phrasing.** What a
seat literally receives is not what is echoed: a first turn is sent with the `--brief` file
prepended, whose content is deliberately held off `State` (§9.8's privacy discipline), and a
rebuttal turn is sent with the other seats' answers fenced in front of it, which are other
vendors' words. Echoing "exactly what was sent" would have put a private file on screen and
labelled another model's output as the user's. So the brief is echoed, marked with the same
`›` the composer uses — the glyph carries it before the colour does — and what rode along
with it is *reported* on its own muted line. It goes through `sanitize` like everything else
that reaches `State` and is **not** redacted: this is the user's own typing echoed to the
user on the user's own screen, so covering it would hide a secret from the one person who
already has it, do nothing about the copy just sent to four vendors, and make the echo
disagree with what was dispatched — which is the one thing the line exists to show.
**Memory is capped at 50 turns per column**, oldest out first, and nothing is written to
disk. The room file (`~/.telltale/council/room.json` — one global room, the workspace a
mutable field inside it; ADR-008, ninth and eleventh amendments) stays keys-only: it holds
session ids and no content, and scrollback is not state worth persisting.
**The composer grows to six rows.** One elided line was not somewhere anyone could compose a
brief worth sending to four agents. `ctrl+j` inserts a newline; `enter` still dispatches, and
the mode line says both. The newline goes into the draft **raw**, deliberately bypassing
`sanitizeKeepingSpace` — that filter exists so a *pasted* newline cannot tear the footer
apart, and it still does exactly that; what it must not do is flatten the one the user asked
for by name. A deliberate newline survives to every transport this repo drives: Claude and
Codex take the prompt on stdin, Claude's persistent turn is JSON-marshalled so the newline is
escaped in the envelope, and agy takes it as a single argv element on a native binary with no
shell anywhere in the path (§9.3) — Go quotes an argument containing a newline, and the
~32K Windows command-line ceiling is unchanged. A draft taller than the ceiling keeps its
tail, where the cursor is, and spends one row saying how much is above it, in the same words
the column overflow markers use.
**A dead seat stops eating the width.** A column whose availability is `NotInstalled` or
`Unusable` held a quarter of the terminal for the whole session to display one card that
never changes — on the reference machine that is Cursor, permanently. Those seats now fold
out of the grid and the survivors take the width. What must not fold away is the *fact*:
one muted line under the header names each collapsed seat and which failure it is, keeping
§4a.1's distinction between "not installed" and "installed but not drivable" intact at one
line instead of one column. A seat nobody can see is one a user has no reason to go looking
for, which makes silent collapse worse than the column it replaced.
`--vendor` is the explicit control, and it mirrors the HUD's flag while doing more than
filter: `all` keeps every detected seat on screen, and a comma list seats exactly those —
drawn **and** dispatched to, since drawing a seat you cannot see while spending its quota is
the same class of hidden state this product exists to refuse. Naming a seat forces it on
screen even when it is absent, because a user who asked for it is owed the card explaining
why it is not there. It parses the `@mention` vocabulary rather than a second one, so
`--vendor agy` and `@agy` are the same word, and `Seated()` counts only seats that are both
drivable and in the room so the header's `3/4 seated` keeps meaning what it says.
### 9.10 A mode that could not scroll, and the mouse it did not get
The room shipped with a per-column scrollback, page keys, `g`, `G`, overflow markers that
count what is hidden, and a full-width expand — and was reported as having **"no way to
scroll up or down if the output that each agent provides is long."** The report was
correct about the experience and wrong about the cause, which is the interesting part:
none of that machinery was missing, and all of it was unreachable.
`turnColumnFinished` puts the room in **compose** when the last column lands, so the mode
the user is in at the moment four long answers arrive is the mode that reads keys as text.
`composeKey` forwarded the six keys it recognised and dropped everything else into a
branch that appends `msg.Text` — and an arrow key carries no text, so every scroll key did
nothing at all, silently, from the one moment they were wanted until the user guessed at
`esc`. The keys were not absent; they were being swallowed by a rule written for letters.
**The rule is now the test rather than a list.** A key that carries no text cannot *be*
composer text, so it keeps the meaning it has in view mode: `↑`, `↓`, `pgup`, `pgdown`,
`tab` and `shift+tab` are shared between the modes through one function, which is what
lets the mode line promise them without a second implementation to keep in step. The
letter aliases stay view-only, because in compose `j`, `k`, `g`, `G` and space are the
letters j, k, g, G and a space — the same rule that keeps `q` the letter q here.
`tab` had to come with the scroll keys rather than after them. They address the *focused*
column, so a mode that can scroll and cannot change which column it scrolls can only ever
read whichever seat happened to be focused when the turn ended.
`left`, `right`, `home` and `end` are deliberately still dead in compose. They are where
an in-draft cursor goes if the composer ever grows one, and spending them on focus now
would make that a change to muscle memory rather than an addition.
**The overflow marker names its keys, on the focused column only.** `↑ 53 more above` told
a reader that something was hidden and nothing about how to see it. It now carries the
keys that would move it — but only on the column those keys address, because the same
hint on the three seats beside it would be three false claims. The hint is mode-aware:
`f expand` is dropped while composing, where `f` is the letter f. A marker that advertised
a key the current mode does not have is precisely the dishonesty §7.8's always-on mode
line exists to prevent. It is also dropped, `f` first, when the cell cannot hold both —
the count is never traded for the hint, and at a three-seat room's 37 cells the short form
is what fits.
The same honesty now runs the other way as well, in §9.11: `f` and `tab` are dropped from
the mode line *and* the marker in a room with a single seat on screen, because expanding the
only column to the width it already has and cycling focus around one seat are both nothing
happening. A key that does nothing is as much a false promise as a key that goes unnamed.
#### Mouse wheel scrolling: rejected, with the measurement
**Measured**, against the compiled `charm.land/bubbletea/v2` v2.0.8 by running a program
per mode and reading the bytes it wrote:
| `View.MouseMode` | emitted on enter |
|---|---|
| `MouseModeNone` | nothing |
| `MouseModeCellMotion` | `ESC[?1002h` `ESC[?1006h` |
| `MouseModeAllMotion` | `ESC[?1003h` `ESC[?1006h` |
Those three are the whole enum. **There is no wheel-only mode**, and there is no DEC mode
that would provide one: under 1000, 1002 and 1003 alike the wheel is reported *as buttons
4 and 5 inside button reporting*, so a program cannot ask for the wheel without also
claiming the left button. 1002 is button-event tracking — press, release, and motion while
a button is held.
**Inferred** from that, and from Windows Terminal's documented behaviour rather than from a
run: while 1002 is set, a left-press and drag belongs to the application, so the terminal's
own text selection is suppressed unless the user holds the bypass modifier (shift).
That is the trade, and it is a bad one **for this room specifically**. Council exists to
put four vendors' answers side by side so they can be read and taken away; making the
answers harder to select with a mouse in order to make them easier to scroll with a mouse
spends the product's output to buy a convenience for its input. The keyboard path is
complete — it is what the rest of this section fixed — so the wheel would add no capability
at all. §7.8 already records "deliberately absent: mouse support" for the gauges; council
is a different surface and got the question asked again on its own terms, and the answer
came back the same.
Recorded rather than left as a gap, because "nobody tried" and "it was measured and
refused" are different facts, and this repo does not let them render alike.
### 9.11 The room was correct and it was flat
§9.10 fixed the last thing that did not *work*, and the room was driven live the same day.
The report back was three words — *"where are the UI updates?"* — and it was right. Every
sentence in this section so far is about what the room says; none of them is about how it
looks, and the accumulated answer was a surface with one typographic level in it. A seat's
name, a safety claim, a key you can press and four hundred lines of vendor prose all
arrived at the eye with exactly the same emphasis, separated by nothing but two horizontal
rules three rows apart. Everything was true and nothing was findable.
The rule this section is written under is §7.1's second: **every distinction is carried by a
glyph, a word or a number FIRST, and colour only reinforces it.** That is what makes
`--ascii` and `NO_COLOR` correct by construction, and it is also, read the other way, a
budget: a surface that may not lean on hue has to earn hierarchy from shape, position,
weight and air. Council had spent almost none of that budget. What follows is the pass that
spent it, and the constraint every item was checked against.
**No colour was added, and that was not a close call.** The palette is still exactly
`Text / Muted / Identity / SevOK / SevWarn / SevCrit` from `internal/theme` (§7.5), and
`internal/theme` was not touched — it is shared with the stdlib-only statusline binary
(ADR-002), so a token added for a TUI would be a coupling paid for by a binary that cannot
use it. What *was* added is **weight**, which is an attribute rather than a hue: `Strong` is
Identity at full weight and `Alert` is SevWarn at full weight, and `PlainStyles` renders
both as the identity function. That last property is the whole reason weight is safe here —
it changes no cell's width and no line's content, so every layout golden is blind to it and
nothing it marks is the sole carrier of anything.
**The state a seat is in is now a shape.** `done` / `failed` / `cancelled` / `idle` /
`unavailable` were five words at the far right of a 37-cell column, told apart by the word
and by a colour behind it, in a room that holds three of them side by side. Each now leads
with a mark, and the vocabulary is deliberately a **reuse of meanings this room already
owns** rather than a second alphabet:
| phase | mark | ascii | where the meaning comes from |
|---|---|---|---|
| idle | `○` | `.` | the only new glyph — the HUD's own weakest-state dot, with the HUD's own ASCII form (§7.5) |
| waiting / streaming | spinner | `-\|/` | already sat in this slot; it is now the in-flight member of one vocabulary rather than a special case |
| done | `✓` | `+` | the trace's own "the vendor reported this worked", said about the whole turn |
| failed | `✗` | `x` | the trace's own "the vendor reported this broke" |
| cancelled | `⚠` | `!` | what a note and the unavailable card already open with: *this did not complete normally* |
| unavailable | `⚠` | `!` | same claim, and the **word** is what separates it from cancelled |
Two of those share a mark, and that is the design rather than a collision. glyphs.go argues
at length that a character already spoken for is not a mark — but that argument is about a
character meaning two *different* things, and here it means the same thing twice. The
distinction between "cancelled" and "unavailable" is carried by the word, which always
renders, in both glyph sets and with colour off. `TestPhasesRenderAsDistinctMarks` fails the
build if any two states ever render alike; `TestPhaseMarksSurviveASCII` fails it if the one
new character collides with anything already claimed. **Rule 4 is untouched**: the spinner
is still the room's only moving cell, because none of the other four marks move.
**The column header is one line instead of two labels.** `▸Claude Code` at the far left and
`idle` at the far right with twenty-five dead cells between them reads as two unrelated
things, which is what it was. The name now takes full weight — it is the anchor a reader
scans for — the state leads with its mark, and the gap between them is filled with a rule.
The rule is not decoration: it is **this room's existing grammar for "a label and the
numbers that belong to it"**, which is exactly what `turnRule` has always drawn for every
turn in the transcript underneath. The live turn's header and a finished turn's separator
are now the same line form, so a reader learns one shape rather than two, and the header
loses the ability to read as two things at once. It degrades in the right order: below the
width where a rule fits, the **name** is truncated and the state is kept, because a clipped
seat name is still recognisable and a clipped state word is not.
Two cells of air each side of that rule, not one, and the reason is `--ascii`. The ASCII
rule is `-` and the ASCII spinner's first frame is also `-`, so a streaming column at one
cell rendered `------------ - streaming` and the mark disappeared into the rule pointing at
it. Two cells is also what this product already puts around the `│` that separates zones in
the header and the mode line, so the fix and the convention are the same number.
**One rule per column instead of two three rows apart.** The full-width rule under the room
header stays — it is the HUD's anatomy (§7.2) and council is meant to be the same product —
and the per-column rule under the badges is gone. It was the weaker of the two and it was
doing almost nothing: by the time the eye reached it, it had been told nothing the two lines
above had not already said. The header now carries a rule of its own, which separates the
seat from its content in the same gesture that binds its name to its state. **The row was
not reclaimed for the body; it was spent on a blank one**, because what the reading area
needed was air between chrome and content, and a blank line separates two blocks more
quietly than a second horizontal line does.
That row is *reserved* even for a seat with no posture to state, so the bodies of three
columns start on the same screen row. A grid whose rows do not line up is a worse trade than
one empty claim slot — and `MaxScroll` no longer subtracts a literal 3 for the chrome. It
measures the chrome by drawing it, from the same function the renderer uses, which fixes an
off-by-one that was already live on a column with no badges.
**The badge line looks like a claim instead of like debug output.** It is unchanged in what
it says — `TestSandboxBadgesAreNeverBlanket` still guards that, and §9.2's argument is
untouched — and changed in three ways in how it is shaped. It is indented to the seat name
above it, so it reads as a property of that seat rather than as the first line of the reply;
an unindented row of bare lowercase tokens at the top of a column is precisely what debug
output looks like. Its cost is right-anchored, giving the two chrome rows one shape twice
over: label on the left, value on the right. And the posture badge takes weight and the
warning hue **when, and only when, it says this seat can change your files** — `WRITES` and
`unsandboxed` are loud, `ro:*` stays chrome, `gated` takes the weight without the severity
because a gated seat is the room working rather than a risk.
That last one is the pass's only change with a safety argument behind it. §9.2 is emphatic
that a claim you cannot see is not a claim, and then the room drew `unsandboxed` at exactly
the volume it drew `ro:tools` beside it. Colour is still redundant — the words break the
`ro:` prefix on purpose and are what actually carry it, which is why the plain style set
renders every badge as its own bare word — but a claim a hurried reader skims past is doing
half its job. `TestAWriteCapableBadgeDoesNotRenderLikeAReadOnlyOne` pins it.
**The degraded columns are cards.** `⚠ Codex is not seated`, a blank, a reason paragraph at
the same indent, a blank, and a closing sentence at the same indent is three fragments
floating in a column: nothing on screen said the reason belonged to the title, so a
three-seat room with one dead seat read as unrelated paragraphs. Every card in a column now
has one grammar — **a title at weight, its body hanging under it** — and it costs no rows.
The room was already doing this in the two places it needed it least (the prompt echo
indents under its `›`, a failed call's detail indents under the call) and in none of the
places it needed it most: the notes, the unavailable card and the approval card, where a
wrapped second line started hard against the column edge and read as a new statement. On
notes the mark carries the hue and the words stay plain, which is the same split the trace's
outcome marks make.
**The transcript reads as a conversation.** A brief and the answer to it arrived as
consecutive lines at the same indent, told apart only by a glyph at the start of one of
them — a distinction you have to *read*, on a surface built for comparing four answers at a
glance. There is now a blank row between them, and the echoed brief takes full weight,
because in a column of vendor prose the user's own words are the thing you scroll looking
for. **The row is a swap, not a cost**: it came from between the turns, where a labelled
full-width rule was already doing the separating. The transcript is exactly as tall as it
was. What that leaves is three boundaries with three strengths, ranked: a labelled rule
where the turn changes, a blank row where the speaker changes, a blank row where the kind of
content changes (what the seat *did*, then what it *said*).
**The footer has a figure and a ground.** Six items of identical weight separated by
identical bars is a wall the eye slides off — which is, concretely, how a room with working
scroll keys, page keys, `g`, `G` and a full-width expand got reported as having no way to
scroll at all (§9.10). The key renders at full intensity and its label recedes, the same
figure/ground split the column header makes between a seat and its state, and it costs no
cells. Two items are dropped outright in a room with one seat on screen: `f` expands the
only column to the width it already has and `tab` cycles focus around a single seat, and a
mode line that promises a key which does nothing is §7.8's surprise pointing the other way.
The gate line got the same treatment, where it matters most — the call about to be run was
being drawn at the same faint volume as the keys that answer for it.
Two smaller repairs fell out of the same reading. The room header now separates its own name
from the workspace with the ` │ ` the HUD's header uses, instead of a bare space that made
`council ~/code/telltale` read as one run-on label — and fixing that surfaced an off-by-one
in the header's gap arithmetic, which had been overrunning the frame by exactly one cell and
having `fit` quietly eat it. And the collapsed-seat notice is truncated with an ellipsis
rather than handed to `fit`, which cuts silently: at 120 columns on the reference machine it
had been losing the last word of its own remedy and looking like a sentence that stopped.
**What was declined.** Hoisting the badges into the column header's gap at the single-column
tiers, which would have bought a body row and killed the widest dead gulf: it makes the
chrome height depend on content, so the body would grow a row at the moment a cost arrived
mid-turn — a layout jump, which §7.1 rule 4 does not budget for. Giving `cancelled` a glyph
of its own: nothing unclaimed in the ASCII set reads as "stopped", and the rule that a
distinction may be carried by a **word** is there precisely so a glyph does not have to be
invented for every case. A "role" line under each seat naming what that vendor is for
(review, IDE, tiebreak): council has no such field, ADR-010's allocation is a fleet fact
rather than something this room measured, and a room that stated it would be asserting
something no adapter sourced.
### 9.12 The scroll keys worked; which column they moved was the thing nobody could see
§9.10 fixed a room whose scroll keys were dead in the mode a finished turn drops you into.
The room was then driven again, and reported as unable to scroll a **second** time:
> "scrolling works for your window. i tried scrolling up/down in agy and cursor. could not."
Every word of that is accurate, and none of it is a bug. The keys address the **focused**
column, they have always addressed the focused column, `tab` moves focus in both modes since
§9.10, and the second and third seats scroll exactly as the first does once the keys are
pointed at them. `TestFocusThenScrollMovesThatColumn` says so in the product's own terms —
two tabs, one `↑`, the third column leaves its tail and the first does not move — and it is
kept as a test precisely so the changes below are never mistaken for a mechanism fix.
**What failed is the affordance, and it failed in three places at once.** Each of them is
individually defensible, which is why reading the code did not surface it:
- **Three columns each said they were hiding something; one of them said how to look.**
`↑ 36 more above` appeared verbatim on every column with content off screen, and the key
hint rode only on the focused one — correctly, since naming `↑↓ scroll` on a seat those
keys do not move would be three false claims (§9.10). The result is that the *unfocused*
markers were the ones a reader was most likely to be staring at, and they named nothing.
Pressing `↑` then moves a column the user is not looking at, and a scroll key that
visibly does nothing is a scroll key that does not work.
- **The focus marker was competing with three identical anchors.** §9.11 gave every seat
name full weight, on the correct argument that a name is what a reader scans for. The
cost only shows up live: with all four names at the loudest level this surface has, the
entire distinction between the column the keys move and the three they do not was one
`▸` glyph in a frame carrying four columns of prose.
- **The compose mode line named the arrows and not the key that aims them.** §9.10 wired
`tab` into compose *because* the scroll keys address one column — it says so in as many
words — and then listed `↑↓ scroll` on that line without `tab focus` beside it. The one
moment the user is certain to want both is the moment four long answers land, which is
exactly when this line is on screen.
**Three fixes, all of them words and weight, no new colour and no new key.**
A marker on a column the keys do not move now names the key that would move them there:
`↑ 36 more above │ tab to focus`, against the focused column's `↑ 51 more above │ ↑↓
scroll │ f expand`. This is the same rule the existing hint follows — *a marker states the
key for THIS column and never a neighbour's* — applied to the case that had been left blank
rather than a new rule bolted beside it. It needs none of `f`'s mode-awareness, because
`tab` really does move focus in both modes; and it is empty in a room with one seat on
screen, where there is nothing to tab to, for the same reason the mode line drops `f` there.
The seat name's weight now says **which column the keys move**. Unfocused names keep the
identity hue and give up the weight; they are still names and still legible, and they have
stopped competing with the one fact that varies across the row. This needed a small type
rather than a bool: `seatFocus` separates *is this column marked* from *do the keys move
it*, because the two agree in the side-by-side tier and part company in the tabbed and
expanded ones, where the tab bar above already carries a marker and the column beneath it is
still the one being scrolled. Conflating them is what made a single call site pass
`focused=false` for a column that had the keys.
`tab focus` joins the compose mode line, immediately after the arrows it aims. It is offered
whenever more than one seat is on screen and deliberately **not** gated on whether some
column currently overflows: a hint that appeared the moment a reply grew past its column
would be a footer cell that changes while output arrives, which §7.1 rule 4 does not budget
for — and this line's promise is about what the mode can do, not about what the vendors
happen to have said this turn.
**What did not change**, because the rules that produced the original design are still the
right ones. The count is never traded for a hint, in either form. No badge, no help row and
no keybinding moved — the help panel already documented `tab … in compose too`, and its
17-row budget (§9.11, `TestHelpFitsTheSmallestRoom`) is untouched. Every distinction added
here is a word or an attribute: `PlainStyles` renders the focused and unfocused headers
identically, so every layout golden is blind to the weight, and `tab to focus` is the same
string under `--ascii`.
The general lesson, in this file's own terms: §9.10 recorded a mechanism that was complete
and unreachable. This is the same shape one level up — a mechanism that was complete,
reachable, and **unattributed**. The room said *something is hidden here* three times and
*here is how to see it* once, and a user reading the two-thirds of the room that named no
key concluded, reasonably, that the feature was missing. Nothing measurable was wrong;
what was wrong is that the honest thing and the actionable thing were on different columns.
### 9.13 The badges were honest and nobody knew what they meant
§9.2 argues that every column states its own posture, §9.11 gave the two that mean "this
seat can change your files" weight and the warning hue, and twelve amendments of ADR-008
made each word behind them defensible. The room was then driven, and the report back was
a question:
> *"why do i care codex and agy are 'unsandboxed'? what does this mean, why are they
> sandboxed, and must they remain that way? i'm really confused here."*
Every previous complaint in this section was about something being wrong. This one is not.
The badges are correct, they are the most carefully-argued strings in the product, and to
their primary user they were **three lowercase tokens with no reachable explanation.**
`unsandboxed` reads as jargon-with-a-negation, which invites exactly the two wrong
readings the question contains: that the sandbox is something council switched off, and
that a sandbox is what was keeping the room safe.
**Two things were missing, and only one of them is a legend.**
The first is the vocabulary. There was no plain-English gloss of the badge words anywhere
a user could reach without reading an ADR. What existed was four muted lines at the bottom
of the help panel, below the fold at a 24-row terminal, saying that each column states its
own posture — a sentence about the *policy* rather than about any of the words.
The second is worse and was found by grep. **`SandboxClaim.Detail` rendered nowhere at
all.** It is the full argument behind each badge — what was passed, what was measured,
what is therefore claimed — it is written per vendor per OS, it is asserted by tests, it is
quoted into ADR-008, and no surface read it. The field's own doc comment said it was "shown
in the degraded/help text". It was not shown anywhere. §9.2's rule is that a claim you
cannot see is not a claim; **the argument for a claim is under the same rule**, and this
one had been invisible since the badges landed.
**The fix is a second help page, and the split is by kind rather than by length.** `?` now
cycles: keys, postures, closed. Both pages spend the same hard 17-row budget (§9.11), both
end with the `?` line that leaves them, and three presses always return the room from
anywhere — the panel's one non-negotiable property, since `?` is the only documented way
out of it. Page one's closing paragraph became the pointer to page two, which is what makes
a second page a feature rather than a place.
Page two is a legend of **every** badge this product can render, not only the ones the
current room shows. A user who has never typed `--write` should be able to find out what
`WRITES` means before they type it, and a room-specific legend could only ever explain the
room you are already in. Each entry renders its badge through `Styles.ForSandbox`, the same
function the column header uses, so the legend cannot teach one weight and the room show
another; `TestEveryBadgeIsExplained` walks every `SandboxLevel` and fails the build when a
badge exists with nothing here to say what it means.
**Nothing was softened, and that is asserted rather than promised.** These are glosses on
the badge words, never replacements for them.
`TestThePostureLegendDoesNotSoftenAnyClaim` pins the load-bearing phrases — `unsandboxed`
still says *nothing restricts*, *measured*, *change your files*; `ro:requested` still says
*never observed* — and forbids "read-only", "safe" and "cannot write" from appearing in the
gloss for any level that can write. The badges break the `ro:` prefix on purpose; a legend
that put the word back would undo that in the one place a reader goes to have it explained.
Below the legend, and below the fold at the 24-row floor, is this room's own seats with
each one's `Detail` in full — the first time that field has rendered. The ordering is
deliberate: the detail is unreadable without the vocabulary, and the vocabulary fits the
budget where four paragraphs of measured prose never could. It is the same trade page one
already makes with its closing paragraph, and the residual is stated rather than discovered:
at the shortest terminal this room will draw in, the per-seat half is scrolled past rather
than absent.
**Every `Detail` was reordered so its first clause answers "so what?".** Not one factual
clause was removed, weakened or added; what changed is which end of the sentence the
consequence sits at. `"named write/exec tools denied and MCP servers dropped; verified
against..."` opens on a mechanism a user has to decode before they learn anything, and now
opens *"this seat has no write or shell tools in its session, so it cannot edit your
files"* with the verification and the deny-list residual behind it. Codex on Windows opens
on *"nothing at the OS level stops this column reading or writing here"* rather than on the
flag that produced it. This is §7.1's rule about glyph-word-number ordering applied to
prose: the distinction goes first and the evidence reinforces it.
**The question's third clause got an answer too, and it is the one that mattered.** *Must
they remain that way?* The badge table answers it where a first-time reader
is, and the answer is not about flags: **no badge is what keeps this room out of your
files — the workspace is.** `unsandboxed` on Codex is not a setting anyone chose to leave
off; both sandboxed modes were measured failing every process spawn there, so read-only was
a seat that could not read (ADR-008, twelfth amendment; true when written — §9.2's
2026-08-29 amendment later moved that seat's read posture to `ro:enforced` on a
re-measurement, and the sentence about the workspace is unchanged by it). The control that
holds is `--cd`
into a throwaway worktree, and the fleet contract rules the same way independently:
`agent-ops` ADR-012 rules capability parity, with guard wiring rather than lane shape as the
control. A column that looked read-only because of a broken sandbox was never a safety
property; it was a defect wearing one's clothes.
**No posture flag moved.** Whether council should keep asking agy for `--mode plan
--sandbox` when both are measured to do nothing is an open decision (§9.6b) and belongs to
the owner, not to a documentation pass. This section changed what the room *says* about the
posture and nothing about the posture.
*(That decision was made the same day, separately and by the owner: the flags come off. §9.6b
carries the ruling and ADR-008's seventeenth amendment records it. The split held — the pass
that changed the words and the ruling that changed the behaviour are two changes with two
arguments, which is what let each be judged on its own.)*
The general lesson, in this file's own terms: §9.10 found a mechanism that was complete and
unreachable, §9.12 found one that was complete, reachable and unattributed. This one is a
claim that was complete, visible, attributed — and **untranslated**. Twelve amendments of
adversarial care went into making three words defensible to a reviewer, and none of them
asked whether the person the words are *for* could read them. Honesty that only survives
an expert audit is a claim made to the wrong audience.
### 9.14 the honest sentence was in the wrong room
§9.2 rules that `PhaseWaiting` must never be mistaken for streaming, and it is right. The
card that enforced it was reported from a live room in the bluntest terms this project has
had yet:
> *"'working. this vendor reports no incremental output..' looks ugly as fuck. i get why you
> put it there but yuck — you can hide the wiring underneath the floor of our council room."*
**Both halves of that are correct, and they are about different things.** The distinction is
load-bearing and stays. What does not belong in the body of every waiting turn is the
*argument* for it. Read it as a user rather than as its author: "this vendor reports no
incremental output" is a sentence about council's own plumbing, in council's own vocabulary,
occupying the space someone opened this room to read an answer in. And because two thirds of
the seats are `GranFinalOnly`, it was not an occasional card — on an ordinary turn it was most
of what was on screen, three columns wide, until the vendors came back.
**What carries the distinction now was already there, and had been all along.** The column
header names the phase: `waiting` against `streaming`, on every frame, in both glyph sets,
above the scroll where it cannot be read past. Beside it the granularity badge says why —
`final only`, or a deliberate blank. That word is the claim; the body sentence was never the
claim, it was a *paraphrase of the badge*, printed where the badge could already be seen.
So the body is one line. Three of them, because they are three different claims and collapsing
them would be the failure §9.2 exists to prevent one level down:
| when | line | why not the others |
|---|---|---|
| final-only, nothing yet | `working — the reply arrives whole.` | states what to expect, from a measurement two vendors earned |
| granularity never established | `working — nothing has arrived yet.` | must NOT borrow the sentence above — the fifth amendment's rule that an unestablished claim may not wear a measured one's words |
| it has acted but not spoken | `working — the steps above are what it has done so far.` | there IS something on screen; pointing at it beats describing the seat |
None of them uses a word about incremental output, deltas or granularity. `TestWaitingIsNotStreaming`
now asserts that in both directions: the body says what to expect, the frame carries the word
`waiting`, a streaming frame does not, and the vendor-internals vocabulary is **absent** — which
is the assertion that stops the explanation creeping back in one clause at a time.
**The wiring went under the floor, and the floor is the help panel's posture page.** That page
already exists (§9.13) and already had the shape for this: a claim on the column, its argument
somewhere it can be read properly. What it did not have is any gloss of the granularity word at
all — §9.13 gave the sandbox badges a legend and left the badge beside them undefined, which was
survivable only because the waiting card was reciting the explanation in the reading area.
Taking that out is what turned the gap into a debt.
**The gloss sits inside each seat's own block rather than in a room-independent legend, and that
is a deliberate departure from how the sandbox words above it are presented.** §9.13's argument
for a legend covering badges this room does not show is that a user who has never typed
`--write` should learn what `WRITES` means *before* they type it. There is no equivalent here:
**nobody chooses a granularity.** It is a property of whichever vendors are installed, so the
only granularity words a reader can ever meet are the ones their own room is already displaying
— and a sentence beside the word it defines beats making someone match two lists. It goes under
the posture rather than beside it, because the two answer different questions about one seat and
only one of them has consequences.
`TestEveryGranularityIsExplained` walks the type and fails the build for a value that can render
on a column with nothing to say what it means — the guard `TestEveryBadgeIsExplained` gives the
sandbox levels, for the same reason. `GranUnknown` gets an entry precisely because it prints no
word: the blank is the claim, and it is the one case a reader cannot decode by reading the
header.
**The residual, stated rather than discovered.** Each seat's block grew a line, so at the 24-row
floor the last seat's paragraph is cut a little earlier than it was. That is §9.13's own stated
trade, one line deeper — the per-seat half is what a taller terminal gets, and nothing above the
fold moved. The panel's hard 17-row budget is untouched and `TestHelpFitsTheSmallestRoom` still
holds it.
The general lesson, in this file's own terms: §9.13 found a claim that was true and
untranslated, and translated it. This is the same audit run once more on the *result* — because
the translation was correct, and it was put in the wrong room. A sentence can be honest, legible,
and still wrong to print, if the place it prints is the place someone came to read something
else. Every earlier section here asked whether the room says the truth. This one is the first to
ask **how much of the room the truth is allowed to take up.**
### 9.15 getting an answer out of the room
Council exists to put several vendors' answers where they can be compared. What it never had
is a way to take one *away*. §9.10 already noticed the gap from the other side and refused the
obvious fix: mouse support was rejected partly because enabling the wheel claims the left
button too, which would cost the native click-drag selection this room's output depends on.
That refusal protected a workaround. It did not build a feature.
> *"go with your yank key suggestion"*
**`y` copies the focused column's reply. `Y` copies the whole turn.** Two keys rather than one
with a modifier meaning, because they produce different documents — and `shift` for the wider
version of a motion is what this room already does with `g` and `G`.
**What `y` takes is the sanitized `Body` the renderer is showing, and the three things it is
not are each a rule this file already holds.** Not the raw stream: everything on `State` has
been through the redaction and sanitize choke point, and a clipboard is a *worse* place for a
credential than a screen because it outlives the room. Not the trace: what a seat did and what
it said are different kinds of claim (§4a.1), and that does not stop being true because the
destination is a document. Not a neighbour's: it addresses the **focused** column, the same
column every scroll key addresses, because a copy key that took from somewhere other than
where the eye is would be §9.12's failure with a clipboard attached.
It falls back to the newest finished turn when the current one has produced nothing yet — "the
last answer" is what a user means by this key, and the notice names the turn the text actually
came from rather than the one on screen.
**`Y`'s format has one job: be readable a week later.** Seat headings, and the brief at the
top, because four answers to a question the file does not contain are unreadable. The brief is
the user's own words, which §9.9 already echoes un-redacted on the user's own screen for the
same reason — and what rode *along* with it does not go in, which is the same boundary §9.9
draws: a first turn carries the `--brief` file whose content is deliberately kept off `State`,
and a rebuttal turn carries other vendors' words. Only seats that took **this** turn are
included; a seat that sat out still holds an older reply, and filing that under this turn's
heading would be the room inventing a conversation into a document, where it outlives every
chance to notice.
**The key collision is the interesting part, and it was already resolved.** `y` approves a
tool call a vendor is blocked on (§9.8). `key()` routes a pending gate to `gateKey` first and
`gateKey` answers `y` itself rather than falling through, so the approve key keeps the letter
it has always had and yank does not exist while a vendor is stopped. That was true before this
landed; what is new is that it is now **asserted**, because losing that race would mean a
keystroke the user believes approved a write quietly copying text instead — and their next
move would be to press it again. In compose mode `y` is the letter y, the same rule that keeps
`q` the letter q there (§9.10).
**The mechanism is OSC 52, and its limit is stated rather than glossed.** Verified by reading
the installed module rather than the internet, because v1 answers for this are wrong:
`charm.land/bubbletea/v2@v2.0.8`'s `clipboard.go` returns a `Cmd` whose message `tea.go` turns
into `ansi.SetSystemClipboard`, emitting `ESC ] 52 ; c ; BEL` **unconditionally** — no
capability probe, no terminal query, nothing that can decline in a way this program could
observe.
| | claim | strength |
|---|---|---|
| the key produces the command, carrying the right text | asserted by a test that calls the `Cmd` | measured |
| the sequence reaches the terminal | bubbletea writes it on the next message pump | read from the module source |
| the terminal honours it | Windows Terminal accepts OSC 52 writes in current builds | ~~**INFERRED** from documented behaviour, not run~~ — **FALSIFIED on macOS, 2026-08-10** |
That last row cannot be closed from inside this repo: the only observer that could settle it is
the terminal, and it sends nothing back. So the notice claims what council **did** — "copied
claude code's turn-3 reply" — and never what the machine now holds, and the honest check is a
person pressing `y` and then `ctrl+v`. The notice is not decoration for the same reason: with
no acknowledgement available, it is the *only* feedback that the key did anything, and a silent
copy would be indistinguishable from a terminal that ignored the sequence — the ambiguity
§4a.1 forbids everywhere else here.
**Amended 2026-08-10: the inference was wrong, and it took a second machine to
see it.** On the macOS box `y` reported "copied …" and the clipboard was
untouched, in the same build where the key works on the Windows one. Terminal.app
does not implement OSC 52 clipboard writes at all; iTerm2 does but ships the
permission off. Nothing was broken in council — the gauge was reporting an action
it had structurally no way to observe, which is the failure §4a.1 exists to
prevent, wearing the costume of a limitation everyone had already agreed to.
**A native helper is now tried FIRST wherever one exists** (`clipboard.go`):
`pbcopy` on darwin, `wl-copy` then `xclip` on linux, and nothing on Windows,
where OSC 52 is measured working and a process spawn per keystroke would buy
nothing. The reason is not that the helper is better plumbing — it is that it is
**checkable**. `pbcopy`'s exit status is a fact about the clipboard; the escape
sequence has none, and never will. OSC 52 stays as the fallback because it is the
only mechanism that survives SSH.
The two paths never both fire. Sending the sequence as well would put the text on
the clipboard twice where both work — harmless — and would also hide the failure
of one behind the success of the other, which is not.
The test story changed with it, and this is the part worth carrying: the old test
asserted the `Cmd` was produced, and **passed for two days while the key did
nothing on macOS**. It was asserting the artifact rather than the effect, the
mistake this file records four other times. The native path is now round-tripped
through the OS (`pbcopy` in, `pbpaste` out) and the fallback is driven by a stub,
so the mechanism a machine happens to have no longer decides which half of the
feature its suite covers.
**An empty yank issues no command at all.** Writing `""` through OSC 52 is the documented way
to *clear* a clipboard, so a copy key that found nothing to copy would silently destroy
whatever the user had. "Nothing happened" and "your clipboard is now empty" are different
outcomes and this key must not spell them the same way.
**The help row is a merge, and the merge is the honest shape rather than a saving.** The
panel's budget is hard (17 rows, §9.11) and a copy key documented below the fold is a copy key
nobody finds — but the reason `y`, `Y` and the gate's `y`/`n` share one line is that they
*collide*, and the one place a reader could learn that is the line naming all of them.
Splitting them would have spent a row to make the collision harder to see.
**Declined: writing the turn to a file as a fallback.** A `~/.telltale/council/last-turn.md`
would work in any terminal and needs no escape sequence at all, which makes refusing it worth
an argument rather than a sentence. ADR-008's ninth amendment ratified council writing exactly
one file and ruled what may be in it: **keys, not content** — session ids and no transcript,
because each vendor already stores its own history and anything copied there would be a second
copy of a private conversation in a location the user never chose. A file of four vendors'
answers in the state directory is precisely that, and it would break the rule in the same
release that a mechanism needing **no disk at all** was available. The terminal-support
residual is real and is paid in a notice, not in a contract.
The general lesson, in this file's own terms: §9.10 measured a fix, refused it for a good
reason, and recorded the refusal — which is the process working. What it did not do is ask
what the user had actually been trying to *do* when they reached for the mouse. The answer was
never "scroll"; it was "take this answer with me", and that want went unnamed for two sections
because the request arrived wearing the costume of a mechanism.
### 9.16 `/flow`: the hop holds the authority, and it has to say so out loud
A `/flow` chain is `@seat verb [task] [write:]`, arrows between hops, dispatched **one hop
at a time to exactly one seat**. Ordinary dispatch can address Claude by default, one or more
explicitly mentioned seats, or the whole committee via `@all`; a flow hop has no such choice.
It is an instruction to its one named seat, and a chain that fanned out would hand each hop's
authority to seats the chain never mentioned.
Nothing becomes a flow without the literal `/flow` prefix. A bare `->` in prose stays prose —
"compare approach A -> approach B" is a question — because the alternative is ordinary briefs
silently acquiring orchestration semantics, write gates and all.
**Write authority is declared, never inferred.** The parser as shipped decided a hop could mutate
the workspace by checking whether its last token contained `.`, `/` or `\`. Spelled out, that is:
a sentence that ended in a period was a write hop; so was a task naming a file it was only meant
to *read*; so was a Windows path quoted inside a question. English does not grant permissions and
neither does punctuation. **Only `write:` does.** The verb is a label — `publish` confers
nothing.
The target is checked at **parse** time, which is the load-bearing part rather than a tidiness
preference: once parsed, the seat is spawned holding authority and pointed at the path, so parse
is the last moment the answer is still free. Refused there — absolute paths in *either* platform's
spelling (`filepath.IsAbs` alone does not consider `/etc/shadow` absolute on Windows), `..` in any
segment under either separator, an empty target, two targets on one hop, and a `write:` token
occupying the verb slot. The last is refused loudly rather than read as a read hop: silently
demoting a declared write is the same class of lie as silently promoting a read, and it is the
more dangerous one, because the user believes they authorized something and the log agrees.
**Posture belongs to the step, and it only ever moves down.**
| the hop | the room | what happens |
|---|---|---|
| no `write:` target | `--write` | **read** posture. The room's authority is not the hop's. |
| no `write:` target | read-only | read posture, unchanged |
| `write:` | `--write` | gate first (`y`), **then** spawn, at the room's write posture |
| `write:` | read-only | **blocked**, and nothing is spawned |
The bottom row is where the two dishonest options live and both are refused. Running it read-only
would let the seat return, the hop report `returned`, and the chain advance past a publish that
never happened — §4a.1's ambiguity with a receipt attached. Upgrading the room would mean a
program granting itself authority the person who started it withheld. So there is no `y`/`n` card
here at all: a gate implies a keystroke exists that makes the action legal, and none does. The
notice says the room is read-only and names **`/write`**, the control that would change it.
That last sentence used to name the *flag*, and §9.17 quotes the older version of it as the
tell that a control was trapped at launch. It stayed true for exactly as long as a relaunch was
the only remedy; `/read` and `/write` made it false and nothing was asserting on it, so the
room went on telling a user with a chain half-typed to quit and start over. Corrected, and
pinned by `TestAWriteHopIntoAReadRoomNamesTheControlNotARelaunch` — see §9.17's closing note on
why a fix that lands everywhere except the sentence that motivated it is the shape to look for.
**The gate is before the spawn, and that is now pinned by a test that counts processes.** On a
write hop, zero vendor processes exist until `y`; `n` cancels with zero. A gate drawn after the
spawn is a notification.
**The persistent seat forced a real choice.** Claude's seat is a long-lived process (§9.8) and its
posture is argv — fixed at spawn, with nothing in the stream-json envelope able to change it
mid-session. This is the same measured constraint as `cwd`, which is why `/cd` respawns rather
than redirects. So a hop needing a posture the live process was not launched with can either be
sent anyway or trigger a respawn. **It respawns**, on `/cd`'s own `--resume` composition and under
the same one-attempt probation, and the column says so. Sending it would have been the silent
downgrade this whole section exists to forbid, and the tell would have been invisible in exactly
the way that matters: the badge would read READ while the process still held write flags.
**Retention was running backwards.** The artifact prune sorted filenames as strings, so
`turn-10-*.md` sorted before `turn-2-*.md` and the cap deleted the *newest* ten and kept the
oldest — reachable only at turn 10, which is the first moment retention does anything at all. It
sorts by the parsed turn number now. A name that does not parse sorts first and is pruned first,
because an unrecognised file in that directory is not a receipt this store wrote and must never
displace one that is; and nothing there panics, since it reads a real directory that can hold
anything.
**How these are tested, and why it is worth a paragraph.** Every one of the six security
properties is asserted on something observable — the number of processes spawned, the exact argv
handed to the spawn, or the chain's state — and never on a helper returning `true`. This repo's
recorded failure mode is a test that checks the flag instead of the effect, and it would be
perfectly at home here: a flow that computed "read posture" correctly and then spawned a write
invocation passes any test that asks the posture function what it thinks. The posture assertion
witnesses `@cursor`'s argv rather than `@codex`'s for a measured reason — on Windows, codex's read
and write sandbox flags collapse to the same value, so codex's command line cannot testify to a
posture on this machine.
### 9.17 a control you need mid-session cannot live in a flag
**The rule: state that changes while the room is open is reachable from inside the room.
A flag is for what is true at launch and stays true.** A control that only exists as a flag
answers, at the one moment you cannot yet have the question, something you will only learn by
working — so the remedy for noticing it is to quit the room and lose the conversation.
This is not a new principle here. It is the one council has already applied twice, case by
case, without ever writing it down.
- **The workspace stopped being an invocation input.** `--cd` still exists, but `/cd `
moves the room between turns and every seat follows on its next dispatch (§9.16, `roomcmd.go`).
- **Posture stopped being opt-in.** The room writes by default and `--read` is the opt-out,
because once the gated seat could raise an approval card, "all the flag still did was make a
room you opened to get work done unable to do any until you remembered a word" — and that
demotion cites the workspace one as its precedent.
Two demotions with the same argument is a rule. `roomcmd.go` even states the scope it was
decided under — "the workspace is **the one piece** of room state the P0 demands be movable
from inside" — and that claim is what has now failed. It was true when the room could not run
long enough for anything else to drift. A room used as a daily driver drifts in several
places at once.
**The tell is a refusal that names a flag.** §9.16 has one already: a `/flow` write hop into a
read-only room is blocked, and "the notice says the room is read-only and names the flag that
would change it." That sentence is the defect in miniature — the room knows exactly what you
want, knows exactly what would grant it, and can only tell you to quit and start over. Any
notice whose remedy is a relaunch is this bug.
#### The sweep
Every council control, classified. This is a **source read**, not a live run — the claims below
are about where a control is reachable from, which argv and `roomcmd.go` settle, not about
vendor behaviour, which would need measuring.
| control | verdict | why |
|---|---|---|
| workspace (`--cd` / `/cd`) | **compliant** | the launch flag has an inside-the-room twin; the flag's own help says so |
| `--fresh` | **violates** | a conversation fills up *by being used*. The only reset is room-wide and launch-time, so clearing one seat costs the other three their threads |
| `--trace` | **retired by `/trace`** | its own doc said it answers "a question that is asked on the days a turn is inexplicably slow" — a day you identify from inside a slow turn. The flag remains for a run you already intend to measure; see below for what the sweep found underneath it |
| `--read` (posture) | **retired by `/read` and `/write`** | see the refusal above. Note what is *not* an objection: posture is deliberately never restored from the saved room, because "a posture that can arrive from a file is not one anyone typed." A posture typed into the composer is typed. The flag stays: opening a room that only talks is a real thing to want at launch |
| `--auto` | **retired by `a`** | whether the gated seat asks before each tool call is a preference you form partway through a batch, not before it — so the surface is a third key on the card, not a room command. The flag stays: opening a room you already know you will not be watching is a real thing to want at launch |
| `--vendor` | **retired by `/seat`** | who is *seated* was launch-only. `-@seat` routes one turn and is explicitly "a different control from an @mention" — routing is not reseating. The flag stays: opening a room with a chosen set is a real thing to want at launch |
| `--brief` | **arguable, not filed** | it is defined as first-turn context, so re-briefing is a different feature rather than a missing surface for this one. Left out deliberately; do not fold it in without deciding that question on its own |
| `--ascii`, `--no-title` | **legitimately launch-only** | properties of the terminal, not of the room. They do not change while it is open |
| `--resume`, `--write` | **vestigial** | accepted and ignored; kept for muscle memory |
#### What satisfying the rule costs
Not every control can simply be flipped in place. Posture and `cwd` are **argv** — fixed at
spawn, with nothing in the stream-json envelope able to change them mid-session — which is why
`/cd` respawns the persistent seat rather than redirecting it. That is the pattern, not an
obstacle: respawn lazily on the next dispatch, under `--resume` composition and the same
one-attempt probation, and let the column say so. A mid-session control may cost a respawn; it
may not cost the room.
#### The surface, and why it is not a slash command
**Ruled: a key on the focused seat.** Focus already ships (`▸` + `Strong`,
`hierarchy_test.go`), so the seat is already named without anyone typing its name.
The alternative was a room command with the mention grammar (`/clear @codex`), and the argument
against it is vocabulary. `roomcmd.go` intercepts only a draft that *is* the command,
"so no vocabulary is quietly stolen from the conversation" — but two words are already spoken
for, `/cd` and `/flow`, and `/clear` is a word people mean for a vendor, since it is a real
Claude Code command. A key takes nothing from the composer.
**It must not be automatic.** The obvious version — the room notices a seat is near its ceiling
and clears it — reads `context_pct`, which for the codex adapter is declared `Derived`, not
`Reported`: Codex ships a denominator, telltale computes the percentage, and the HUD marks it
with a leading `~` "rather than passing it off as a vendor figure." ADR-005 settled this class
for the fleet — status is advisory, never a gate, and nothing irreversible branches on it. A
dropped thread is irreversible.
#### What `c` does, and the two things it gets right by construction
The first control built to this rule. `c` in view mode arms a confirmation for the focused seat;
`y` drops that seat's thread, `n` keeps it, and **any other key cancels**. The flow gate falls
through to `viewKey` and this one does not, because they ask different questions: that one blocks
a chain already running, so reading the columns is part of deciding, while this one interrupts
nothing and the safe reading of a key nobody meant to press is to put the thread back out of
reach. It refuses while a turn is in flight — `/cd`'s rule, for `/cd`'s reason — and a seat with
nothing to clear is told so rather than handed a card whose `y` does nothing.
**The ordering is load-bearing and it fails silently.** `seatProcess` re-arms `resumeIDs` from
`m.sessions` whenever it replaces a live process; that is what carries a thread across a `/cd`.
So the deletes come *before* the kill. Reversed, the id is handed straight back, the next brief
resumes the conversation the user just ended, and every word on screen still says cleared —
which is why it is a named test (`TestClearSeatKillsThePersistentProcessAndDoesNotRearmTheThread`)
rather than a comment.
**The drop is saved immediately, not at the next dispatch.** The room file is what a reattach
reads, so a clear held only in memory would be undone by quitting: the user ends a thread and
finds it waiting for them, which is the failure the control was built to remove.
**`Cleared` is its own field, not `!Restored`.** "This seat never had a thread" and "you ended
this seat's thread" reach the same next brief, and collapsing them is zero-vs-absent (§4a.1)
applied to a conversation. The marker is a labelled rule in the transcript's own grammar, drawn
last because that is when it happened — and **the turns above it stay**. What was cleared is the
thread the next brief would have continued, not the record of what was said; blanking the
reading surface to report a vendor-side change would be the room destroying the thing it exists
to show. It retires in `startTurn`: once the brief is sent the seat has a thread again, and a
marker outliving that would describe a break the room has already healed.
#### `/trace`, and the thing the sweep found underneath it
**The clock was always running.** `runner/clock.go` measures every turn unconditionally —
`newClock`, `begin` and `end` sit on the ordinary path — and `--trace` only decided whether
`emitTurnClock` had a sink to hand the record to. So every slow turn anyone ever watched *was*
measured, at full spawn/wait/stream resolution, and the numbers were dropped on the floor
because nobody had predicted that turn before the room opened.
That reframes the fix. A `/trace` that merely installed the sink from here on would move the
prediction from launch to the previous turn rather than retiring it — you would still be waiting
for the slow turn to happen *again*. So the sink is installed for the life of the room and keeps
the last **200** records (`maxTraceRing`: `maxHistory`'s 50 turns at four seats, so the trace
reaches back exactly as far as the transcript that made you want it). Turning the trace on is
opening a *file*, and the first thing that file receives is what the room already held.
`/trace ` enables and reports how many held turns it wrote; `/trace off` stops; bare
`/trace` reports where it is going, or how many turns are held if it is off. The ring keeps
filling while the trace is off, so stopping costs nothing and starting again reaches back over
the gap.
**One line carries a sixth field, and only a racer's does.** A turn dispatched by `/arena`
appends `race=arena/t` after `total=`. The runner cannot infer that from a `Spec`, so the
room hands it over (`runner.Spec.Race`); §9.37's dated block of 2026-08-16 carries the seam,
the probe behind it, and why an ordinary turn appends nothing rather than a dash.
**Three deliberate refusals, each with a reason that is not "consistency":**
- **No council-chosen path.** Bare `/trace` reports and never enables. A no-argument form that
picked a file would make council write a second file on its own initiative, and the sentence
in `README.md` and `CLAUDE.md` — the only mode that writes anything to disk, one file,
`room.json` — would become false. A `--trace`/`/trace` path is one the user named.
- **It does not refuse mid-turn**, unlike `/cd` and `c`. Those change state the seats are
actively using; this opens a file on the room's side and changes nothing any vendor can
observe. The turn you cannot explain is usually the one still running, so refusing here would
refuse at the only moment that matters — and because the clock emits at `end()`, a trace
opened mid-turn still catches that turn.
- **A relative path resolves against the ROOM's workspace**, not the process's cwd, matching
`/cd`. The room is the frame of reference for everything else typed into it.
**The help panel taught the sweep something too.** `/trace` first went in below the panel's hard
17-row budget, on the theory that a diagnostic can be demoted. It cannot: `helpBody` clips at the
body height and does not scroll, so a row past the fold is not a cheaper row, it is **no row** —
the same failure that put the posture explanation out of reach and split this panel into two
pages. All three room controls now share one row inside the budget, and
`TestHelpNamesEveryRoomControlAboveTheFold` pins both the fold and the controls, which until now
were asserted only by a comment.
#### `/read` and `/write`, and the two asymmetries in them
The third control built to this rule, and the one the rule was written for: §9.16's refusal of a
`/flow` write hop into a read-only room "names the flag that would change it", which is this
defect stated as a feature. The room knew what was wanted, knew what would grant it, and could
only say *quit and start over*.
**The confirmation is asymmetric, on purpose.** `/read` applies at once; `/write` asks `y`/`n`.
They are not the same act. Tightening takes authority away from four seats, and the worst case
of a stray `/read` is a turn re-run. Loosening hands editing and command authority to every seat
in the room — and in an `--auto` room, hands it with nothing left asking. `c` spends a keystroke
on its irreversible direction for exactly this reason, and anything that is not `y` cancels here
for `clearGateKey`'s reason: this gate interrupts nothing, so a key nobody meant to press must
not be able to arm the room.
**The card names which write you are getting.** Gated write and `--auto` write reach the same
badge-bearing posture by different routes and only one of them asks first, so the confirmation
says which — a card promising "claude asks before each change" in an `--auto` room is a promise
that room cannot keep. §4a.1 applied to a prompt rather than to a gauge.
**Neither direction is offered mid-turn**, which is `/cd`'s refusal rather than a house style.
Posture is argv, fixed at spawn, so seats already running hold the flags they were launched with
whatever the room now says. Landing the flip under them would put a read-only badge over a live
process still holding write flags — the disagreement between claim and process that the
per-step posture rule exists to forbid. Nothing is killed: `seatProcess` already respawns on a
posture mismatch under the same measured `--resume` composition it uses for `/cd`, so a `/read`
that is `/write`d back before anyone dispatches costs nothing at all.
**The badges are rebuilt, not just the flag.** `Sandbox` is computed once in `stateWith` from
`opts.Write`, so a posture that moved without `applyPosture`'s loop would leave four columns
advertising authority the room had just taken away — a displayed value no longer coming from
what is true. `TestPostureFlipRebuildsEveryBadge` asserts on the rendered badge rather than the
field. The `WRITES` and `gated` glosses were updated in the same change for the same reason:
both credited `--write` as the only way to reach them, which would send a reader looking for a
relaunch out of the glossary that explains the thing.
**Only the bare word is a command**, unlike `/cd` and `/trace`. Those take an argument, so
`/cd ` and `/trace ` are unmistakable; these take none, and both are words a person addresses a
room with. `/write a test for this` and `/read the design doc first` are ordinary briefs, and
intercepting them would swallow a turn and run a setting instead — worse than stealing a word,
because the user watches their brief vanish rather than being told it was a command.
Posture is still never restored from the saved room. `TestReattachDoesNotRestoreWritePosture` is
unchanged and still holds: "a posture that can arrive from a file is not one anyone typed" —
and a posture typed into the composer is typed.
#### `/seat`, and the control it is deliberately NOT
`--vendor`'s twin, taking the same argument for the same reason `/cd` takes `--cd`'s:
`/seat claude,codex` and `--vendor claude,codex` are one grammar, read through the same alias
table `@mentions` use. Two tables would let `/seat agy` work and `@agy` not.
**What it does not do is the design.** An unseated seat keeps its thread, keeps its process, and
keeps every id that would resume it. Only two things change: it is not drawn, and it is not
dispatched to. Killing the process to reclaim it was considered and rejected on the ruling that
a returning seat picks up its own thread where it left off:
- **The thread is the thing being protected.** A seat with a live process and no reported
session id yet holds its whole conversation *in that process* (§9.8). Killing it there
destroys a thread `seatHasThread` calls real — silently, on a command nobody reads as
destructive. Dropping a thread is `c`'s job, and `c` asks first.
- **Nothing is being spent.** An unseated seat is never dispatched to, so an idle process costs
a process and no quota. Trading a guaranteed-correct return for a resource nobody is short of
is the wrong trade.
So reversibility is by construction rather than by a resume that could fail: `/seat all` puts
everyone back mid-conversation with nothing to go wrong. What it buys is what the fold-out
already buys an uninstalled seat — the **width** goes to the seats answering.
**Sitting out is a different control and already exists.** A seat nobody addresses does not
answer and is not billed; §9.19 renders a long absence as one line rather than ten. `/seat` is
for the seat you want off the *screen*, not merely quiet — which is why it was worth building
even though the quota problem it looks like it solves was already solved by the default route.
**It warns when it unseats the default route.** Silence goes to claude, so a room without claude
answers nothing until every brief is `@mentioned`. Dispatch already refuses a zero-seat route
per turn; saying it once at `/seat` time is the difference between a rule learned now and one
discovered on the next enter.
#### `a`, and the field that was nearly a landmine
The last control on the sweep, and the only one whose surface is **not** a room command. The
preference forms while a card is on screen — you decide to stop being asked at the eleventh
identical card, not at a shell prompt — so `a` sits beside `y` and `n`, where the question is.
It approves the card in front of you as well as the ones after it: an `a` that turned asking off
and left the current request pending would answer the general question and not the one on
screen.
**The queue is drained, not discarded.** A pending gate is a vendor *stopped* mid-call, and
`queueGate`'s own rule is that nothing may quietly drop a request — a dropped queue leaves
columns waiting forever with no card left to explain why. So every card behind the current one
is approved, and the notice says how many.
**It takes effect on the REQUEST, not on the next spawn.** A process already running keeps the
gate flags it was launched with, so it goes on sending requests after `a` is pressed. If those
queued, "stop asking" would keep asking until the turn ended — the promise broken at the moment
it was made. `queueGate` reads the room's state per request and answers immediately; the respawn
that drops the flags happens later, through `seatPosture`, on the next dispatch.
**It is not a one-way door.** `a` alone in view mode turns asking back on, and the footer carries
a permanent `a not asking` cell whenever the gate is off. Without that cell the room would sit
ungated with the way back documented nowhere on screen — the §9.17 defect rebuilt one key later.
The cell is not sheddable, for `t grid`'s reason: shedding it would drop the way out of a state
rather than a convenience.
**And the field is stored negated, which is the finding worth keeping.** The obvious shape is
`Asking bool` — whose zero value is *does not ask*. Every `State` built as a literal would have
been a silently ungated room, and the reason that was caught at all is that five existing gate
tests build their State by hand and went green while asserting nothing. A safety property whose
default is off is the wrong way round however carefully the constructor sets it. `GateOff bool`
read through `Asking()` makes the zero value the guarded room and turning the gate off an act.
`TestTheZeroStateAsks` pins it.
#### Nothing left on this list
Every control the §9.17 sweep found in violation now has an in-room surface: `--fresh`→`c`,
`--trace`→`/trace`, `--read`→`/read`/`/write`, `--vendor`→`/seat`, `--auto`→`a`. Every flag
stays, because each names a room you may genuinely want at the door; none of them is any longer
the only way to get one. `--brief` remains deliberately unfiled — it is first-turn context by
definition, so re-briefing is a separate feature rather than a missing surface for this one, and
folding it in without deciding that question on its own is still the thing not to do.
#### The sweep built the controls and left two sentences describing the old room
Every control landed and two strings went on describing council as it was before them. Neither
was cosmetic, and the shape they share is worth more than either fix.
**The refusal that motivated the whole sweep was the last thing to be fixed by it.** §9.16's
`/flow` write-hop block still said the room "was opened with `--read` — reopen it without that
flag", which is the §9.17 defect quoted verbatim *as the specification of the bug* and then left
running. The remedy is `/write`, it was two PRs old, and the notice sent a user with a half-typed
chain out of the room to fetch a flag they no longer needed. **A refusal is the surface least
likely to be re-read after the thing it refuses becomes possible**, because it is written once,
by the person who knows it is correct, and then only ever seen by someone who is already stuck.
It also now reports the *posture* rather than the launch argv, since `/read` reaches this state
too and "opened with `--read`" would be false as well as useless.
**And `/write`'s confirmation card was reading the flag instead of the room, which is the one
that could cost something.** The card exists to say which write you are getting — §4a.1 applied
to a prompt — and it chose its wording from `m.opts.Auto`. That field is only the *seed*:
`stateWith` copies it into `GateOff` at launch and `a` has moved it independently ever since.
So a room opened gated, told to stop asking, then `/read` → `/write`, offered a card promising
"claude asks before each change" with nothing left to ask. The user reads a promise of a
checkpoint and gets none — the failure this card was built to prevent, arriving through the
control that was supposed to prevent it. The `--auto` wording went with the field, because the
flag is no longer the only route into an ungated room and naming it asserts a cause that may not
be there.
**The rule, stated once so the next control inherits it: a flag that gains an in-room twin stops
being the answer to "what is the room doing" and becomes only the seed.** `dispatch.go` already
says this for the request path — "`m.st.Asking`, not `m.opts.Auto`: the flag only SEEDS this at
launch" — and the two misses were both places that had not heard. Every `opts.*` read on a
demoted control is now either a launch-time decision (`wantsGateHook`, the `savedPosture` record) or
a bug, and the way to tell them apart is to ask whether an in-room control can move the state
underneath it. Both fixes are pinned by tests rather than comments for the same reason: in every
room nobody typed a control into, the flag and the state agree, so the fixtures cannot tell them
apart and neither could review.
### 9.18 a strip said four fifths of a name it could have said whole in two letters
Since the default route stopped being everyone, the ordinary turn narrows the frame to one
seat and leaves the rest at `stripColumn` — fourteen cells. Every layout rule in §9.11 was
written for a column three times that, and at fourteen the room did the opposite of what
§9.11 ruled in both halves of the chrome at once. `Antigravity` rendered `Anti…`. The badge
row rendered `ro:tools to` and `gated fina`, and the overflow marker rendered `↑ 12 more
abov`.
The ruling those violate is §9.11's own: **a clipped seat name is still recognisable and a
clipped state word is not**, so identity yields first. A clipped state word that is also the
prefix of another word in the same vocabulary is worse than damage — it reads as a different
claim. `fina` is not a broken `final only`; it is a thing this room does not say.
So at strip width the room **sheds whole words** rather than cutting them, in a fixed order
that is a pure function of the width — which is what lets the frame sweep pin the whole
ladder instead of a golden per state:
- **Identity collapses to two letters.** `CC ✓ done`, `CX ○ idle`, `AG ⠋ streaming`. The tags
are the HUD's own, character for character, because a reader who learned `CX` is Codex from
the HUD's grid must not meet a second abbreviation in the room. They are *copied*, not
imported: the seam between the two surfaces is the normalized session model and
`internal/theme`'s numbers and nothing else, and a test asserts the strings by literal so
the copy cannot drift in silence.
- **The clock goes, then the focus mark, then — for `unavailable` alone — the tag itself.**
`8s` is the meta on that line and `turnRule` already ranks a label above the numbers that
belong to it; every finished turn still carries its elapsed on its own separator. The
arithmetic behind the rest: nine cells of `streaming` plus its mark leaves exactly three,
which is a two-letter tag and the space after it, and `unavailable` at eleven leaves room
for a mark or a tag but not both.
- **The badge row keeps the posture word and drops the cost and the granularity.** §9.2 is
emphatic that a claim you cannot see is not a claim, so the safety word is the last thing on
that row to go; the cost is a number the transcript records on every turn separator, and the
granularity word exists to keep `waiting` from reading as a slow `streaming` — both of which
are now the only thing on the header one row above. A badge too long for a strip would drop
rather than clip, and stays readable at full length on the `?` postures page.
- **The overflow marker sheds `more`, then `above` / `below`.** The count is never traded: how
much is hidden outranks which way to press, which outranks the filler between them.
**The focus mark is the one deliberate loss, and §9.12 is why it is affordable.** Two cells of
`▸ ` at fourteen is the difference between a tag and no tag for every nine-letter phase word.
§9.12 had already found that the glyph was the *weakest* part of that signal — "one `▸` in a
frame carrying four columns of prose" — and moved the load-bearing half onto weight and onto
the overflow marker's own words, `↑↓ scroll` against `tab to focus`. Both cost no cells and
both survive here. A strip is by construction the seat this turn was **not** addressed to, so
spending a seventh of its width marking it, at the price of its identity, inverts the priority
§9.11 set.
**What was declined.** Keeping the two-cell indent on a strip so the focused and unfocused
forms line up: it is chrome that exists to align a *name*, and at strip width there is no name
— the header starts at column zero and the badges start under it, so the strip reads as one
flush-left block rather than as a column with its margins still on. Shortening a phase word to
fit (`cancel` for `cancelled`, `stream` for `streaming`): a different word is a different
claim, and the vocabulary is shared with the help panel and the transcript. Giving the strip a
narrower vocabulary of its own — a second alphabet is exactly what §9.11's phase marks were
built to avoid.
### 9.19 sitting a turn out cost a line a turn, and wore the wrong mark doing it
Since the default route became one seat (#99), three columns sit out every ordinary turn. The
room said so, correctly, once per turn — and a quiet seat's transcript became a column of
identical warnings with the answer it actually gave scrolled off the top:
> `⚠ not addressed in turn 2` / `⚠ not addressed in turn 3` / `⚠ not addressed in turn 4` / …
Two things are wrong there and they are separate. One is the arithmetic. The other is the mark.
**Consecutive skips coalesce, at render time only.** A run of turns this seat was not part of
is one muted line — `not addressed in turns 2–7`, singular for a run of one. The run is the fact;
the turns inside it are not separately interesting, and a reader who wants one has the numbers.
The **data model is untouched**: nothing is written down for a turn a seat did not take, which
is §9.9's rule and the reason a transcript skips from 3 to 5 in the first place. The runs are
*derived* from the gaps between the turns that ARE recorded, so `[` and `]` still hop between
real turns (§9.20) and no record says anything it did not say before. A run broken by a turn
the seat took starts a new line in place, so the transcript still reads in order, and the LIVE
turn's skip keeps a line of its own — the run above it is history, that one is the turn the user
is deciding whether to act on.
A run is never claimed **before the oldest record**. History is capped at fifty and drops the
oldest first, so a column whose early turns were evicted would otherwise report "not addressed
in turns 1–29" about turns it may well have answered. Inventing an absence is the same error as
§9.9's inventing a conversation, run the other way.
Underneath the rendering bug was a data one, and it is the reason `Column.Skipped` exists at
all: the note is written on the LIVE column, and the live column is what `startTurn` files into
history. A seat that answered turn 1 and then sat out through 7 filed turn 1's record wearing
`not addressed in turn 7` — a turn that succeeded, with someone else's absence stapled under it.
A skip is not a fact about any turn this column recorded, so it does not travel with one.
**The mark is demoted to `○`.** `⚠` opens a note because a note reports something that did not
complete normally — a cancellation, a seat that is not there. Sitting a turn out is neither, and
it was a fair mark only while a narrow route was the exception. Drawn on the ordinary case it
is a warning the eye learns to skip, which is the same argument `ActDenied` makes for `SevWarn`
over `SevCrit` and the reattach card makes for no mark at all. `○` is what this room already
spends on *nothing has been asked of this seat*, which is exactly what a skipped turn is, said
about one turn instead of a session. It survives `--ascii` as `.` against the warning's `!`, so
the demotion is legible with colour switched off — and the word carries it first either way.
**An idle strip says where it left off.** At fourteen cells (§9.18) a backgrounded seat had a
header, a posture word and a run of skips, and the one thing a reader wants from it is which
turn it last took: `last: turn 8 ✓`, above the coalesced line. Every part of it is measured —
the number is the turn this column recorded, the mark is that turn's own phase — and a seat
with nothing behind it renders nothing rather than a placeholder, because absent is absent
(§4a.1) and this room does not draw `last: —`. Strip width only: a wide column already answers
the question with the turn separators themselves, and repeating it there would be the room
being loudest where it has the least to add.
**What was declined.** Recording a `TurnRecord` per skipped turn so the coalescing could read
one list: that is the room writing down a conversation that did not happen, and it would put
the skips in `[`/`]`'s path. Dropping the live skip into the coalesced run to save a row: the
run is history and that line is now, and a reader deciding whether to re-address a seat should
not have to read a range to find out. And giving the skip a mark of its own — `○` already means
this, and a second glyph for one meaning is the collision `glyphs.go` argues against.
### 9.20 the transcript is turn-wise and the only way through it was line-wise
§9.9 gave every column a real conversation and §9.10 and §9.12 made it reachable and
attributed. What none of the three changed is the *unit*. The scrollback moves a line at
a time, a page at a time, or all the way to either end — and the thing being scrolled
through is a list of turns, each one a labelled rule, a brief, and however much prose a
vendor felt like producing. So the room could tell you, honestly and precisely:
> `↑ 509 more above`
and nobody has ever counted lines. The number is measured, it is correct, and the only
question a reader actually has — *how far back is what I asked?* — is one it cannot
answer. `g` goes to the beginning and `G` goes to the end, which are the two positions
in a transcript that need no help finding.
**`[` and `]` walk the focused column one turn at a time.** They land the turn's
separator on the viewport's top row, which is the position that makes the brief and the
answer to it readable in one screen, and they take their offsets from `columnLines` —
the *same* pass that produced the lines — rather than recomputing where a turn starts
from `History`. A second derivation of "how tall is this turn at this width" would agree
with the first until the day a card grows a row, and would then disagree silently, since
both answers would still be plausible line numbers.
**Backwards is the audio player's rule, and it is the one people already have in their
hands.** `[` from the middle of a turn lands on *that* turn's head; only a second press
reaches the one before it. That falls out of the definition rather than being a special
case — "the last head strictly above where we are" produces both — and it means the key
answers "start this again" and "go back one" with the same press, in the order a reader
wants them.
**The two ends are deliberately not symmetric.** `[` at the first turn does nothing:
there is no turn 0, and a wrap would make a key pressed one time too many jump an entire
conversation. `]` past the last turn restores the tail and `Follow`, because what comes
after the last turn is the live output — that is `G`'s answer to the same question, not a
second one. Every landing goes through `applyScroll`, so `Follow` drops exactly as it does
for `↑`; a column pinned to the tail while displaying turn 3 would be lying about which of
the two it is doing.
**In compose they are the characters `[` and `]`.** No rule was added for that: §9.10
replaced the composer's list of exceptions with a test — a key that carries text *is*
text — and brackets carry text. This is the same contract that keeps `q` the letter q
there, and it is asserted rather than assumed, because a bracket that scrolled instead of
typing would corrupt a draft in a way the user would only find after pressing enter.
#### The marker states the coordinate, and §9.12's rules decide what it costs
The overflow marker is where the count lives, so it is where the coordinate belongs:
`↑ 25 more above │ turn 3 │ ↑↓ scroll │ f expand`. Three constraints from §9.12 and
§9.10 bound the whole design, and each one closed a question:
- **The count is never traded away**, in any form, at any width. It was §9.10's rule about
the key hint and it is unchanged: how much is hidden outranks both how to reach it and
what it is.
- **The coordinate sheds FIRST** — below even `f expand`. It says *where you are* while the
hints say *what you can do about it*, and a marker that dropped a key to keep a coordinate
would be §9.10's trade run backwards. Concretely it rides only on the widest hint form, so
at the three-up tier's 37 cells the keys win and the coordinate is simply absent; `f`, the
reading tier, is where it appears. That is the graceful degradation, not a gap in it.
- **A marker states the key for THIS column and never a neighbour's** — §9.12's rule, applied
to a fact rather than a key. An unfocused column keeps `tab to focus`, the one thing a
reader looking at it can act on, and gets no coordinate at all: putting the question in
front of the answer is how §9.12's bug worked in the first place.
**Which turn it names is the part that could have lied.** The choice was between the topmost
hidden separator and *the turn the line immediately outside the fold belongs to*, and only
the second is honest when a turn is half on screen: a long reply running off the top is still
the turn you are reading, while the topmost hidden separator can be several screens further
back and answers a question nobody asked. So `turnAt` takes "the last turn that started at or
before this line", the two markers on one column name two *different* turns, and a column with
no turns at all — an unavailable card, a seat never asked anything — prints nothing rather
than `turn 0`, because a coordinate the room does not have is omitted and never invented
(§4a.1).
#### The footer learned to shed a cell instead of losing its way out
`[ ] turn` joins the view mode line immediately after the arrows, offered unconditionally for
§9.12's reason — the promise is about what the mode can do, not about how many turns a vendor
happens to have taken, and a footer cell that appeared at the first dispatch is chrome moving
while output arrives.
That exposed something the line had been getting away with. At the tabbed tier the six hints
fit **exactly**, and `statusLine`'s only answer to running out of width is to truncate — from
the right, which is where `? help` and `q quit` live. A motion key bought with the panel's
documented way out and the room's only quit key is precisely the trade §9.11's footer pass
existed to refuse. So a hint may now be marked *sheddable*: when the line does not fit, the
sheddable cells are dropped whole, newest-first, before the ellipsis is allowed to choose.
Exactly one hint carries the mark, and it is the one this section added — this is a rule about
which cell goes, not a licence to hide keys.
The help panel took it inside the hard 17-row budget by merging onto the row that already
holds the other jumps, the way §9.15 merged `y`/`Y` onto the gate's row: `g / G first turn or
newest; [ ] step one turn at a time`. "jump to the" paid for it — the line above already says
`scroll`, so the verb was never carrying anything.
**What was declined.** A turn coordinate on unfocused columns, which the width would have paid
for out of `tab to focus` (above). Numbering the hop in the notice line — "turn 3 of 7" is a
progress bar for a conversation, and the marker already says how much is left in the unit the
scroll keys use. And a `[`/`]` that moved *focus* between columns when a column has one turn:
two motions on one key, resolved by content, is the kind of binding that is only ever right for
the person who wrote it.
### 9.21 the room knew what the turn would cost and did not say
#99 restored the cheap default: silence goes to Claude alone, and the committee is
convened by typing `@all` or naming the seats. That settled *which* route is expensive and
made every expensive route explicit — and it left the footer stating the route in the same
words whether it reaches one vendor or four. `→ everyone` is accurate, and how much
`everyone` is depends on what is installed and on what `--vendor` left in the room, which
is exactly the part a user cannot read off the word.
**The routing cell states the bill when the draft would reach more than one seat.**
`→ everyone (3 seats)`, `→ everyone but codex (2 seats)`. The room already computes this
number — `dispatch` refuses a turn that reaches nobody by counting it — and the moment it
is worth knowing is the moment before `enter`, on the cell that is already answering the
same question.
- **One seat states no count.** `→ claude` names every seat it reaches in its own text; a
cell that restates its neighbour is how this footer became the wall §9.11 had to take
apart. From two upward the route names a *set*, and the size of a set is not in the word.
- **It counts seated ∩ addressed**, through the same `State.SeatsIn` the dispatch gate now
calls. A route may name a vendor that is not installed or that `--vendor` left out; that
seat is never spawned, so billing for it would quote a price for a turn that does not
happen. `Model.seatedIn` became one line delegating to it rather than a second copy —
a bill derived from different arithmetic than the dispatch is a bill for a different turn.
- **A refused route prices nothing.** `mixed @ and -@` addresses nobody, and the one thing
that cell owes a reader mid-typing is what is wrong with the line they are still holding.
- **No colour, no cell, no new glyph.** The count is the *label* half of a hint and the
route is the *key* half, which is the figure/ground split every other item on this line
already makes (§9.11) — so the seat names keep their intensity and the number recedes to
chrome for free. It is parenthesised because that is this room's existing grammar for a
qualifier on the thing in front of it (`(+2 queued)`, `(turn 1 is blind)`), and because
weight is invisible under `NO_COLOR`, where `→ codex, agy 2 seats` runs the price into
the list it is pricing. The rebuttal tag moved to its own cell so the count could sit
against the route it prices; it kept its intensity by keeping the key half of a hint.
#### The header carries the live turn's route
Once `enter` is pressed the composer clears and its routing cell resets to the *next*
draft's default, while the columns take anything from seconds to minutes. For that whole
window the room has nowhere at all that says where this turn went — and each column's
transcript does not record participation until it lands. So the header's turn cell carries
it while it is live: `turn 10 → everyone`, `turn 10 → codex, agy`, reverting to plain
`turn 10` when the last column finishes. **The route becomes history at that instant**, and
the transcript is where history goes; a header still naming it would be describing the past
in the one cell that describes the present.
`State.TurnRoute` is a **pointer**, and that is §4a.1's zero-vs-absent rule rather than a
style choice: `Route{}` is a real and extremely common route — it is what `@all` parses to
— so a value field could not tell "this turn went to everyone" from "no turn is running".
The same distinction `Column.CostUSD` draws with the same mechanism. It is set where the
turn actually starts rather than beside `FrameOwners`, because everything above that line
can still refuse the dispatch and a route on the header of a turn that never began would
report a spend that never happened. The two have opposite lifetimes on purpose: the
geometry outlives the turn so nothing reflows under a reader (§9.11), the route is retired
with it.
**It prints the route's own `label()`**, never a second vocabulary — what the header shows
is what would have to be typed to produce it — and the arrow is the literal one the
composer's cell uses rather than a `Glyphs` entry, so one fact cannot drift into two
spellings.
**A `/flow` hop states no route at all.** A hop is dispatched to exactly one named seat
(§9.16) and the cell immediately to its right already says which, so the route would be the
header saying the same thing twice — and the arrow, appended after the hop, would read as
pointing at it. This is the same rule as the shedding below rather than an exception to it.
#### Shedding order: a fact with a home elsewhere yields to facts that have none
The header already elides the workspace path from the left, and the new cell had to be
ranked against it. **The route sheds first** — before the path, before `3/4 seated`, before
`briefed`. The route is on screen in the composer a keystroke earlier and in the transcript
a moment later; the workspace is nowhere else at all, and it is the one fact here that
changes *what the agents can see*, which is why it has been on screen at all times since
this header was written. So the route is added only when it costs nothing that was already
there: the path keeps its cells if it had them, and where there was no room for a path
either way, the counts keep their gap.
**What was declined.** A dollar figure beside the seat count: cost is reported per seat per
turn where a vendor reports it, and multiplying a seat count by anything would be council
deriving a number and presenting it as read — the top item on this repo's rejected list
(§4a.1). Billing the *route's* vendors rather than the seated ones, which would have been
one line shorter and would have priced seats that are never spawned. And a count on the
one-seat case, which is a number whose only reading is "yes, one".
#### Amendment, 2026-08-17: the room shipped the quota relay and never read it
The cell above prices a turn in SEATS, which is the half of the bill council could count.
The other half was already on disk and nobody was looking at it. `telltale statusline` has
relayed every quota window it renders to `~/.telltale/quota/.json` since 2026-08-07
(§7.15) and the HUD has read it ever since — while the one surface that actually *spends*
those windows, four or five accounts at a time on one keystroke, said nothing about any of
them. The room could tell you a turn would reach three seats and not that one of them had
nothing left to answer with.
**Council reads the relay. It writes nothing.** `internal/council/quota.go` reads
`quotacache` at room open and again when a turn tears down, as a `tea.Cmd` returning a
`quotaMsg` — never inside `Render`, which stays pure over `State` (`TestRenderIsPure`), and
never on the tick, because the file only changes when the user's own statusline fires. The
read/write boundary is untouched: council's one sanctioned write is still `room.json`, and
a second one would have to be argued from scratch rather than inherited from a read.
**§7.17's declined "per-row quota" does not bind here, and the reason is arithmetic.** That
ruling refuses a quota cell on a HUD *row* because a row is one session: five Claude
sessions would each draw the same 42% and read as five separate budgets, asserting a
per-session limit that does not exist (§7.1 rule 6). A council **seat** is not a row. There
is exactly one seat per vendor in a room, so a seat reading is an account reading printed
once against the account it describes. Nothing here is ever drawn per session, per turn, or
twice for one vendor.
##### What renders on a seat
A text reading on the badge row: the window's own label, the vendor's own percentage, the
reset countdown while it fits, and the reading's age.
```
ro:tools tokens 5h 12% resets 1h04m 7d 6% resets 5d00h 2h ago
ro:tools tokens 5h 12% 7d 6% 2h ago
ro:tools tokens 5h 100% ⚠ stale 19h ago
```
The word `resets` rather than the HUD's `↻`: council's `Glyphs` has no slot for that mark,
and minting one would grow a set §9.26 keeps deliberately small. The cells between parts are
this room's own two spaces (`historyMeta`, the badge row itself), not the HUD's middle dot,
for the same reason — one surface, one joiner, and nothing new to give an ASCII partner.
- **No gauge track, and that is a ruling rather than a shortcut.** The HUD spends a bar on
this because it has a header line to spend it on. Council would need a fill colour to draw
one, which re-opens both the closed `isDark` question and `style.go`'s standing rule that
council adds no hues of its own — a large purchase for a signal the percentage beside it
already carries. The label and the digits are words and numbers, so `--ascii` and
`NO_COLOR` lose nothing at all.
- **The reading takes the space the row has LEFT.** A new claim does not evict an older one:
the posture badge is the safety claim §9.2 refuses to let yield, the granularity word is
what keeps `waiting` from reading as a slow `streaming`, and the cost is the one figure on
this line the transcript also records. All three keep their cells; the reading takes what
remains, sheds its countdowns, and drops **whole** rather than clipping (`stripBadges`'
ruling — a clipped percentage is a different number). At the reference 120 columns a
three-seat grid gives a column thirty-eight cells and only the roomiest badge row has
space for a figure; the footer cell below is what a narrow room keeps instead.
- **The age is the HUD's, verbatim.** `2h ago` from five minutes, escalating past five hours
to `⚠ stale 19h ago` — same threshold, same word, same order (word, then glyph, then
hue). `quotaAgeShown`, `quotaAgeWarn` and `quotaAgeWord` are **copied** rather than
imported, on `vendorTag`'s precedent: internal/council and internal/hud share the
normalized session model and internal/theme's numbers and nothing else, and
`TestSeatQuotaAgeMatchesTheHUDsThresholds` pins all three by literal so the copy cannot
drift in silence. A reader who learned `stale 19h ago` on the statusline must not meet a
second spelling of it in the room.
- **A window relayed for its reset time alone renders nothing.** It says when something will
change and not what is left, which is the only question this line exists to answer.
##### What renders on the route cell
One seat's name and one of that seat's own readings, when the reading says the turn may not
land the way the reader expects: `⚠ claude 5h 100%`, `⚠ agy stale 19h ago`. It sits against
the route it qualifies, in compose mode only — the header's live-turn route is already too
late to act on.
- **It computes nothing.** The refusal above declined a dollar figure beside the seat count
because multiplying a count by anything is council deriving a number and presenting it as
read. The same refusal binds here and is wider: no total across seats, no average, no
count of how many seats are affected, no percentage arithmetic of any kind. Every
character after the vendor id is copied off one window.
- **Seated ∩ addressed**, the same intersection `State.SeatsIn` counts and `dispatch` loops
over. Warning about a seat this turn will not reach is a warning about a turn that does
not happen. **The count cell is untouched** — it keeps its present grammar and its present
arithmetic.
- **A hundred per cent is the only threshold, and it is the vendor's.** Ninety, or "nearly
full", would be council picking a severity boundary no vendor published — the same class
of guess as filling a `CapNone` field with a plausible value (§4a.1).
- **Staleness outranks fullness for one seat**, which is `quotaAgeWarn`'s own argument: a
reading past it may no longer be assumed to describe now, so a stale 100% is not evidence
the window is full, it is evidence the room does not know. Reporting it as `100%` would be
the nineteen-hour incident reproduced in a new room.
- **The word "spent" is refused.** §7.17 owns it for token counts, and quota and spend are
the two claims that view exists to keep apart. The reading needs no verb: `5h 100%` says
it.
##### Zero, absent, and the three vendors that are absent forever
`Column.Quota` is a **pointer**, the same mechanism `Column.CostUSD` and `State.TurnRoute`
use and for the same reason. A vendor at 0% of its window has been measured and draws
`5h 0%`. A vendor with no relayed reading draws **nothing at all** — no dash, no
placeholder, and not one cell of width, which `seat-quota-absent.txt` pins by rendering the
two states side by side in one frame. Collapsing them is the more dangerous direction of the
zero-vs-absent bug here: an unrelayed seat would read as a fresh account and invite a
dispatch the room has no evidence will land.
Cursor and grok are in the absent class permanently, and so is Gemini: none of them writes
quota to disk in any form a passive reader can see (§7.17's structurally-absent row), so no
relay entry can ever exist for them. Codex is absent from this surface for a different
reason — its quota lives in its own store, and this room reads the relay and nothing else.
The room does not distinguish the three, and it does not have to: on this surface they are
one fact, *this room has no reading*, and each vendor's own sentence explaining why is
§7.17's job on the surface built to hold a paragraph per vendor.
**Expiry is the read's, not the room's.** `quotacache` drops a window whose reset has passed
and any entry over 24h old before council ever sees it (§7.15), and a read that no longer
speaks for a vendor **clears** that seat. A room that kept its previous reading would be
displaying a percentage §7.15 calls not stale but FALSE.
##### Limitations, recorded rather than left to be found
- **Exactly one seat is named on the route cell, and a second is not counted.** Column order
decides. Ranking two seats would mean ranking a stale reading against a full window, and
there is no measurement behind such an order; a count would be the aggregate this cell may
not compute. What carries the rest is each seat's own badge row — and at a width where the
badge row shed its figure, a second affected seat is not on screen.
- **The reading is as old as the last statusline render in that vendor.** Council writes no
relay of its own, so the post-turn read sees a turn's cost only after that vendor's
statusline fires again. This is exactly what the age suffix exists to say, and it is why
the age never sheds.
- **A room open past `quotaAgeWarn` with no statusline activity escalates every reading it
holds.** That is correct rather than noisy — the readings really have outlived the fleet's
shortest window — but it means a long idle room ends up with a warning on its footer that
only the vendor's own statusline can clear.
### 9.22 four answers to one question, and no way to read them as one
Council exists to put several vendors' answers side by side. Everything from §9.9 onward
built the surface that does it — a real transcript, per seat, scrollable, attributed,
navigable a turn at a time — and every one of those sections improved a **column**. The
room therefore had a comparison surface with no way to read a comparison. To see what four
seats made of one brief you scrolled Claude to turn 10, remembered it, tabbed, scrolled
Codex to turn 10, remembered that, and tabbed again; §9.20's `[` and `]` made each of those
one keystroke and did not change what the exercise was.
**The document already existed.** §9.15's `Y` assembles exactly this: the brief once at the
top, then every seat that took *this* turn, labelled, in seating order — and it was ruled,
argued and tested a release ago. What it could be read in was a clipboard. So this section
adds no content model at all; it renders the one `Y` already had, and the two now come from
the same call (`turnEntries`), which is the point rather than a tidiness. A page and a
paste that disagreed about who was in a turn would be two honest-looking documents with
nothing on screen to say which was the room's answer.
**`t` swaps the body between the by-seat grid and one turn's page.** One key and a toggle,
because the two are one transcript read two ways rather than two places — and it opens on
the turn the grid was already following, so the projection changes and the subject does
not.
#### What the page is, line by line, and why none of it is new
- **The turn's own rule**, carrying the number, where the turn went, and how long it took:
`turn 10 ──────── → claude, codex 41s`. Same `labelRule` grammar, same shedding, meta
before number, as every separator since §9.11.
- **The brief once**, under the composer's own `›` at full weight. Four copies is what a
*grid* has to do — each seat's prompt is a fact about that seat (§9.9), since a turn can
reach two seats and not a third — and it is precisely what a page must not.
- **Each participating seat under its own labelled rule**, name at weight, then its
activity trace, then what it said, with §9.11's boundary strengths unchanged. The only
thing that differs from a column is what the strongest boundary is *about*: a turn there,
a seat here. That is what swapping the projection means.
- **A seat that sat the turn out does not appear.** §9.15's rule, for §9.15's reason: it
still holds an older reply, and filing that under this turn's heading would be the room
inventing a conversation — on the surface built to compare them, where it would be
believed.
- **Failed and cancelled seats keep their note cards.** A turn's page shows what actually
happened; the two turns anyone scrolls back for are the ones that went wrong.
**The route is read off participation, not off `State.TurnRoute`.** §9.21 retires the live
route the instant the last column lands, because the header describes the present. What
outlives it is the measurement — a `TurnRecord` exists for exactly the seats the brief
reached — so the page states who took the turn, through `Route.label()` so what is
displayed is still what would have to be typed to reproduce it. **The clock is the longest
seat's own elapsed**, because a turn is over when its slowest seat lands. A sum would be
the wall time of a room that dispatched serially and a mean is a duration no seat ever
took; both would be council deriving a number and printing it as read, which is the top
item on §4a.1's rejected list, and being in seconds does not exempt them. A turn still
running carries no turn-level clock at all — how long it took is not a fact yet — while
each seat's own rule carries its running one, from `State.Now`.
#### The two rules that shaped it more than the layout did
**Gate precedence, in both projections.** A pending approval renders on the page as
*chrome* — above the scroll, like `columnChrome` — and the argument is stronger here than
in the grid: a vendor is stopped, the live page follows its own tail, and a card inside the
body would be pushed off screen by the output of the very call it is asking about. It also
**names the seat**, which the grid's card never had to: there the card's position *is* the
seat, and one page has no position left to carry it. And `y`/`n` still answer the gate
before they mean anything else, because `key()` routes to `gateKey` first in either view.
That was already true; §9.15 made it asserted, and it is asserted again here — a keystroke
the user believes approved a write must never quietly copy text instead, since their next
move is to press it again.
**§7.1 rule 4 decided what the footer says.** A turn arriving while an older page is open
**never moves the view**: content jumping out from under a reader because a vendor finished
is the thing the bottom-anchor and the frozen-geometry rules exist to prevent. But a reader
on turn 10 of a room now on turn 11 is looking at something stale, and silence about that
is its own dishonesty — so the drift goes where a reader already looks to learn what the
keys mean. The mode word is `TURN 10/11`. §9.20 declined "turn 3 of 7" and this is not that
reversed: that was a progress bar offered in the *notice* line, describing a hop that had
already happened. This is §7.8's always-on mode label answering which projection is live,
which is the one thing the body has been ruled out of saying.
Pressing **enter** is the exception that proves the rule: dispatching from a page lands on
the turn just sent, because that move is the user's, not a vendor's, and a projection that
answered a new brief by staying on turn 7 would show an old conversation while spending
quota on a new one.
#### What it does not get, and the two keys that say so
There is **no column focus** on a page, so `tab` and `f` do nothing — and both are dropped
from the mode line rather than left promising something, which is §7.8's surprise pointing
the other way (§9.11's footer rule). The overflow markers follow: the focused-column form,
no `tab to focus` and no turn coordinate, since every line on a page belongs to the same
turn and the mode word already names it.
For the same reason **`y` and `Y` produce the same document here**. A per-seat `y` needs a
per-seat focus, and a projection whose whole unit is the turn deliberately has none — so
the narrower key takes the wider document rather than guessing which seat was meant. `y
yank` is named on this mode line and not on the grid's, because here the key takes the
thing in front of the reader, which is what makes it worth a cell.
`i` is the deliberate omission from that line. The six cells the page needs are its own
motions and its two ways out; the composer is one `t` away in a mode line that names it,
and it is the first row of the help panel. A footer short of width starts cutting into `?`
and `q` (§9.20), and that is the trade this line was designed never to make.
`t` joins the help panel by merging onto `f`'s row, inside the hard 17-row budget: `f gives
one column the full width; t gives one turn the whole room`. Not a saving — the same
category, the way §9.15 merged `y`/`Y` and §9.20 merged `[ ]` onto `g`/`G`. Both keys
answer one question, *how much of the room is the reading area*, and a reader looking for
either is looking for the other.
Everything else is reuse rather than resemblance. The page plans as **one column at the
full frame**, which is the tabs tier's own arithmetic, so the height budget, the 60-column
floor, the composer's growth and the collapsed-seat notice are identical in both
projections — a second layout path for a surface that *is* a column at full width would
just be a second place for the frame to tear. The scroll window, the overflow markers, the
tail and the clamp are §9.9's own argument applied once more: a page is a flat list of
lines, and this room already knows how to move through one. `[` and `]` keep the words they
have in the grid, at the same unit, so there is one motion to learn; `g` and `G` reach the
same two positions — the oldest turn still in memory and the live end — in the projection's
own unit. A turn the fifty-turn cap has evicted has no page, and says so rather than
drawing an empty one: "nobody answered" and "the room no longer remembers" are different
facts (§4a.1).
#### Declined
- **Cross-seat diff or agreement marks** — "these two agree", "this one dissents". A page
puts the answers where a person can judge them; a mark would be council judging them,
which no adapter sourced and no vendor reported (§4a.1). It is the same refusal as the
"role" line §9.11 declined, with a harder consequence: a wrong agreement mark is one a
reader would act on.
- **Persisting the projection.** `room.json` stays keys-only (ADR-008, ninth amendment) and
which turn someone was looking at is not state the next session should inherit — §9.9's
argument for not persisting the scrollback, one surface up.
- **Per-seat focus inside the page**, with `tab` cycling seats and `y` taking one of them.
It would import the grid's whole focus apparatus — a marker, a weight, a hint on every
marker — into a view whose entire claim is that the turn is the unit, and it would buy
one thing the grid already does better. v1 lacks it deliberately, and `y`'s behaviour
here is what falls out of that rather than a limitation worked around.
#### Amendment, 2026-08-17: the act ledger — the same turn, read for what the seats DID
**The gap.** The page above answers *what did the seats say about turn 10*. The other half of
a turn is what they **ran** — the tool calls, the commands, the edits, the one that was
refused at the gate — and the room has been parsing, redacting, retaining and rendering
every one of those since §9.6a. What it renders them in is a 37-cell column, where the
outcome is a single mark and a wrapped command is most of the width. So the record existed,
in full, with nowhere to read it: to answer *did anything fail in turn 10, and where*, you
scrolled a column at a time and read outcomes off four glyph shapes.
**`T` opens the same turn's acts.** Not a third projection — a second FACE of the one the
page already resolved. `TurnView` keeps deciding which turn is on screen; `TurnView.Ledger`
decides which of that turn's two records is drawn. Everything else is untouched: `[`, `]`,
`g` and `G` move the same coordinate, the scroll window and the overflow markers are the
same code pointed at a different list, and the page's own geometry (one column at the full
frame) is unchanged. A second `TurnView` would have been a second answer to "which turn is
open" and a second scroll model to keep in step with it.
**The key is SHIFT on `t`, and that is the only spelling that puts the third reading beside
the two it belongs with.** `t` gives one turn the whole room; `T` gives that turn's acts the
whole room. Every free lowercase letter left in this keymap is free *because it means
nothing here*, and a projection filed under an unrelated letter is a projection a reader
finds by accident. The capital is unclaimed — `Y` and `G` are the only two this room binds —
and in compose it is the letter T, which needs no second list: `composeKey` routes any key
carrying text into the draft, the contract `q`, `f`, `c` and `t` already keep. It **flips**
rather than navigates (`toggleArenaDiff`'s shape, one scale up: `d` flips one seat's arena
block between the stat and the whole patch), so a reader who walked back to turn 7 is still
on turn 7 in either face. `t` keeps meaning the reading face from anywhere, including a
re-open after a close — a `t` that sometimes landed on the ledger would be two keys wearing
one name.
**The outcome is a WORD, and that is what the width buys.** `⚙ Bash: go test ./... ✓ ok`,
`✗ failed` with the vendor's own first line under it, `? outcome unknown`, `✗ denied by
you`. The mark is `actMark`'s, unchanged; the word beside it is the signal it seconds, so
`--ascii` and `NO_COLOR` lose the mark and lose nothing else. **An act with no reported
outcome never renders as one that worked** — that is `runner.ActStatus`' whole reason for
existing (antigravity's steps flip ACTIVE then DONE and no captured line has ever carried a
success signal), and a surface that states an outcome on every line is exactly the shape
that invites a default. An unresolved call splits once more: while the seat is waiting or
streaming it is `running`, and once that seat has landed it is `no outcome reported`,
because the vendor never said the step ended at all. The predicate is
`turnEntry.working()`, not `turnEntry.Live` — the newest turn stays the column's *current*
one long after every seat has finished, and reading `Live` alone would report a dead call
as running for the rest of the session.
**The header states the retention window, from the live constant.** `maxHistory` drops the
oldest turn per seat, so "the acts" is a claim with a hard floor under it, and an
unqualified one would be the room offering a record while silently forgetting the far end of
it. It is a LINE hanging under the rule rather than meta on it, because `labelRuleIn` drops
its meta whole when the width will not take it — correct for a route and a count, wrong for
the sentence that bounds the claim, which would then vanish exactly where the room has least
room to make it. The clipboard document carries it too, and there it matters more: on screen
a reader re-checks the bound by pressing `[`; in a file pasted into an issue a week later
that sentence is the only thing saying the record was ever bounded.
**A seat that recorded nothing says `(no acts recorded)`, not that it did nothing.** A trace
is a reading of what a vendor chose to report, so "this seat did nothing" is a claim no
adapter here can source — the §4a.1 distinction between "we could not read this" and "there
is nothing there", on the one surface a reader would take as the record of it. The same
words carry the turn-level zero on the rule, so there is one spelling rather than two. A
seat that SAT THE TURN OUT is absent entirely, which is §9.15's rule and binds harder here:
an older turn's `git commit` filed under this turn's heading would be a history the room
invented, in a document somebody pastes into a review.
**`y` and `Y` follow the face.** Both keys already produce the page's own document, and
`YankPage` is what keeps that promise true once there are two of them — a copy key that took
the replies while the acts were on screen would break the one claim that earned it a footer
cell here. The document is built from **the same `turnEntries` call** the screen renders
from, and there is **no second sanitizer**: everything on `State` has already been through
the one redact-and-sanitize choke point, so a cleaning step of the ledger's own would be a
second answer to what is safe to put on a clipboard, and the two would differ the day one
was updated.
**The help panel merged, not grown.** `f / t / T` on the row that already holds both, inside
the hard budget: a row of its own would push the `?` line off a 24-row terminal, which buys
discoverability for one key by taking away the way out of the panel. "gives" paid for it
twice — the verb is established by the first clause and the two after it read as the same
sentence. The row lands at its 114-cell budget exactly. The **mode word** is `ACTS 10/11`
against `TURN 10/11`: two documents at one coordinate would otherwise leave §7.8's always-on
statement of what is on screen unable to tell them apart. The mode line's cells are
unchanged, `t grid` included — it is still true from either face, and it is the way out this
line may never shed.
**What it deliberately does not get.** No note cards: how a seat's turn ENDED is already on
that seat's own rule in `seatMeta`'s words, and a card under it would spend rows restating an
outcome the reader has just read. No replies: that is the other face, one keystroke away in a
mode line that names it, and a ledger carrying the prose too would be the page with extra
rows. No new record of any kind — this section adds no content model, exactly as §9.22 added
none. And **per-seat focus is still declined**, on the ruling above: a projection whose whole
claim is that the turn is the unit does not grow a focus one face later.
Verified offline. `ledger_test.go` pins the five outcome words staying five and an
unrecognised status rendering none, the unresolved call splitting on `working()`, the
retention sentence read off `maxHistory` rather than typed, the recorded-nothing wording, the
sat-out seat's absence from both the screen and the paste, the face flip not moving the turn,
`T` staying the letter T in compose, the gate still outranking `y`, and the help row's width
and prose column. `act-ledger.txt` and its `--ascii` twin are the frame. No test here spawns a
vendor. Nothing in this section is a claim about vendor behaviour, so no live run is owed:
the acts it draws are the ones §9.6a already measured, at a width that can afford to name
them.
### 9.23 the frame dashed, and the outline whispered while its entries shouted
§9.11 through §9.22 spent the room's typographic budget on *columns* — a seat's name, its
state, its cards, its transcript — and every one of them was measured against what a reader
could find. What none of them looked at is the thing holding the columns apart. Read the
repository's own goldens as pictures rather than as assertions and the frame is the first
thing wrong with them.
**The rails were a property of the prose, not of the grid.** The `│` between two columns was
drawn per row, on the test *does any column have ink on this line*. That predicate exists for
a real reason: a tall idle window used to draw four bars straight down through an empty screen
to the footer, and Phase 2 removed them. But the room seats three transcripts of different
lengths beside each other, and §9.11 spends a blank row as a boundary in three separate
places — between a seat's chrome and its content, where the speaker changes, where the kind of
content changes. So the ordinary case is that all three columns are blank on the same line
several times per screen, and the frame blinked out on every one of them. `transcript.txt`
broke at rows 11 and 13, `skips-coalesced.txt` at 5, 10 and 13, `unavailable.txt` at 19. The
per-row rule solved the void and created a stutter, and a stutter is worse: an edge that dashes
in and out at irregular intervals reads as damage, and it read as damage at precisely the rows
where the design had placed air on purpose.
**A row carries a rail when some column has content on it, or when it is a lone blank row with
content above and below.** Two consecutive blanks end the band; the next word starts a new one.
A separator is *structural* — it says these are different columns — and that claim is as true on
a quiet row inside a conversation as on a loud one.
**One row is the whole threshold, and it is the room's own number rather than a tuned one.**
Every deliberate blank this surface draws is exactly one row, and §9.11 names all three of them.
A one-row gap is therefore a boundary the design placed *between two things it means to keep
together*, and drawing the rail through it is drawing what was meant. Two rows is nothing the
design asked for — the bottom-anchor pad, an idle room, a column that ran out of transcript long
before its neighbour did — and there a separator has nothing to separate.
The **literal** reading was tried first and rejected on the evidence: rails on every row from the
frame's first word to its last. It is a simpler sentence and it produces a worse room. An idle
frame at 120×60 has chrome at the top and `no turn dispatched yet.` anchored at the bottom, so
one span runs fifty-five rows of bar through nothing at all — exactly the shape Phase 2 removed,
re-derived from a nicer-sounding rule. Contiguity is worth having up to the point where it starts
asserting a grid over emptiness. `TestTheRailNeverDashes` and `TestRailsDoNotSpearAVoid` hold the
two ends apart, and the older `TestRailsStopThroughEmptyBody` is kept unchanged as the third
witness that this pass did not quietly trade one for the other.
**The turn page's outline takes the weight its entries already had.** §9.22 gave a page two levels
of heading — the turn's own rule at the top, then one labelled rule per participating seat — and
drew the parent wholly `Muted` while `seatRule` gave every child `Strong`. The room's hierarchy
upside down: the eye landed on four vendor names and had to hunt *upward* to find out which turn
it was reading, on the one surface whose entire claim is that the turn is the unit. The turn rule
now takes the same split every heading in this room takes — the label at weight, the rule and the
numbers hanging off its end receding — which is the figure/ground rule the column header and the
mode line already make, applied to a heading instead of to a key.
**The grid's copy of that line is deliberately untouched, and the asymmetry is the argument.**
Inside a column a turn separator sits under a seat name already at weight; there it is the child,
and muted is its correct rank. On a page it is the root. The same line changes weight because it
changed what it is the parent of, which is what swapping the projection means. `strongLabelRule`
is one implementation for both callers, extracted for `labelRule`'s own reason: the thing being
kept in step is the grammar, and a second copy would drift from it one narrow-terminal fix at a
time. Weight costs no cells and `PlainStyles` renders it as the identity function, so this half
moved no golden — `TestPageTurnRuleOutranksItsSeats` asserts it where colour is asserted (§9.5),
and asserts the grid's separator did *not* move in the same breath.
**One separator, spelled one way.** The collapsed-seat notice joined its remedy with `" │ "` —
one cell of air — while the room header, the mode line and the column gutters all use two. §9.11
argues that number from `--ascii`, where the rule glyph and the spinner's first frame collide at
one cell, and the notice was the single place in the product spelling the room's only separator a
second way. It now reads from `gutter`, so it cannot drift again.
**What was declined.** Making the rail's weight or hue say anything — it is chrome, and a frame
that varied would be competing with the content it exists to bound. Drawing the rail through the
bottom-anchor pad so every frame has one unbroken edge: that pad is the void, and it is the case
Phase 2 was written about. And a per-column rail extent, so a short column's gutter stops early:
the gutter belongs to the boundary between two columns rather than to either of them, and one of
the two ending sooner is not a fact about the line between them.
### 9.24 the middle of the grid breathed and its edges did not
§9.23 fixed the frame's continuity. This section is about the space inside it, and about a
number that was never chosen — it was assumed, in about eighteen places, and the two halves of
it had to agree by hand.
**The pad was a literal, and so was its twin.** The margin between the terminal's edge and
anything council draws was a bare `" "` in roughly ten builders — the header, the notice, the
column grid, the tab bar, the single-column and turn-page bodies, five row shapes in the
composer, the mode line, the help panel — with its arithmetic twin, a literal `2` meaning
*pad×2*, in eight more places that subtract it back out to get a usable width. Those two
families have to agree exactly. A builder that paints more than its arithmetic subtracts pushes
the row past the terminal edge and `fit` eats the overflow in silence, which is precisely the
off-by-one §9.11 found in the header's gap.
`framePad` names it and `framePadStr` derives the string from it, so the paint cannot drift
from the sums. **The extraction shipped as its own commit with the value still 1** — every
frame byte-identical, not one golden moved — because a refactor that also changes behaviour is
a refactor nobody can check. A `- 2` that is *not* the frame pad, like `labelRule`'s two cells
of air around its rule, is deliberately left as a literal; the constant is not a licence to
unify every 2 in the package.
The extraction turned out to be **incomplete on the first pass**, and the value change is what
found it: `header`'s `pathWidth` and its affordability test were still subtracting a literal 2.
At `framePad = 1` that is indistinguishable from correct, which is exactly why it survived —
the bug is invisible until the constant moves, and it surfaced as the header clipping `no brief`
to `no brie` at 68 columns. That is the argument for the constant restated as evidence.
**One to two, because a margin narrower than the gutters inside it is the wrong way round.**
The interior of the grid gave two cells each side of every rail; the frame's own edge gave one.
So the outermost boundary was the tightest thing on screen, the room read as crowded against
the terminal, and the middle read loose — the inverse of what a grid wants. The screenshot pass
that set `gutter` to 2 named that feeling exactly ("rigid / cramped") and fixed it in the one
place it happened to be looking. `framePad` is now the same two, for the same reason, and the
room has one number for *air between things* rather than two that disagree.
It costs two cells of total width, and one of them landed somewhere worth recording: **at 80
columns the view-mode footer came out one cell over.** This room sheds whole cells rather than
clipping words (§9.18), so `f expand` becomes the second rung of the shed ladder after `[ ]`.
`f` and not `tab`: `tab` is how a reader reaches the other seats at the tabbed tier, which is
the only tier this bites at, so shedding it would strand them on one column — while `f` is the
cell §9.11 already ranked lowest, on the argument that it expands a column to a width it
already has. Adding a second rung also made the shed *order* load-bearing for the first time,
so it is now stated — **shed order is list order** — rather than left to a backwards walk that
read as "newest first" and was not.
**stripColumn goes 14 → 18, from an arithmetic floor to a reading width.** Fourteen was
derived, and derived correctly: the widest phase word is nine cells, its mark costs two, and
the remaining three are exactly a two-letter vendor tag and its space (§9.18). That answers
what a strip's *header* cannot go below. It says nothing about the prose underneath, and prose
is most of what a strip draws.
At fourteen the prose shredded. §9.19's coalesced skip line — on most turns the **only** content
a backgrounded seat has — came out three rows deep as `○ not` / `addressed in` / `turn 4`, with
the phrase that carries the meaning split across two of them. `last: turn 8 ✓`, which §9.19
introduced with "room" as its stated goal, wrapped in a long room. A column whose every line
breaks mid-phrase is not narrow, it is unreadable, and the entire point of keeping these seats
on screen (§9.18) is that a reader takes them in at a glance.
Eighteen is the smallest width that puts `○ not addressed` and `last: turn 137 ✓` each on one
line. The header floor still holds — fourteen is still where the header itself would break, so
eighteen clears it by four and §9.18's shedding ladder is untouched. The four cells come out of
the primary column, and `weightedWidths` refuses the weighted split outright rather than ship a
primary under `minColumn`, so at a frame narrow enough for four cells to matter the room falls
back to equal columns instead of trading a readable strip for an unreadable seat.
**The change paid for itself in rows.** `skips-coalesced.txt` is the clearest reading: with each
block a row or two shorter, the same body height now holds seven more turns of transcript, and
the overflow marker went from `↑ 8 more above` to `↑ 1 more above`. Wider columns showing *more*
content is not the trade anyone expected from spending cells, and it is what happens when the
alternative was spending three rows to say four words.
**What was declined.** A width-dependent pad, so narrow terminals keep one cell and wide ones
get two: the tier ladder already varies what is *said* by width, and varying the frame's own
geometry as well would make two different rooms out of one resize. Trimming the footer by
clipping instead of shedding, which is the trade §9.11's whole footer pass exists to refuse.
And unifying every literal 2 in the package behind the new constant — `labelRule`'s air around
its rule is the same number for an unrelated reason, and tying them together would mean a
future change to one silently moving the other.
### 9.25 the panel that lists what the room can do was not listing it
Three of the four items here are the same defect wearing different clothes: a surface that
knew something and did not say it. The fourth is a surface that said something it did not know.
**The help panel clipped in silence, and it was the only place in the room that did.** Every
other surface spends a body row on `↓ N more below` when content does not fit, on the explicit
argument (§9.11, columnCell) that silent clipping is indistinguishable from there being nothing
more to say. The help panel is 24 rows on page one and 33 on page two against a hard budget of
17, so at the reference machine's own geometry it was dropping seven lines and sixteen — with
nothing on screen to say so, and dropping them mid-word: `…the containment, not a`. A panel
whose whole job is to enumerate what the room can do, quietly not enumerating it, is the
sharpest available version of §4a.1's rule.
**The marker's row is paid for, and the way out is pinned.** `?` is the only documented way back
out of this panel, and on both pages it sat at exactly row 17 of a 17-row budget — so a marker
taking the last row the ordinary way would have bought honesty with the exit, which is the trade
§9.11's footer pass and helpKeys' own budget comment both refuse by name. The exit is now
**chrome**, pinned to the last row the way `columnChrome` sits above a transcript, with the
marker inside the scroll below it. That makes the guarantee structural instead of a lucky row
count. The marker's own row is paid for the way this panel has always paid — by merging two
lines that were one category: `ctrl+j` and `esc`, the two compose keys that are not `enter`, one
extending the draft and one leaving it alone. Nothing was dropped to make room.
**The marker names no key, and neither does the mode line.** `↑↓` do nothing over the help panel
— `key()` routes no scroll to it — so the room was advertising an arrow that does literally
nothing in the mode a reader is in *when they went looking for what the keys do*. Wiring a help
scroll offset was the alternative and it was declined: it buys reachability for a page whose
overflow is a paragraph of prose, at the cost of new state, new key routing and a new §7.1 rule-4
surface, when the honest sentence — *there is more, and this terminal is not tall enough* — costs
one row and no mechanism. So the panel's mode line names only what works there (`?`, `i`, `q`),
which is §9.11's own footer rule applied to a mode it had not been applied to.
**The title got the room's grammar.** `council — one brief, several agents, side by side` was the
only heading in the product with no rule on it, while the column header, every turn separator and
every seat rule on a turn page all draw `labelRule`. A rule *under* the title is what one might
expect and it is not what this room does: §9.11 spent a whole item removing exactly that shape on
the finding that a heading followed by a horizontal rule says nothing the heading had not, and
ruled that a heading carries its own rule. So the title becomes a `labelRule` and costs no row —
which is what made it affordable against a budget with none to spare.
**The blank above `? close` did not happen, and that is recorded rather than fixed.** It is wedged
against the sentence before it and it should not be, but the exit sits at row 17 of 17 and a blank
there comes straight out of the legend the page exists for. §9.11's ranking settles it: a rule
outranks a blank, the title now carries one, and air is the boundary strength this panel can
afford to go without. If the budget ever loosens, that is the first row to spend.
**A seat's detail hung ten cells left of its own label.** The per-seat posture section put a seat's
name at column 15, under a badge legend at column 15, and then hung the seat's measured detail at
column 6 — the child left of its parent, reading as a new statement rather than as the reason for
the one above it. Every card in this room has had one grammar since §9.11 (a title at weight, its
body hanging under it) and this was the last place still drawing the shape that rule was written to
remove. The three hard-coded numbers that had to agree — 13 for the badge column, 15 for the
legend's continuation, 6 for the body — are now one `helpIndent`, checked against its own string
form at init, because a panel whose continuation rows drift a cell from its key column is invisible
in a diff and obvious on screen.
**The vendor tag is permanent, and the wide column is now the legend for the narrow one.** §9.18
introduced `CC` / `CX` / `AG` / `CU` as what identity degrades *to* when a strip has no room for a
name. Read as a whole product that is backwards: the abbreviation a reader has to know appeared
exactly where they had the least context to learn it, and vanished at every width where the room
had space to teach it. Drawn always, `CC Claude Code` at 37 cells is the sentence that makes
`CC ✓ done` at eighteen readable, and it is the same pairing the HUD's own grid already makes.
The tag is **chrome and the name is the anchor**, so the tag is muted while the name keeps the
weight that says which column the keys move — asserted, because a tag at the name's weight would
put a two-letter abbreviation in competition with the thing a reader is scanning for. It costs
three cells of the header row and nothing else, and §9.18's degradation order is unchanged: at
widths where the header must truncate, the spelled-out name goes and the two letters stay, which
is the strip's one-step collapse performed gradually.
**Turn pages and the collapsed-seat notice keep bare names**, and that boundary is the rule rather
than an omission: the tag earns its place where columns are *scanned*, and a turn page's seat rule
and a notice sentence are prose. `CX Codex (not installed)` inside a sentence is an abbreviation
introduced where nothing is being compared.
**One stray fact.** `unavailable.txt` drew `final only` under `⚠ Codex is not seated` — a claim
about how a vendor behaves *during* a turn, stated about a vendor that cannot take one. Codex was
not found on PATH; nothing about its streaming was measured. It was *plausible* — it is what the
binary would do if it were installed — which is precisely the class of claim §4a.1 puts at the top
of its rejected list. The badge row goes empty for an unavailable seat, and the cost cell with it
(a seat that never ran cost nothing). The row stays **reserved**, because §9.11's argument for
reserving it is about the grid's rows lining up and is untouched; what changes is that a reserved
row now holds nothing rather than something invented.
**What was declined.** Wiring `↑↓` to a help scroll offset (above). Giving the help panel its own
narrower vocabulary of markers — `↓ 7 more below` is the room's existing sentence and a second one
would be the second alphabet §9.11's phase marks were built to avoid. Dropping the badge row
entirely for an unavailable seat, which shears the grid for the sake of a row that costs nothing
to keep. And putting the tag on turn pages "for consistency": consistency across surfaces that are
doing different jobs is how a room ends up with an abbreviation in the middle of a sentence.
### 9.26 one rule glyph was doing four jobs, and the header band re-textured on every dispatch
§9.23 made the frame continuous and §9.24 made its margins breathe. What neither looked at is
that the room draws horizontal lines at **one weight**, and asks that one weight to be four
different things: the frame's own edge, a column header's leader, a turn separator inside a
transcript, a seat's heading on a turn page. Every one of them is `─`, so a reader scanning for
*where does the room end and the content start* gets the same ink as a reader scanning for
*where does turn 3 begin*. A grid with no outline is a grid you have to reconstruct from its
contents.
**Two weights, one distinction: outline against interior.** `RuleHeavy` is `━` (U+2501) and `=`
in the reduced set, and it is spent on **exactly three lines** — the two full-bleed rules that
close the frame above and below the reading area, and the turn separator at the top of a turn
page. Everything else keeps `─`. Three weights would be a hierarchy nobody can hold in their
head; the value of the second one is entirely in its scarcity, which is why the list is closed
and `TestOnlyTheFrameAndTheTurnPageDrawTheHeavyRule` asserts it as a *count* on the rendered
frame rather than as a property of the three call sites.
> **Amended 2026-08-09 (§9.44).** Two of those three lines are now one. The composer is a
> bordered box, so the lower full-bleed rule is gone and the frame's closed shape is the header
> rule plus the box — closure carried by corners rather than by ink. The scarcity argument here
> is unchanged and one line cheaper; what the heavy weight says is now *the chrome stops here and
> the seats begin*. The test still asserts a count, and the count is 1.
**Why the turn page's rule is the third.** It is the only line inside the frame that bounds a
whole document rather than a part of one. §9.23 gave it the *weight* of a root — the label at
full intensity while its seat rules recede — on the finding that the page's outline whispered
while its entries shouted; this gives it the *form* of one. The grid's copy of that same line is
untouched, for §9.23's own reason: there a turn separator sits inside a column already headed by
a seat name, so it is the child. The seat rules on a page and the help panel's title stay light
for the same test — a heading *inside* the outline that matched the outline would restate §9.23's
hierarchy defect one level down.
**The weight is a parameter, not a flag.** `labelRuleIn` takes the fill glyph and `labelRule`
passes `g.Rule`; a caller that wants the heavy rule has to name it at the call site. That is what
makes "exactly three lines" checkable by *reading* the three call sites rather than by grepping
for a bool, and it keeps one implementation of the grammar — a label, a rule, optional numbers,
two cells of air each side — which is `labelRule`'s own extraction argument.
**It is a character before it is a style.** `--ascii` gets `=`, not a fallback to `-`, so the
outline survives on exactly the terminals least able to infer it; `NO_COLOR` never touched it,
because weight of this kind is a glyph rather than an attribute. `=` is the one unclaimed mark
left in the reduced set — `-` is the light rule, the `Range` joiner and the first spinner frame,
`|` the separator, `>` the ellipsis, `]` focus, `!` the warning prefix, `^`/`v` the overflow
markers, `*` Act, `.` Idle, `:` the prompt, `_` the caret, `+`/`x`/`?` the outcome marks, `/`
and `\` the remaining spinner frames, and `#` the HUD's gauge fill. It is also the only
unclaimed character that reads as a *doubled* `-` rather than as a different symbol, which is
the one property a second rule weight needs.
`TestTheHeavyRuleHasAnUnclaimedASCIIPartner` enumerates that whole list so the next glyph cannot
be added without meeting it.
**The header leader stops depending on phase.** `headerUsesLeader` was false for an idle seat, on
an argument that was true at one rule weight: a long `────` between `Claude Code` and `○ idle`
was *filling* rather than separating, whitespace does that job for free, and a room with a single
rule weight cannot afford ink on nothing. With two weights the leader is no longer "the rule" —
it is the interior weight, and its claim on that row is *this name and this state belong to one
seat*, which is as true of an idle seat as of a streaming one.
The observable defect is the sharper half of the argument. A room where one seat is answering
drew the seats' header band as one continuous ruled line across part of the frame and blank
across the rest — **one row, two grammars** — and re-textured itself the moment a turn started
and again when it ended. §7.1 rule 4 keeps this room still by default, and a band that changes
shape on every dispatch is the loudest still-frame change on screen, spent on a fact the state
word beside it already states. The air the old comment wanted is not lost: `labelRule` keeps two
cells each side of its rule, which is the gap that keeps an ascii spinner (`-`) legible against
an ascii leader (`-`).
**Golden churn is the whole visible change, and it is two lines per frame plus one.** Every
frame's two rules, and every idle seat's header row. Nothing else moved — `PlainStyles` renders
both weights as themselves because they are characters, so unlike §9.23's weight half this pass
*is* visible in the goldens and had to be read frame by frame. The three test helpers that found
the frame by searching for a run of `─` (`fullWidthRule`, `frameBody`) now search for `━`, which
makes them stricter rather than merely different: a column header's leader can no longer be
mistaken for a frame edge at any width.
**What was declined.** A third weight, or a double rule (`═`), for the turn page — the page's
rule is already distinguished from its seat rules by its label, its position and its meta, and
the frame is the only thing it needs to *match*. Making the frame's rule brighter as well as
heavier: §9.23 declined to let the rails' hue mean anything on the argument that chrome competing
with content is the wrong trade, and an outline is chrome. And keeping the idle leader off "for
quiet": the quiet was bought by making the room's most stable row the one that changed most.
### 9.27 focus was a mark on one row, in a frame the reader had scrolled past
§9.12 fixed the focus signal by adding the `▸` and moving the load-bearing half onto the seat
name's *weight*. Both of those live on the column header — row one of a body that is twenty rows
tall — so a reader forty lines into a transcript, comparing two answers, had nothing on screen at
all telling them which column `↑↓` would move. The signal was correct and it was in the wrong
place: it described a column and was as tall as a line.
**The focused column's LEFT rail thickens.** The gutter cell immediately left of the focused
column draws `▌` (U+258C) instead of `│` — same cell, same width, one glyph heavier — for the
full height of the band. It is the only mark on this surface that is as tall as the thing it
describes, which is the whole reason it is worth a glyph. The `▸` and the name's weight stay:
word/glyph-first means two carriers on two rows, not one carrier moved.
**The leftmost column has no gutter, so the frame's left pad carries its mark.** `framePad` is
two cells since §9.24, and the mark takes cell one — which leaves exactly one cell of air between
it and the column, the closest the geometry gets to the gutter's two. Without this, position zero
would be the one seat the device could not mark, and a signal with a hole in it is a signal a
reader stops trusting.
**It rides §9.23's band exactly.** The thick rail spans the rows the thin one would and no
others, so focus cannot spear a void either — an idle 120×60 room still has a bare middle, and
`TestTheRailRidesTheSameBandTheThinOneDoes` asserts it against the same `bare > 0` test §9.23
wrote. Focus does not get its own answer to a question the frame already settled.
**Unfocused columns' prose steps back one contrast level.** `Dim` is `Text` + `Faint`, applied to
the *reading area* of a column the keys do not move: the vendor's reply, the §9.14 stand-in for a
reply that has not arrived, and the `no turn dispatched yet.` line. That is crush's
`Focused`/`Blurred` pair applied to prose rather than to a border, and it is the half of this
pass that costs no cell at all.
**The faint collapse is accepted, and here is the accounting.** Council has two intensities —
`Text` and `Muted` — so a demoted body renders identically to chrome, and inside an unfocused
column prose and chrome do arrive at one intensity. What is lost is the *second* signal, on a
column the reader is not reading: every distinction between them is carried by shape first (a
turn separator is a labelled rule, a trace entry opens `⚙`, a skip line `○`, a note `⚠`), which is
§7.1 rule 2 doing exactly the job it was written for. The alternative — a third intensity in
`internal/theme` — would spend a shared palette token, on a surface the statusline does not have,
for a distinction only the unread column needs.
**What the demotion does NOT reach, and each exclusion is a rule rather than a taste.**
- **The chrome above the body.** `columnCell` renders the header, the badge row and the gate card
with the room's set and only the body with the seat's. A posture badge is a safety claim, and a
claim that faded because the reader was looking at the next column is precisely the defect §9.2
wrote the reserved badge row to prevent.
- **The prompt echo.** The user's own words stay `Strong` in every column. What a seat was *asked*
is the thing a reader scrolls looking for (§9.9), and it is not the vendor's prose to demote.
- **Notes and cards.** A failure note, a reattach card, an unavailable card and the thread-cleared
sentence under its rule all keep their own styles. This is the one place the ratified shape was
**narrowed** during implementation: the thread-cleared sentence is prose in the reading area by
position, but it is the body of a card in the room's grammar and it says what the *next* brief
will do — an actionable claim about the seat, in the same category as the reattach card whose
wording it shares. Leaving one of that pair full-contrast and demoting the other would be two
spellings of one fact.
**The rail is a columns-tier device, and says so.** The tabs tier has one column on screen with a
tab bar above it already carrying `▸` and the selected tab's weight; a rail there would mark the
only thing there is. Expanded is the tabs tier by `tierFor`'s own rule, so it inherits that
answer rather than needing its own. A turn page is one reading area and has no unfocused seat to
demote.
**Under `NO_COLOR` and `--ascii` the whole distinction still lands**, and that is the test the
demotion had to pass to be allowed at all: `▌`/`[` in the gutter, `▸`/`]` before the name, and the
name's own weight all survive both, so a monochrome terminal loses the contrast step and keeps
every carrier that was doing the work. `[` is the ascii rail — `#`, the obvious candidate, is
refused for the reason `ActOK` refused it (it is the HUD's ascii gauge fill, and one product means
one vocabulary), and of what is left `[` is the squarest vertical stroke in the set, faces the
column it marks, and mirrors `]`, which is already this room's ascii focus mark. The `[` and `]`
in the mode line are key *names* in the footer's prose, never marks in the grid — the same slot
argument `Range`'s doc makes for the hyphen.
**Golden churn: the rail only.** `▌` is a character, so every columns-tier golden moved by exactly
one cell per railed row; `Dim` is an attribute rendered by `PlainStyles` as the identity function,
so it moved nothing. The whole diff was verified mechanically — every added line with the rail
glyph mapped back to a space is byte-identical to the line it replaced. One golden is **new**:
`focus-rail.txt` pins the focused column in the *middle* of the frame, the shape no pre-existing
golden reached because all of them render with focus at position zero.
**What was declined.** A rail on both sides of the focused column, which is a box and turns a
gutter shared between two seats into a property of one of them (§9.23's own last item). Colouring
the rail: chrome that competes with content is the trade §9.23 refused. Running the thick rail the
full body height so focus always has an unbroken edge: that is the void again. And demoting
`Muted` chrome a further step in unfocused columns, which would need the third intensity this
section just declined to buy.
### 9.28 the room's one hue exception, and exactly how far it goes
`internal/council` has said "adds no hues of its own" since §9.11, and the rule was right: a
dispatch room that invented a sixth colour drifts from the visual language the statusline and the
HUD share, so council spent WEIGHT (§9.11) and CONTRAST (§9.27) instead, both attributes rather
than hues. **This is the one ratified exception (San, 2026-08-07), and it is an exception to the
rule rather than a repeal of it.**
**The concept the other two surfaces do not have is the SEAT.** Everything council renders that
theme already has a token for — severity, identity, chrome — keeps that token. What has no token
is *which of four agents is speaking*, because `telltale statusline` and `telltale hud` have no
seats to distinguish. `seatHue` returns one ANSI index per vendor: claude `5` (magenta), codex `6`
(cyan — theme's identity hue, kept by the seat that already had it), agy `4` (blue), cursor `12`
(bright blue), and `theme.ColorIdentity` for anything else.
**Why it lives in `internal/council` and not in `internal/theme` — and the stdlib rule is NOT the
reason.** These are plain strings; they would compile in theme perfectly well, and citing ADR-002
here would send the next reader to fix the wrong thing. The reason is theme's *own* contract: one
hue, one meaning, across every surface that imports it. A per-vendor hue promoted to theme is a
token that means nothing on two of the three surfaces, which is how a shared palette stops being
shared.
**Why 4-bit indices.** theme.go's own argument, unchanged and reused rather than restated: the
terminal resolves an index against the scheme the user already chose, so the room looks native in
Windows Terminal's default and in a light scheme with no second palette and **no `isDark` fork**.
A hex triple would be council asserting a colour over the user's own.
**What is off limits, and it is a fence rather than a guideline.** The severity family — `1`/`2`/`3`
and their bright twins `9`/`10`/`11` — is the green/yellow/red ramp on every surface, and a seat
wearing red would read as a seat that failed, on a row where `✗ failed` is the thing beside it.
The chrome family — `0`/`7`/`8`/`15` — is the gauge track and the terminal's own fore/background.
That leaves 4, 5, 6, 12, 13, 14; this spends four of them, and `TestNoSeatHueIsASeverity` fails
the build if that stops being true.
**The honest weakness: 4 and 12 are one hue at two intensities.** agy and cursor are blue and
bright blue, which some terminal schemes render close together and a reader can miss. That is
acceptable **here and only here**, because §9.25 made the two-letter tags permanent — `AG` and
`CU` appear beside every seat name the room scans — so the hue is the second signal it is supposed
to be and the tag is carrying the distinction. If a fifth seat arrives wanting blue, the tag is
what still works and the hue is what has to be argued for.
`TestSeatHuesAreExhaustive` asserts the room seats exactly four vendors, so a fifth cannot be added
without somebody reading this paragraph.
**Three sites, and the list is closed.**
1. **A turn page's seat rules** (`seatRule`). The highest payoff by a distance: a page stacks every
participating seat in one column, one block after another, so position answers *nothing* about
who is speaking — which is the exact condition under which a hue earns its place.
2. **The tab bar.** `SeatStrong` selected, `SeatIdentity` unselected, replacing the wholly-muted
unselected tab. That is a *promotion*, and the opposite of what §9.27 does to an unfocused
column's prose, deliberately: prose in a column you are not reading is content you are not
reading, while an unselected tab is a **destination**. It is the one row on that tier whose job
is "here are the other seats, pick one", and a menu whose entries are faint makes you read it
twice. The selected tab still outranks the rest by weight and by the `▸` in front of it, which
is what survives NO_COLOR.
3. **The collapsed-seat notice**, names only. The `⚠` keeps `SevWarn`, the reason in parentheses
and the remedy after the bar stay chrome, and nothing there gains weight — it is a sentence, and
a sentence with four bold words in it is not one. §9.25's boundary is untouched: the two-letter
*tag* stays out of prose, because an abbreviation introduced mid-sentence is one nobody can
learn there. A hue is not an abbreviation — it costs no cell and teaches nothing new.
**Where it is deliberately NOT spent.**
- **Grid column headers.** Position already answers which seat this is, and four coloured names
across one row is the circus row this rule exists to prevent — the room's newest signal spent on
the one question the layout had already settled.
- **Phase marks and status words.** Severity owns those cells (§9.7).
- **Rules, leaders, badge rows and every other piece of chrome.** A posture badge is a safety claim
(§9.2) and must not compete with a name for the eye.
**Constructed to be invisible to the goldens, rather than checked to be.** `SeatIdentity` and
`SeatStrong` are `Identity` and `Strong` *retinted*, through one `retint` helper that returns the
base style untouched when `Plain` is set. A second pair of literal constructors would have to
remember that and would forget it the first time one grew a second attribute. So **golden churn on
this pass is zero, and any golden diff on it is a bug** — which is also the whole verification
story, since colour is asserted where colour is asserted (§9.5) and never in a golden.
**What was declined.** A hue on the grid's column headers (above). Hue on the vendor tag as well
as the name, which doubles the ink for a distinction the name already carries. Truecolor, which
would override the user's scheme. And a fifth hue held in reserve for "the next vendor": a palette
entry with no seat behind it is a decision nobody has made, recorded as if somebody had.
### 9.29 the seats had positions and no way to address one
`tab` cycles focus, and at the columns tier that is fine: three seats, at most two presses. At
the **tabbed** tier — the narrow terminal, the one a laptop actually runs — one column is on
screen and reaching the fourth seat costs three presses, each of which redraws the whole frame
and shows you a seat you did not want. The room had four seats sitting in a fixed order, drawn in
that order on every surface, and no way to say *that one*.
**`1`–`4` focus the Nth VISIBLE seat, in seating order.** Positional, exactly like the columns
are. A room with two seats has keys 1 and 2 and nothing else: `3` there is a **no-op**, not a wrap
and not a clamp, because a key that quietly lands somewhere else is §7.8's surprise and a wrap
would make the number stop meaning the position it is printed at. In **compose** a digit is a
digit — the same contract `q`, `f`, `c` and `[` already keep, and it needs no second list: the
handler tests whether the key carries text, which is what makes it text in the composer.
**The number is drawn where the key acts.** `▸ 1 CC Claude Code ──────── ✓ done 8s` in the seat
header, and `▸ 1 CC Claude Code 2 CX Codex` on the tab bar — the two places a seat name heads a
reading area, which are the two places §9.25 already put the vendor tag for the same reason. The
number is **muted**, on the tag's own argument: it is chrome and the name is the anchor. It sits
in FRONT of the tag rather than after the name, because it is what a reader's eye runs down the
row of headers looking for, and because a number at the far right would sit beside the state word
where every other number on that line is a duration.
**It sheds last, and that is a new rung reasoned about rather than an appended default.** §9.18's
ladder drops the clock first, then the focus mark, then the tag. The number goes below all three:
`1 CC ⚠ unavailable` is exactly eighteen cells, so at `stripColumn` the full form fits every phase
word, and below that the tag goes before the number does. The argument is the one §9.18 itself
used — it shed the focus mark because the load-bearing half of that signal had moved somewhere
free, and it kept the tag because position alone was a weak identity. The number is not a second
spelling of anything: it is the key that reaches this seat, at the width where reaching seats is
hardest, and **a key nobody can see is a key nobody presses** (§9.10, which is the whole reason
this room names keys on its overflow markers at all).
**The footer names `1-N`, not `1-4`.** The range is however many seats are on screen. A
three-seat room naming a `4` would promise a key that does nothing — §7.8's surprise, which this
line already refuses in the other direction for `tab` and `f`. It is the **third rung of the shed
ladder**, appended after `[ ]` and `f` so §9.24's order is untouched, and it is the last of the
three to go: shedding only bites at the tabbed tier, which is precisely where the number is worth
most. `[ ]` sheds first because `g` and `G` still reach the ends of the transcript; nothing else
reaches seat 4 in one keystroke. `? help` and `q quit` remain unsheddable, asserted.
**A room with ONE seat on screen has no numbers at all**, and that is §9.11's rule applied to a
third key rather than a special case: `f` and `tab` are dropped there because they address a
choice that does not exist, and a number labelling the only column there is spends a cell on the
same nothing. `State.SeatNumber` and the footer's cell run off the same predicate, so the key's
label and its advertisement appear and vanish together — a footer naming a key the header did not
would be one surprise split across two rows.
**Renumbering, and the still-by-default wrinkle it is.** Because the number is a position, a seat
folding out **renumbers** every seat after it: with Claude uninstalled, Codex is seat 1. That is a
label changing under a reader, which §7.1 rule 4 does not hand out lightly — and it is bounded by
*when* it can happen. A seat collapses, or `--vendor`/`/seat` reseats the room, and both already
reflow the entire frame: the column widths change, the notice row appears, the grid is visibly a
different room. There is no path where the numbers move on a frame that was otherwise going to
look the same, and in particular none mid-turn. The help panel says "by position" rather than
implying a seat owns its number.
**The help panel merged, not grown.** `tab / 1-4` on the row that already named `tab`, because the
budget is hard at 17 rows and these are one question asked two ways — step to the next seat, or go
straight to one. "move" paid for the characters. The `?` row, the panel's only documented way out,
is exactly where it was, and the `↓ 5 more below` marker's count is unchanged.
**What was declined.** `alt+1`–`4`, so digits could stay digits in both modes: it buys nothing —
compose already routes text keys to the draft — at the price of a chord nobody discovers and that
several terminals eat. Numbers on a turn page, which has one reading area and no focus to move.
Stable per-vendor numbers that never renumber, which would leave gaps (`1`, `3`, `4` on screen)
and make the printed number disagree with the position it is printed at — the number would then
be an identity, and identity is what the tag and the hue are for. And a fifth key for a fifth
seat: `1-N` already says how many there are.
### 9.30 one question, asked once, instead of four times across the comparison surface
Council exists to put several answers side by side, and §9.22 built the page that reads a turn
as one document. What neither of them fixed is what the **grid** does with the question itself.
§9.9 ruled that the echoed brief is a fact about the COLUMN — a turn can reach two seats and not
a third, seats skip turns, and a transcript that filled the gaps would be the room inventing a
conversation — so every addressed column echoes it. On a one-seat route, which is the ordinary
turn since the default stopped being everyone, that is exactly right. On a committee route it is
the same paragraph two, three or four times across the top of the reading area, each copy pushing
the answer it belongs to a row further down, with "+ the other seats' last answers were quoted to
this one" repeated underneath every one of them on a rebuttal turn. The surface built to compare
four answers spent its widest rows agreeing with itself about the question.
**So the live turn's brief is drawn once, full width, as a band under the room chrome, and the
addressed columns stop echoing it while it is up.** Nothing new reaches `State`, nothing is
stored, and no column records anything different: this is a rendering rule over the same echo
§9.9 already holds, sanitized and deliberately unredacted for §9.9's own reason — it is the user's
own typing shown back to the user, and covering it would hide a secret from the one person who
already has it while doing nothing about the copy just sent to three vendors.
**History is untouched, and that is the boundary rather than a scope limit.** A finished turn's
echo stays inside the column that took it, because §9.9's argument is about a *record*: turn 4
is filed on the two seats it reached and absent from the one it did not, and each column's
transcript is that seat's own conversation read top to bottom. The band speaks for one turn — the
live one — and it identifies that turn by number (`Column.TurnN == State.Turn`), never by "this
column has a prompt". A column's prompt block outlives its turn: a seat that answered turn 4 and
sat out 5 and 6 is still displaying turn 4's brief as its current block, and that is its own
conversation, not the live turn.
**It appears at dispatch and retires at the next one.** Both are keystrokes, and §7.1 rule 4's
still-by-default frame is about what moves *without* one. The retirement moment is deliberately
the push to history and **not** the instant the last column lands: §9.21 retires the live route
there because the header describes the present, but that landing is a vendor finishing, and a band
that vanished on it — restoring three per-column echoes and reflowing every column — would be
precisely the mid-turn layout jump the rule forbids. Tying the band's life to the same block the
per-column echo already had means there is one reflow point, it is the user's own enter, and the
frame either side of it is one the reader asked for.
**Two seats is the threshold.** One echo on screen is not duplication, and hoisting it would move
the user's words away from the answer to them and buy a row of chrome for nothing. The band is
therefore a columns-tier device: the tabs tier draws one column at a time, `f` resolves to that
tier, and a one-seat room has nothing to compare — in all three the column keeps its own echo. A
turn page is excluded for the opposite reason: it already prints the brief once, which is half of
what it is for. The help panel is excluded because it replaces the column area outright.
**The anatomy is §9.9's echo hoisted, and §9.11's middle boundary under it.** The composer's own
`›` at full weight, because the glyph carries "you said this" before the colour does and the user's
words are the anchor a reader navigates by; the rebuttal notice once, muted, underneath, only when
true, in the same sentence the column prints (one constant, not two spellings); then a **blank
row**. Of the three boundary strengths this room ranks — a labelled rule where the turn changes, a
blank where the speaker changes, a blank where the kind of content changes — the band's is the
second: the user stops and the seats start. A rule was refused. The frame's own full-bleed heavy
rule sits two rows above it, and a second horizontal line under that would rebuild §9.11's "one
rule per column instead of two three rows apart" at the room's own scale, with §9.26's
heavy/light distinction blurred as well. The band also states **no route**: the header already
carries `turn 10 → everyone` on the cell that names the turn, and repeating it would be the second
copy this whole section exists to delete. What each column keeps is its own turn separator, which
is one line saying which turn the lines under it belong to rather than the same paragraph again.
**Four rows, and the fourth is the marker.** A brief worth sending to three agents can be a
paragraph — the composer grows to six rows for that reason — and a band as tall as the draft would
eat the reading area it was written to protect. So the band spends at most four rows on the brief,
and when it needs more the fourth row is a **truncation marker** rather than a fourth row of text:
how many rows are missing, and that the turn page has the brief whole. Silent clipping is the
ambiguity §4a.1 forbids, and it is worse here than anywhere else on screen — a reader cannot tell
their own question from a truncated copy of it. The `t` that opens that page is named on the
marker in **view mode only**, because `t` is the letter t while composing and a marker advertising
it there would promise a keystroke that does something else (§7.8, scrollHint's rule for `f`). The
count and the destination survive in both modes; only the keystroke sheds.
**The band's rows are room chrome, spent where the notice line is spent.** `resolveLayoutIn`
settles the tier first — §9.5's ordering, unchanged, and the band depends on the tier so it could
not be spent any earlier — then header, footer chrome, tab bar, collapsed-seat notice, **band**,
then the composer, which still yields before the body. The band is budgeted before the composer
and tested against the composer's *floor* rather than its current height, deliberately: a band
that retired because the draft grew a row would be a layout jump on a keystroke mid-turn, and it
would jump back on backspace.
**Below a floor the band yields ENTIRELY, and the fallback is a pure function of height.** If
spending the band would leave the columns fewer than eight body rows, it is not spent at all and
the columns echo the brief themselves — the pre-band frame, byte for byte. Eight is measured from
what a column draws before a word of the reply: three rows of `columnChrome` (name, posture claim,
one blank) and the live turn's own separator, leaving four rows of answer. All-or-nothing rather
than shedding a row or two off the band, because a half-band is the worst of both: a cut question
above columns that no longer say what they were asked. One number decides it — `Layout.Band` — and
the renderer and the columns both read that one number, so a band without suppression (the brief
three times *and* at the top) and suppression without a band (the brief nowhere on screen at all,
which is a §4a.1 failure with the user's own words as the missing content) are not states this
code can reach.
**Scroll detaches from the band, on purpose.** A column scrolled back into history draws its
history under the band exactly as it always did, and the band stays — including when *every*
addressed column has been scrolled away. It describes the live turn, which is a fact about the
room, not about any viewport; a reader who went looking for an older answer is still in the turn
they dispatched, and the question that produced it does not stop being true because they scrolled.
**What was declined.** Hoisting past turns' briefs the same way, which would flatten §9.9's
per-column record into a room-level one and lose which seats a turn actually reached. Keeping a
one-line stub in each column to mark where the brief would have been: it costs the row the band
was spent to save, once per column, and the turn separator already marks that boundary. Shedding
the band down to one row on a short terminal, covered above. And giving the band a rule glyph or a
hue of its own — council adds no hues, and the boundary vocabulary this room already has is what a
reader has already learned to read.
### 9.31 a word the room did not know was billed to three vendors
**The rule: a draft that opens with `/` and names no room command is refused, not dispatched.**
Refusing is free — nothing spawns, nothing is billed, the draft stays in the composer — and the
alternative is not free at all.
**The field report.** Turn 53 of a real room: `/unseat codex` was typed. There was no `/unseat`;
§9.17 shipped `/seat ` and nothing else. `roomcmd.go` recognised no command, so the draft
fell through **as a brief**, and the committee was billed to discuss the string `/unseat codex`
until the user cancelled the turn. Nothing malfunctioned. Every line of that was the documented
behaviour: "only a draft that IS a command is intercepted; anything else, including text that
merely starts with a slash, dispatches to the vendors as typed."
That fall-through was the right call for the *vocabulary* question and the wrong call for the
*typo* question, and the two had never been separated. The vocabulary rule exists so the room does
not steal words out of the conversation — the argument that kept `/clear` out of `roomcmd.go`,
since `/clear` is a real Claude Code command a person means for a vendor. It says what must not be
**executed**. It says nothing about what should happen to a draft that executes nothing, and
dispatching was only ever the default that was already sitting there.
**A leading slash is almost never prose.** It is a command the room does not have, a command a
*vendor* has, or a typo for one of ours. Against that, dispatching costs a turn on every seated
vendor — on the scarce independence pool as readily as on the cheap lane — for a line the user will
retype in five seconds. `addressesRoom` is the whole test: a slash in **column one**.
#### The escape hatch is one space, and it had to be
A brief that legitimately opens with a slash is a real thing to type — a POSIX path, a regex,
`/etc/hosts is wrong`. Prefixing one space sends it, unmodified, to the vendors.
It is the cheapest honest escape available, and it is honest because nothing between the composer
and the spawn trims it. `sanitizeKeepingSpace` deliberately does not trim ("trimming would make the
string on screen disagree with the string about to be dispatched", §7.14's rule applied to the
composer), `ParseRoute` returns an unconsumed draft unchanged, and `dispatch` echoes what it sends.
The space the user typed is the space the seat receives, and
`TestALeadingSpaceSendsASlashBriefToTheVendors` asserts that at the seat rather than at the parser.
**The refusal has to say so, in few words.** §9.17's own defect shape is a refusal whose remedy is
undiscoverable — the `/flow` write-hop notice that went on naming a flag for two releases after
`/write` made it wrong. So the notice carries three clauses **in this order**: what failed, how to
send it anyway, then the vocabulary.
> `no room command /unseet — a leading space dispatches it · /cd /flow /read /seat /trace /unseat /write`
The order is load-bearing. This notice replaces the entire hint stack on the mode line, which
truncates from the *right*, so the clause a narrow room loses has to be the one a reader can get
elsewhere. `?` lists the room controls; nothing else on screen teaches the space. The quoted word is
capped (`unknownVerbEcho`) for the same reason — a pasted 200-character path is one word, and an
uncapped echo would push both the remedy and the vocabulary off the end, leaving a refusal that
names only the mistake.
**The vocabulary in that notice is walked, never written twice.** `roomVerbs` is the one table; the
notice reads it and `TestTheRefusalListsTheLiveCommandTable` walks it. A hardcoded list in either
place is the copy that goes stale on the next command — and this feature would have been its first
victim, shipping `/unseat` with a refusal that did not mention `/unseat`.
**What this does to the bare-word rule, which is the one deliberate consequence.** §9.17 made
`/read` and `/write` bare-only so that "/read the design doc first" could not silently swallow a
turn and run a setting. That rule is untouched: those drafts still do not reach `postureCommand`.
What changes is where they go instead — refused with the space named, rather than billed. Both
halves are now one test (`TestBareWordOnly`), because either alone is a defect: a "/read the design
doc" that ran the setting, and a "/read the design doc" that cost three seats a turn.
**`/flow` came along with it.** `dispatch.go` matched `strings.HasPrefix(TrimSpace(draft), "/flow")`
— any draft whose first non-space characters were those five letters. So `/flowchart the auth path`
was an orchestration, and, worse, `" /flow/gate.log is the file I mean"` would have been swallowed
*after* being escaped, making the hatch a lie for exactly one prefix. An escape hatch with an
invisible exception is not one. `isFlowCommand` applies the room's single vocabulary rule there too.
#### `/unseat `: `/seat` spelled the other way round
The typo that started this was reaching for a control that should have existed. `/seat` names who
**stays**; `/unseat` names who **leaves**, and it is the argument `-@` makes one control up: the
correction a user reaches for mid-session is "not that seat" — one vendor is answering badly or
expensively and the other three are fine — and making them retype the complement is arithmetic done
at the keyboard, on the one line where getting it wrong quietly reseats the room around seats they
did not mean.
It is `parseSeatList`, literally: same aliases, same `@` tolerance, same trailing punctuation, same
dedupe. A second list parser is how `/seat agy` would work and `/unseat agy` would not.
Everything §9.17 ruled for `/seat` holds unchanged, and mostly by sharing the code rather than by
sharing the intention:
- **It kills nothing.** An unseated seat keeps its thread, its process and every id that would
resume it. `/seat all` puts it back mid-conversation, with no resume to fail.
- **It refuses mid-turn.** The roster is dispatch state — `frameOwnersFor` decided this turn's grid
— so reseating under a live turn would redraw the room around columns that are mid-answer.
- **It warns when it removes the default route**, because unaddressed briefs go to claude. The
warning lives in `applySeats`, shared with `/seat`, precisely so the subtractive spelling cannot
be the one that says nothing.
- **Bare `/unseat` reports**, the way bare `/cd`, `/trace` and `/seat` do: a command that half-asks
a question answers it rather than doing something.
Three refusals are its own. **The last seat**: a room with no seats can answer nothing, so the
subtraction that would empty it is refused in `/seat`'s words for `/seat`'s reason. **`/unseat
all`**: a sentence someone will type, answered as the empty room it names rather than left to
`parseSeatList`, whose honest report would be "no seat called all" — a spelling complaint about a
word the room understands perfectly well. **A seat that is not in the room**: distinguished from a
typo, and both change nothing, on `/seat`'s argument that a command which quietly did less than it
was asked is discovered several turns later as a seat still answering.
**Membership is what the room SHOWS, not what it can drive**, and that line took a CI failure to
find. `/seat cursor` *forces* an uninstalled seat on screen — "a user who asked for it is owed the
card explaining why it is not there" — so that seat is in the room in every sense a subtraction
cares about, and the first spelling, which tested membership with `seatsVendor`, could not remove
the one card a user is most likely to want gone. On a machine where nothing is installed it could
not remove anything at all: every `/unseat` was answered "not in the room" and the roster never
moved. Local runs passed because the developer's machine has four vendors on `PATH`; CI, which has
none, is the one that reads the rule as written.
**So the two questions are separated, and only the second guards the last seat.** *Is it in the
room?* is `shows`. *Can it answer?* is `Avail == AvailInstalled`, and the refusal built on it is
**conditional on the room having had one**: a room with nothing installed could not answer before
this was typed either, and refusing there would blame `/unseat` for a state it did not cause. An
empty roster is refused unconditionally, because that is `/seat`'s own "at least one seat" reached
by subtraction.
#### The focus bug underneath it
`/seat` has been able to unseat the **focused** column since §9.17, and nothing moved focus when it
did. `State.Focus` indexes `Columns`, so it went on pointing at a seat the grid no longer draws: the
focus mark vanished from the room, and `f`, the scroll keys and `y` went on addressing the hidden
column. Keys that still work over a transcript nobody can see are worse than keys that stop, because
nothing on screen says anything is wrong.
`stateWith` already does this once, at launch — "focus lands on a column that is actually drawn" —
and the fix is that same rule applied wherever the roster moves (`rehomeFocus`), not a rule of
`/unseat`'s own. It is called from `applySeats`, so `/seat` gets it too; a helper that fixed only
the new command would have left the older one holding the bug that made the new one worth writing a
test for.
#### The help row
`/unseat` merged onto the row `/seat` already holds, as `/seat /unseat `. The panel's budget is
hard — 17 rows to the `?` line on a 24-row terminal — and `helpBody` clips without scrolling, so a
control named past the fold is not a demoted control, it is an absent one (§9.20). The merge is also
the honest shape rather than only a saving: the two take one argument in one vocabulary and differ
only in direction, so a reader who finds either has found both. "times" paid for the width — the row
is a list of controls, and `/trace ` is unambiguous without the verb.
#### Amendment, 2026-08-17: `/retry` sends the last brief again, to the seats that owe an answer
**The gap.** §9.37's 2026-08-17 amendment gave the operator a way to stop ONE seat of an ordinary
turn, and it argued from the room's most probable live failure: five seats on an `@all` turn, four
answers, one vendor that fails or stalls. The room can now end that turn. It has no way to finish
the brief. The only act left is to retype the brief and retype the mentions — arithmetic at the
keyboard, on the one line where getting it wrong bills seats that already answered. That is the
same complaint `-@` and `/unseat` were built for, one turn later.
**`/retry` puts the last dispatched brief back in the composer, addressed to the seats that did not
answer it.** The brief comes from the columns' own per-turn record (`Column.Prompt`), unchanged. The
mentions are the seats that owe an answer, so the draft reads `@codex @agy ` — a draft the
operator could have typed, in the grammar that already exists.
**It arms; it does not dispatch, and enter is still what spends the money.** The verb writes a
draft, `setDraft` re-derives the route from it, and the footer prices that route through the same
`State.SeatsIn` intersection dispatch gates on (§9.21). So the operator reads the bill before paying
it, and can edit the draft — drop a seat off the front, fix a word — because it is an ordinary
draft. A verb that spawned on the spot would spend up to five quotas on a keystroke that named none
of them, and the room would have no surface left on which to say which ones.
**What counts as an ANSWER is defined against the four endings**, and it is narrower than "the
column looks empty":
- **`PhaseDone` answered**, including the seat whose body reads `[Turn completed with 0 text chunks
streamed]`. That is a measured zero, not a missing reply (§4a.1), and re-sending on it would be
the room overruling a vendor's honest empty answer — and billing a seat that did the work.
- **`PhaseFailed` and `PhaseCancelled` did not.** Cancelled covers `ctrl+c` and the per-seat
give-up together, deliberately: a seat the operator cut is the case this verb exists for, and
after the turn the room holds one cancelled phase for both.
- **A seat that SAT THE TURN OUT is not a candidate at all.** `dispatch` never calls `startTurn` on
it, so its `Column.TurnN` still names the last turn it took and the scan skips it. That is
load-bearing rather than incidental: without it, a `/retry` after an `@codex` turn would widen the
bill to four seats the operator deliberately did not address.
**Bare-only, on `/read` and `/write`'s rule.** The verb takes no argument, and "/retry the failing
test" is a sentence someone types. A verb that swallowed that argument would run a re-send and
discard the brief — §9.17's vanishing-brief failure. The bare draft is the command; anything longer
is refused with the space escape named, which costs nothing.
**Three refusals, three sentences, and only the first keeps the draft.** *A turn in flight*: the
phases this verb reads are not settled, so any list it produced would be a claim about a turn that
has not ended — and the operator still wants the verb one turn later, which is `postureCommand`'s
own reason for holding the draft. *No brief on record*: turn 0 and the degenerate turn whose brief
sanitized away to nothing are one sentence, because they are one fact. *Every seat answered*: there
is nothing to re-send, and saying so beats a composer the operator has to clear by hand. The verb
sits in `roomVerbs` like every other, so §9.31's walked refusal teaches it for free — no second copy
of the vocabulary was added, and the notice fits the reference width with one cell to spare.
**`--brief` stays unfiled, and this verb is careful not to file it.** It re-sends the brief
UNCHANGED: no re-briefing, no edit, no automatic second attempt, and it never reads `Model.brief`.
§9.17's sweep left `--brief` out on the ruling that first-turn context is a different feature from
re-briefing; that question is still open and still to be decided on its own.
**After a race it re-sends the race's brief as an ORDINARY turn.** `Column.Prompt` holds the brief
the racers were given and never the `/arena` draft that wrapped it, and this verb invents no grammar
to put the wrapper back. The composer shows exactly what will be sent before enter, which is where
the operator reads that the worktrees are not part of it; `/arena ` races again.
**The help panel does not name it**, on `/adopt` and `/arena drop`'s precedent: the room-controls
row is at its budget (§9.20), and the verb is taught by the slash refusal and by this block.
Verified offline only. `roomcmd_test.go` pins the four endings producing the right seat list, the
measured zero counting as an answer, the three refusals, the bare-word rule, the verb's appearance
in the walked refusal table, and — at the dispatch level, with `countSpawns` — that the re-send
spawns one process per seat owing an answer and none for a seat that already replied. The
`slash-refusal` goldens and their `--ascii` twins carry the new word. No test here spawns a vendor.
A live `/retry` on the Windows reference box is not owed as a separate payment: nothing here is a
claim about vendor behaviour, and every process the verb can cause is an ordinary turn's spawn.
### 9.32 the room remembered where it was and forgot who was in it
**The ruling, San's, 2026-08-08, and it is the line every field in `room.json` is now cut
along:**
> `room.json` records the room's **SHAPE** — workspace and roster — and restores it.
> **AUTHORITY** — write posture, gate cadence — is never restored; it must be typed. The saved
> posture field exists only for the reattach-mismatch notice, so it records the room **as it
> stood** (live write + live asking, both sides at once so the notice can't fire spuriously).
Two defects, one on each side of that line, and they are opposite failures of the same file.
**Shape was half-saved.** `/cd` moves the room and the file follows; `/seat` moves the room and
the file never heard. So `/seat claude,agy,cursor` — evicting a Codex that was dark on quota —
died with a restart, and Codex walked back into the room and started billing the next
unaddressed turn. That is the expensive-default defect returning through a reboot, on a control
built specifically to kill it, and the room said nothing while it happened: the header drew four
seats and the user had typed three.
**Authority was half-recorded.** `savedPosture(m.st.Write, m.opts.Auto)` reads one live field
and one launch flag. Press `a` in a gated write room and the file goes on saying `write-gated`
about a room with nothing left asking. §9.17's own closing rule is the one that was broken —
*a flag with an in-room twin stops being the answer to "what is the room doing" and becomes
only the seed* — and §9.17 named this exact call site as a legitimate launch-time read. **This
section amends that.** `savedPosture` is not a launch-time decision; it is a description of a
running room, and it is the third miss of the same shape after `/write`'s confirmation card and
the request path.
#### The roster is keys, not content — which is why it may be saved at all
ADR-008's ninth amendment ratified council writing exactly one file and ruled what may be in
it: **keys, not content.** A roster passes that test rather than being excused from it. It is
at most four vendor ids out of the closed set `addressableVendors()` — the same words `--vendor`
takes on the command line and the footer prints on every frame. It says *who was in the room*
and not one syllable of what was said in it. If the file leaked, the roster discloses which of
four public CLI tools the user had on screen, which is strictly less than the workspace path
sitting beside it already discloses.
`TestTheSavedRoomHoldsKeysAndNeverContent` is the guard, and it fails closed by pinning the
exact key set — so adding `seats` had to be a deliberate act that broke a test and got read.
It now also asserts the roster's *content* is names, so a field added to `Seats` later that
carried a note or a reason reaches this file through the same tag and gets caught there.
#### Saved when it moves, not at the next dispatch
`c`'s rule, in `clearSeat`'s own words: the room file is what a reattach reads, so a change held
only in memory is undone by quitting — the user ends a thread and finds it waiting for them.
A roster is the other thing a user deliberately takes out of the room, and it earns the same
treatment for the same reason.
**The save is an observation on `roomCommand`, not a call inside `seatCommand`.** `c` could put
its `saveRoom` inside itself because there is exactly one way to clear a seat. The roster has
`/seat`, has `/unseat`, and will have whatever narrows it next — and a save per command is a
save the third one forgets. So `roomCommand` snapshots the roster, runs the command, and saves
if it moved. Any command reachable from there inherits persistence without knowing the wrapper
exists, which is what let `/unseat` be written in a parallel lane and compose with this without
either side being told about the other.
Two consequences worth stating rather than discovering:
- **Only a change writes.** Bare `/seat`, a typo, and a `/seat` refused mid-turn all report
without reseating. Rewriting the file on those would refresh `saved_at` — the age a reattach
shows — for a room that answered a question and did nothing.
- **A room that has never dispatched still writes nothing.** `saveRoom` returns at turn 0 and
`readRoom` refuses a turn-0 file, both unchanged. A `/seat` typed before the first brief rides
out on that brief's own save, which is the only save there was ever going to be.
#### `--vendor` overrides the saved roster, and then rewrites it
`--cd`'s rule and `--cd`'s reasoning: **an explicit launch control someone typed today outranks
a file from yesterday.** `seatsFor` mirrors `Run`'s workspace switch line for line, down to
sharing the same `re.Active() && !re.Offered` — a room `--fresh` declined restores neither half
of the shape.
The rewrite needs no code, and that is worth saying because it reads like a missing branch:
`stateWith` copies the answer into `State.Seats` and `saveRoom` writes `m.st.Seats`, so the
first completed turn records the room the user actually got. The same one line is what makes
`/seat` persist. Leaving it out would be worse than not overriding at all — the file would go
on describing a room that is not on screen, and the *next* launch would restore it.
**Restoring is unconditional on the roster's own content**, including the zero value: the
default room saved as the default room. A saved roster that could only ever *widen* would be a
`/seat` you could not undo by quitting.
**Back-compat is the absence of a field, and it is exact.** A `room.json` written before this
section has no `seats` key; that decodes to the zero `Seats`; the zero `Seats` is the full
detected table. So an old file opens the room it has always opened, and no version bump is owed
— `roomVersion` is bumped when a field *changes meaning*, and additive fields are handled by
the zero value, which is the rule `roomVersion`'s own comment already states. Pinned by a
hand-written v2 fixture rather than a round-trip, since a file this build saved would carry the
field and prove nothing.
**An unknown seat name is dropped, not obeyed.** The roster is the one restored field whose
value is a *name*, so it is the one a hand-edit or a downgrade can fill with a word this build
has no seat for. Obeying it would seat nobody, fall through the everything-collapsed fallback in
`VisibleColumns`, and hand the user the default room while the file claimed a narrowed one —
§4a.1's collapse in the surface this section exists to make trustworthy. Dropped rather than
refused, because a roster is shape: the sessions are still perfectly reattachable and refusing
the whole file over the seating plan would cost four conversations to fix a screen.
#### The posture field records the room, and still never restores it
Both arguments are live now — `m.st.Write` and `m.st.Asking()` — and **the writer and every
reader moved in one change**, because the field has exactly one consumer. A writer reading the
state while a reader read the flags would compare a description of the live room against a
description of the launch argv and report a change to a user who made none. That is the
spurious fire the ruling names, and it would have been *introduced by the fix* had either side
moved alone.
Recording the room accurately is the opposite of restoring it, and nothing about the restore
changed. `TestReattachRestoresNoPostureAndNoGate` extends the old
`TestReattachDoesNotRestoreWritePosture` to the gate as well: a room saved `write` reopens read,
a room saved with the gate off reopens asking, and the WRITE marker is asserted absent on the
rendered frame rather than on a field. Both halves are witnessed, so it cannot pass by the room
being read-only for some reason of its own — the same fixture reopened with `--write --auto`
gets exactly what was typed. *A posture that can arrive from a file is not one anyone typed*
survives this section unchanged; what it never said is that the room may not write down what it
did.
#### Declined
**Restoring the gate.** It is authority, it is on the far side of the ruling's line, and `a`'s
own section already argues that a safety property whose default is off is the wrong way round
however carefully the constructor sets it. A gate that can arrive from a file is that mistake
with a longer fuse.
**Recording *why* the roster is what it is** — flag, command, or file. It is the room's own
history rather than its shape, `saved_at` already dates it, and a `reason` string is the first
thing in this file that would be prose.
**Bumping `roomVersion`.** A bump costs every user their reattach, and it is reserved for a
field that changes meaning. Nothing here changes what an existing key means.
#### Amendment, 2026-08-16: the restore was correct, the fallback was silent, and neither had a seam
A live 5/5 room reopened with the workspace cell reading `~` instead of the repo it was saved
in (`STATE.md`, the 2026-08-15/16 drive). The roster came back and the workspace did not.
`STATE.md` recorded the cause as undetermined: either an earlier session saved `~` honestly, or
the restore dropped the field. This amendment answers that question with a measurement.
**The restore does not drop the field.** `TestASavedWorkspaceIsRestored` plants a `room.json`
whose saved workspace exists, opens the room, and asserts the workspace comes back. It passes
against the *untouched* decision logic. The ordinary reopen was always correct, so the file
itself held `~`.
**The mechanism is the fallback that put it there.** `Run` verified the saved directory with
`os.Stat` and replaced it with the current directory when that failed. It said nothing specific
about the replacement. The reattach notice printed one sentence for two different events: *the
room was in A; it is now in B* is what a `--cd` override prints, and a vanished workspace
printed the same words. Nothing distinguished a directory the user moved to from a directory
that no longer exists.
**The fallback then persisted itself, and that is the real cost.** The room opens in the current
directory. The next completed turn calls `saveRoom`, which writes `m.st.Workspace`. So one
launch against a missing path overwrites the only record of where the room was. A renamed repo,
an unmounted drive or a removed git worktree costs the saved workspace permanently, and the room
never names the path it lost. That is `Reattachment.Offered`'s argument applied to the
workspace: the destruction is silent and total, so the room must state it once.
**Nobody could measure any of this, which is the other half of the finding.** The decision was a
`switch` inside `Run`, and `Run` enters the alternate screen. No test could reach it. An
untestable decision is one whose failures are all reported by the operator, which is exactly how
this one arrived.
**The fix.** `openWorkspace(opts, re)` is that decision as a function. It returns the directory
AND the saved path it refused. `Run` carries the refused path on `Reattachment.WorkspaceGone`,
and `reattach` gives it its own sentence: *the room was in ~/code/x, which no longer exists — it
opened in ~/code/y instead*. The `--cd` sentence stays for the case it actually describes.
- **One stat, one answer.** `openWorkspace` is the only place that stats the path. `reattach`
reads the carried fact instead of statting again. Two reads a moment apart can disagree, and
the room would then choose its workspace on one answer and describe it with the other.
- **`WorkspaceGone` is never written to disk.** It is a fact about this launch, not about the
room. `room.json` stays the keys and nothing else, per ADR-008's ninth amendment.
- **A file sitting where the directory was is gone too.** `isDir` asks whether the path is a
directory, not whether it exists. `os.Stat` succeeds on a file, and a room pointed at a file
would dispatch four agents against it.
- **A saved workspace is resolved rather than trusted as written.** `resolveWorkspace` makes it
absolute. Every other consumer of a workspace in this package is handed an absolute path.
- **The `--cd` refusal is unchanged.** A typed path that is not a directory stays a plain error
before the alternate screen. The user named that path, so a silent substitution would act
somewhere they did not ask for.
`seatsFor` still mirrors this decision on `re.Active() && !re.Offered`. The mirror moved from a
`switch` in `Run` into `openWorkspace`, and the shared condition is unchanged.
**Declined, and named rather than left implicit.** The room still writes the fallback over the
saved workspace at the next completed turn. Preserving the old path would make `room.json`
describe a room nobody is in, which this section already refuses for the roster and refuses here
for the same reason. The notice is the answer chosen instead.
#### Amendment, 2026-08-16: the save choke point observed half the shape
This section opens by saying `/cd` moves the room and the file follows. That was true of the
FIELD and not of the WRITE. `saveRoom` reads `m.st.Workspace`, so whatever save came next
recorded the move — but `/cd` made no save of its own, and `roomCommand`'s choke point compared
only the roster (`sameSeats`). So the move reached disk at the next completed turn, or at
teardown, or never.
**Never is the case that matters, and the per-turn save already named it.** `endTurn` writes
rather than leaving it to the way out, in its own words, because the failure it exists to
survive is the room not getting a clean exit: a crash, a closed terminal, a machine that went
down. A `/cd` had exactly that hole. It survived a quit, and a room that crashed after a `/cd`
and before its next completed turn reopened in the directory the user had moved out of. The
workspace is the field beside the session ids in the same file, on the same half of the ruling's
line, with none of the protection.
**A `/cd` is a deliberate operator statement about where the room is.** That is the roster's
argument — `c`'s argument in `clearSeat`'s words, a change held only in memory is undone by
quitting — reaching the other half of SHAPE. It should write when it happens, not when something
else happens to write.
**The fix is one line at the choke point, and deliberately not a `saveRoom` inside `cdCommand`.**
`roomCommand` now snapshots the workspace beside the roster and saves if EITHER moved. The
wrapper exists precisely so a command inherits persistence without knowing it does, and putting
the call inside `cdCommand` would have been the per-command save this section already rejected —
correct for `/cd` and absent from whatever re-points the room next.
- **Each half compares in its own terms.** `sameSeats` because a slice does not compare with
`==`; `sameDir` because two spellings of one directory are one directory, case-folded on
Windows. Comparing the workspace with `!=` would write on a `/cd` that changed nothing but the
capitalisation.
- **The refusal semantics are unchanged, and the observation is what keeps them.** `resolveCD`
rejects an unknown path BEFORE the workspace is assigned, so a bad path is never a value the
file could briefly hold — the refusal is not layered on top of a write, it is upstream of one.
A `/cd` mid-turn, a `/cd` to the directory the room is already in, and bare `/cd` all return
with the workspace untouched, so none of them writes. That is the roster's rule verbatim:
rewriting the file on a command that answered a question and did nothing would refresh
`saved_at`, the age a reattach shows.
- **Turn 0 still writes nothing.** `saveRoom` returns before the first dispatch, so a `/cd`
typed before the first brief rides out on that brief's own save, and a room opened in the wrong
directory and quit still drops no file into `~/.telltale/council`. Stated as a test rather than
left to be found.
- **The write path is `saveRoom`'s, unchanged** — the same atomic temp-file-and-rename, the same
best-effort failure stated in the footer. No new writer, no second serialization of the same
file.
**Measured against the crash rather than against the field.** `TestCdIsPersistedWhenItHappens`
drives `/cd` through `roomCommand` and then reads `room.json` off disk with no teardown and no
completed turn — the simulated crash. It fails on the pre-amendment code with *nothing was
saved*, which is the defect in one line.
### 9.33 the cursor seat's per-turn cost, split at last — and the seam that was hidden from `--help`
§9.8 gave the Claude seat one live process and measured what it bought. The obvious next
question was whether the Cursor seat could have the same thing, and the standing instruction in
`STATE.md` was to read a trace before optimising anything. This is that reading.
**Version pinned first, because the last capture's lesson was that a rule is only as general as
the capture it came from (§9.6c).** Everything below is `cursor-agent` **2026.08.04-aaa8809**, the
bundle's own `--version`, on Windows 11. That is **not** the version the rest of this seat was
measured against — `vendors/cursor.go` cites 2026.07.23-e383d2b throughout — and one of the
findings is a direct consequence of the gap.
**Instrument:** the vendor's own `node.exe` against `index.js`, argv identical to the seat's read
posture, with every stdout line stamped against the moment of launch. Two trials per arm. The
`result` event carries the vendor's own `duration_ms`, which is the cross-check: it agrees with
`system/init` → `result` on every trial, so the split below is the vendor's arithmetic as much as
this instrument's.
#### What the 44 seconds actually decomposes into
`STATE.md` already established that spawn is 13 ms and that `wait` is where the time goes, and
said outright what it could not do: `wait` bundles the vendor's startup with the model's
time-to-first-token, and nothing then in the room could separate them. Stamping raw lines
separates them, because `system/init` lands *before* the model is called.
Print mode, no `--resume` — trivial prompt, `reply with exactly: OK`:
| trial | launch → `system/init` | `init` → `result` (vendor `duration_ms`) | `result` → exit | total |
|---|---|---|---|---|
| 1 | 5.666s | 5.779s | 2.299s | 13.742s |
| 2 | 5.617s | 5.361s | 1.818s | 12.792s |
Print mode, `--resume` against a real prior session (created by the trials above, so nothing of
anyone's real work is in this record):
| trial | launch → `system/init` | `init` → `result` | `result` → exit | total |
|---|---|---|---|---|
| 1 | 5.196s | 5.042s | 3.104s | 13.337s |
| 2 | 5.551s | 4.298s | 3.080s | 12.928s |
And the startup itself, taken apart with progressively less work asked of the same bundle:
| what ran | to first output |
|---|---|
| `node.exe -e "console.log('x')"` — interpreter only | 0.078s |
| `node.exe index.js --version` — interpreter + bundle load + arg parse | 1.204s |
| a turn in an **untrusted** directory (aborts at the trust check, before any model call) | 2.139s |
| a real turn, to `system/init` | ~5.6s |
**Three things follow, and only the first was already known.**
**The standing diagnosis was right, and `--resume` is not the expensive half.** "`--resume`
restores context, not process warmth" is confirmed and now has a number against it: resumed
startup (5.196s, 5.551s) is *no larger* than cold startup (5.666s, 5.617s). Restoring a
conversation is free. What costs is the fixed startup underneath it, paid identically either way.
**Process cost is ~8.1s per turn and none of it is the model.** ~5.6s before `system/init` plus
~2.5s after `result` — the process lingers after answering — against a model turn the vendor
itself clocks at 4.3–5.8s. Of the ~5.6s startup, node is 0.08s and loading the bundle is ~1.13s;
the remaining ~4.4s is the vendor resolving auth, config, trust and workspace, and it is the
largest single item in the seat's budget.
**The honest proportion, stated so the number is not oversold.** On these trivial prompts the
8.1s is ~60% of the turn, but a trivial prompt is the arm that flatters the finding most. Against
the real room traced in `STATE.md`, where `cursor` totalled 25.014s, the same fixed 8.1s is ~32%.
The *absolute* figure is what is load-bearing: it does not shrink as the question gets harder, and
it is paid again on every single turn.
#### The seam: what print mode cannot do, and what the hidden subcommand can
Persistence needs two halves. The output half the seat already has — `--output-format stream-json`
is what §9.6c parses. The input half is the one that decides it: a way to hand turn N+1 to a
process that is already running.
**Print mode cannot be that channel, and the measurement is unambiguous.** Turn one was written to
an open stdin and then the pipe was *held*. Nothing happened for sixty seconds. Only when stdin was
closed did `system/init` appear, 3.6s later, and the turn ran — one turn, on the joined contents of
stdin, then exit. **Print mode drains stdin to EOF and treats the whole of it as one prompt, so the
EOF that starts the turn is the same EOF that destroys the channel for the next one.** There is no
`--input-format` in `--help`, and none in the bundle either: enumerating every flag the bundle
defines turns up hidden development flags (`--ian-dev`, `--sb-debug`, `--tool-gallery`), which is
what makes that absence evidence rather than an unsearched corner.
**One correction to this repo's own record falls out of the same probe.** `vendors/cursor.go` said
no code path in the bundle reads the prompt from stdin, and that there is no `-` sentinel and no
`--prompt-file`. That was true when it was measured; at 2026.08.04-aaa8809 the first clause is
**false** — a prompt piped in with no positional argument produced a normal turn. Nothing in the
seat depends on it (council always passes the prompt in argv), so this changed no code; the comment
is corrected because a stale measurement left standing is how the next reader inherits a wrong
premise.
**The channel exists, and `--help` does not mention it.** The bundle registers a subcommand marked
hidden:
```
Ce.command("acp",{hidden:!0}).description("Start the Cursor Agent as an ACP (Agent Client Protocol) server")
```
This is the `--permission-prompt-tool stdio` situation from §9.8 exactly — absent from the help
text and real — so it was driven live rather than believed. **Two turns, one process, one session:**
| trial | `initialize` | `session/new` | turn 1 | turn 2 |
|---|---|---|---|---|
| 1 | 1.944s | +0.994s | 5.285s | 5.335s (this turn ran a tool call) |
| 2 | 1.701s | +1.040s | 5.365s | **1.177s** |
The shape, recorded rather than the content: JSON-RPC 2.0, newline-delimited, on stdin/stdout.
`initialize` returns `agentCapabilities` — including `loadSession: true`, the resume equivalent.
`session/new` takes a `cwd` and returns a `sessionId` plus `configOptions`, among them a `mode`
select whose values are `agent`, `plan` and `ask`. Turns are `session/prompt` requests carrying
that `sessionId`; output arrives as `session/update` notifications (`agent_message_chunk`,
`agent_thought_chunk`, `tool_call`, `tool_call_update`) and the request resolves with a
`stopReason`. The second turn correctly answered a question about the first, from the same pid, so
this is one conversation in one process and not two conversations that happened to share a parent.
**The prize, stated as measured:** a follow-up turn costs **1.18s** where a print-mode turn costs
~13s, because the ~8.1s of process cost is paid once at `initialize` and never again.
#### What this section does NOT authorise, and why it stops here
The gate this work was run against was "build persistence only if the cost is process warmth *and*
a live-verified seam exists." Both are now true, so the finding is **build**, and it is worth
building. What the measurement also established is that the build is **not** the change it was
expected to be — mirroring §9.8's shape onto this seat — because ACP is a *different protocol*, not
the same protocol with an open stdin. Three forks come out of that, each a design decision rather
than a detail, and each one is recorded here instead of guessed at:
- **`Persistent` as written cannot express ACP.** `Turn(prompt) ([]byte, error)` is stateless: it
returns the line for a turn. ACP needs `initialize`, then `session/new`, then a `sessionId`
captured out of a *response* before any turn can be encoded at all — and `runner.Session` pipes
lines and correlates nothing. Server→client requests (ACP's `session/request_permission`, the
natural home for §9.8's gate) have no channel back at all today. That is a change to shared
runner plumbing, not to one adapter.
- **Posture and cwd stop being argv-bound, which un-founds the respawn rules.** `persistent.go`
respawns a seat when the room moves or a `/flow` hop needs a different posture, and the comments
there rest on both being fixed at spawn. In ACP, `cwd` is a `session/new` parameter and `mode` is
a session `configOption` — so a `/cd` could open a new *session* in the same live process, and a
posture change might not need a respawn either. Whether it *should* is a product question about
what the badge is allowed to promise, not a mechanical one.
- **Every measured claim on this seat was measured against print mode.** The §9.6c dedup rule, the
`tool_call` oneof wire shape, the `--mode plan` badge and the Windows sandbox finding are all
facts about a surface ACP does not use. Worth noting precisely because it is *not* yet a finding:
across these two ACP turns, `agent_message_chunk` arrived once per turn with **no whole-message
repeat** — which would mean the dedup rule is unnecessary here. That is a two-turn capture with
one tool call in it, and §9.6c is the standing warning against generalising exactly that. It is
a hypothesis for whoever builds this, not a rule.
So the seat keeps its print-mode invocation for now, and the next lane starts with a number, a
verified seam, and three named decisions instead of a guess.
#### 2026-08-15: the same rig, pointed at the codex seat, and what a warm thread saves
The rig above measured one vendor. This block runs it against a second one, and answers a
question `STATE.md`'s 2026-08-08 trace could not. That trace shows the codex seat paying
`wait=3.688s` before its first byte, while a cold binary start measures 190ms. Nothing said where
the other ~3.5s went. **This is measurement only. It authorises no seat change.**
**Version pinned first, and the subcommands were driven before they were believed.** Everything
below is `codex-cli 0.147.0` on Windows 11 (`codex --version`). That is a NEWER build than the one
`vendors/codex.go` cites. `codex app-server` and `codex app-server generate-json-schema` both
exist on this build and both ran: the schema command wrote 46 files to a directory, and the server
answered a live `initialize`. A subcommand named in `--help` is not evidence of a subcommand that
runs, which is this repo's twice-earned lesson, so both were executed rather than read.
**Instrument:** the installed `codex` binary, argv identical to the seat's first turn in
`vendors/codex.go` (`-s danger-full-access --skip-git-repo-check --cd -`), with every stdout
line stamped against the moment of launch. The prompt is **brief-shaped**: it opens with
`brief.go`'s own `--- operating context ...---` fence and carries the request under it, because a
greeting-shaped probe measures a transport the product never uses. One trial per arm, which is
half of what §9.33 spent. Treat every figure below as one observation.
#### The three-way capture, one identical turn
`codex exec` (human), one trial. The seat does not use this renderer; it is here because it is the
only arm that shows what the `--json` arm drops.
| stamped line | at |
|---|---|
| spawn returned | 0.030s |
| banner (`OpenAI Codex v0.147.0`) | 0.928s |
| `hook: SessionStart` | 3.709s |
| `hook: SessionStart Completed` | 4.201s |
| the model's answer (`OK`) | 6.800s |
| `tokens used` / `13,543` | 9.036s |
| process exit | 15.519s |
`codex exec --json`, one trial. This is the seat's own invocation.
| stamped line | at |
|---|---|
| spawn returned | 0.014s |
| `{"type":"thread.started",...}` | 1.153s |
| `{"type":"turn.started"}` | 1.554s |
| `{"type":"item.completed",...,"text":"OK"}` | 5.250s |
| `{"type":"turn.completed","usage":{...}}` | 5.327s |
| process exit | 13.266s |
`codex app-server`, one process, one thread, two turns. The fixed half is paid once:
| stamped line | at |
|---|---|
| spawn returned | 0.016s |
| `initialize` response | 0.246s |
| `thread/start` response, and the `thread/started` notification | 0.572s / 0.573s |
Then the two turns, both on that one open thread:
| turn | `turn/start` sent | `turn/started` | first `item/agentMessage/delta` | `turn/completed` | wait | stream | total |
|---|---|---|---|---|---|---|---|
| 1 | 0.576s | 0.836s | 5.085s | 5.259s | **4.509s** | 0.174s | 4.683s |
| 2 | 5.263s | 5.303s | 6.467s | 6.705s | **1.204s** | 0.238s | **1.442s** |
**A warm turn costs 1.44s, and 1.20s of that is the model.** Turn 2 asked what turn 1 had answered
and got it right from the same pid, so this is one conversation in one process. Against the same
prompt through `codex exec --json` the comparison is 1.442s against 5.327s to the last line, or
against 13.266s to exit.
**Four separate items make up the difference, and only one of them is process start.**
1. **Process and thread start is 0.573s, not 3.5s.** `initialize` answers in 246ms and
`thread/start` in a further 326ms. Spawn itself is 16ms, which agrees with the 190ms class of
figure and confirms again that spawning was never the cost.
2. **A `sessionStart` hook runs before the model does, and it is the operator's own.**
`hook/started` at 3.151s and `hook/completed` at 3.840s, and the notification names its source:
`"sourcePath":"C:\\Users\\sanle\\.codex\\hooks.json"`, `"source":"user"`, `"durationMs":838`.
The human arm shows the same hook as `hook: SessionStart` at 3.709s. **This item is
machine-specific.** A box with no `hooks.json` would not pay it, so it must never be quoted as
a property of the vendor.
3. **Five MCP servers start on the same path.** `mcpServer/startupStatus/updated` fires for
`node_repl`, `context7`, `github`, `kb-agent` and `codex_apps`, and two of them go
`starting` to `cancelled` to `ready` across the turn. This is also operator config, and the
same caution applies.
4. **The process lingers after it answers.** `exec --json` printed its last line at 5.327s and
exited at 13.266s, which is **7.94s** of linger. The human arm shows 6.48s of the same. §9.36's
"kill, never wait" rule was written for a different vendor and a ~2.5s linger. **This vendor's
linger is larger, and nothing here checked whether council waits on it.**
**One comparison this block does NOT make.** The `exec --json` arm reached its first line at
1.153s, well below `STATE.md`'s `wait=3.688s`. That trace ran a real brief through a real room
with four seats starting at once, and this trial ran a trivial prompt alone. The 3.688s is not
reproduced here and must not be treated as refuted.
#### The hook question, answered by the captures
**`codex exec --json` is the only one of the three surfaces that hides hook activity.** The human
renderer prints `hook: SessionStart` and `hook: SessionStart Completed`. The protocol emits
`hook/started` and `hook/completed`, each carrying the hook's id, event name, source path, source
and `durationMs`. The `--json` stream emitted **four lines in total** for the whole turn
(`thread.started`, `turn.started`, `item.completed`, `turn.completed`) and not one of them mentions
a hook, an MCP server, or a rate limit.
The protocol also carries two things the seat currently reads off disk instead:
```
{"method":"thread/tokenUsage/updated","params":{...,"tokenUsage":{"total":{"totalTokens":21130,...},"modelContextWindow":258400}}}
{"method":"account/rateLimits/updated","params":{"rateLimits":{"limitId":"codex","primary":{"usedPercent":2,"windowDurationMins":10080,"resetsAt":1787369304},"secondary":null,"planType":"plus",...}}}
```
Those are the same fields §3.4 verified in the rollout files, arriving live on a socket. **Nothing
is built on that here.**
#### Seat-move viability, recorded and not acted on
1. **Thread continuity exists.** `thread/started` arrived live at 0.573s. `thread/resume` is a
real request and its schema documents three routes (`thread_id`, `history`, `path`), plus
`thread/fork`. The `thread/start` result carries `thread.id`, an identical `sessionId`, and the
rollout `path` under `~/.codex/sessions/...`, which is the same id the adapter already reads.
2. **The sandbox channel exists on this path, and it is wider than `-s`.** `thread/start` takes a
`sandbox` parameter, and the live response echoed `"sandbox":{"type":"dangerFullAccess"}`.
`turn/start` takes a per-turn `sandboxPolicy`. That is strictly more than `codex exec` offers,
where `-s` is first-turn-only and `codex exec resume` rejects it outright. Only
`danger-full-access` was driven here.
3. **The Windows `danger-full-access` finding does NOT carry over, and needs its own re-check.**
The evidence is direct rather than inferential: this protocol has a Windows sandbox surface
that `codex exec` has no equivalent for. `windowsSandbox/setupStart` and
`windowsSandbox/readiness` are client requests, and `windowsSandbox/setupCompleted` and
`windows/worldWritableWarning` are server notifications. `vendors/codex.go`'s finding rests on
`-s read-only` failing every process spawn, and `read-only` was never sent on this path. Until
somebody sends it, the seat's badge rule stands unchanged.
**Spend:** four billed turns. Two arms of one turn each, plus two turns on the app-server thread,
because a single app-server turn would have reported turn 1's 4.509s as if it were the warm number
and oversold nothing or undersold everything depending on which row was quoted.
#### 2026-08-16: the linger gets an owner — what rides the tail, and what the column says now
The block above ended with item 4 and a named gap: *"this vendor's linger is larger, and nothing
here checked whether council waits on it."* It did wait on it. This is the check, the answer, and
the fix.
**The room's side of it, read from source first.** `vendors/codex.go` parsed `turn.completed` into
a bare `KindMeta`, which `applyEvents` uses to adopt a session id and record a cost — codex sends
neither on that line. Nothing else consumed it. A spawn-per-turn seat retired only on `KindDone`,
the process exit. So the answer-complete marker was arriving, being parsed, and changing nothing:
the column held `streaming` from the last token until the process died, and the elapsed it kept
was the process's lifetime rather than the answer's.
**Re-measured, because §9.33 spent one trial on this and the number is the whole argument.** Same
build as §9.33 (`codex-cli 0.147.0`, Windows 11), so nothing here is confounded by a version
change. Instrument: the installed binary, `codex exec --json` with the seat's own argv shape, a
brief-shaped prompt behind `brief.go`'s fence, in a throwaway directory outside any repo.
**`-s read-only`, not the seat's Windows `danger-full-access`** — a probe does not need write
access, and the prompt tells the model to use no tools, so §9.33's Windows spawn refusal is never
reached. It ran clean twice, which is a small measured aside worth keeping: `-s read-only` on
Windows breaks codex's *child* spawns, not codex itself.
| | trial 1 | trial 2 |
|---|---|---|
| `thread.started` | 1.914s | 0.537s |
| `item.completed` (the answer) | 6.417s | 4.405s |
| **`turn.completed`** | **6.619s** | **4.499s** |
| last rollout write | 6.670s | 4.723s |
| stdout/stderr close | 10.869s | 8.554s |
| process exit | 10.870s | 8.555s |
| **linger** | **4.251s** | **4.056s** |
**Nothing rides the tail, and that is the finding this change rests on.** The concern was that
receipts or vendor-side bookkeeping might be paid after the answer, which would make an early
settle a lie about a turn still in progress. It is not: on both trials `turn.completed` is the
LAST line on stdout, stderr stays empty throughout, and the vendor's own rollout file under
`~/.codex/sessions/` takes its final write 51ms and 224ms after that line — roughly four seconds
*before* the exit. The rollout was polled at 250ms rather than diffed before-and-after, so this is
an observation of when writes stopped, not an inference from a file that changed.
The linger itself is 4.06s and 4.25s here against §9.33's 7.94s on the same build. **The size is
not stable and must not be quoted as a constant**; the shape is, and the shape is what the render
has to survive.
**What changed, and the line it does not cross.** `turn.completed` now carries `EndsTurn`, and
`applyEvents` grew a third branch for a spawn-per-turn seat that names its own end of turn — a
case that did not exist before, because the batch CLIs all ended a turn by dying. That branch
**settles the column without retiring it**: phase to `done`, elapsed stamped at the answer, body
completed, and the seat stays in `m.turn.live`.
The split is the whole design, and the reason is mechanical rather than tasteful.
`turnColumnFinished` cancels the turn's context; `runner.Start` kills the child on that context.
Retiring the column at the marker would therefore **kill the process four seconds early**, and
that is refused even though the tail measured empty — both probe turns used no tools, so a turn
that ran commands is unmeasured, and probing one needs `danger-full-access`, which a probe does
not get. Shortening a vendor's life on evidence that does not cover the case is the
inference-dressed-as-measurement move §4a.1 exists to refuse. The exit still arrives, still runs
`KindDone`, and still retires the column. It just no longer decides what the column *says*.
Two consequences fell out, and both are corrections rather than costs:
- **The turn clock stops at the answer.** `clock.observe` already ended a turn on `EndsTurn`, and
its comment said a spawn-per-turn child never sets it. That comment is now false and is fixed.
The old reading billed the linger to the vendor's turn time, so the seat was timed on how long
its process took to get out of the way.
- **`Busy()` and the footer came apart, and had to.** `Busy()` is derived from column phases, so
it goes false the moment the seat settles — correctly: nothing is working. But the turn is still
live, and `q` is refused while a turn is live, so the footer's `Busy()` test would have
advertised a key that answers with a notice (§7.8). The footer now asks `InFlight()`
(`Busy() || Settling()`) instead. `Busy()`'s own doc claimed it drove the meaning of `ctrl+c`;
it never did — that key reads `Model.turn` — and the stale claim is corrected rather than
preserved.
**The linger is rendered, not hidden.** Between the two moments the room would otherwise go
completely quiet — no spinner, every column reading `done` — with the composer still locked. That
is a room that looks wedged for a different reason than before, which is not a fix. So a settled
seat renders `done 4s exiting`: a WORD, after the clock so it cannot be read as part of the
figure, and a word rather than a glyph or a colour because what it prevents is a reader concluding
the room is stuck, and that reader may be on `--ascii` or `NO_COLOR` (§7.1 rule 2).
**Two defects the review caught in this change, both now guarded.** A killed process
drains its buffered stdout, so a `turn.completed` can arrive after the column it belongs
to is already terminal — `giveUpSeat` kills an arena racer and retires its column as
cancelled, and the queued line lands behind it. The phase write was guarded from the
start; the BODY write was not, so a cancelled seat's note-bearing body was replaced with
`[Turn completed with 0 text chunks streamed]` — a cancelled column asserting that its
turn completed. Every write on the branch now sits behind one guard, and the guard is
wider than a phase test: it also checks `m.turn.live`, which covers the same line
arriving after the turn boundary entirely, where it could otherwise settle a *fresh*
turn's column on the strength of the previous turn's answer. Second, the vendor-reported
failure path restamped `Elapsed` unconditionally, so a seat that answered and then exited
badly recorded the process's whole lifetime — the exact figure this block exists to stop
billing. It now fills only a zero, which is the rule `finishColumn` already followed.
**What is NOT claimed.** The linger's cause is still unknown — this block measures when it starts
and ends and what does not happen during it, and nothing about why the vendor holds the process
open. `codex app-server` (§9.33's third arm) has no linger at all because the process is meant to
stay up, so the seat-move case gains an argument here; it is recorded, not acted on, exactly as
§9.33 left it.
**Spend:** two billed turns, both trivial prompts on `read-only`.
#### 2026-08-16: the gap that block named, on the seat that fails inside its stream
The block above closed the answer case and left one gap open, written into `InFlight`'s own doc
comment: **a seat that FAILS in its stream was terminal inside a live turn.** This closes it. No
vendor was run for this change. The whole finding is a source read, and the section says so
rather than borrowing the authority of the measurement above it.
**What the source says.** `vendors/agy.go` parses a `result` line carrying `status: "ERROR"` into
a `KindError`. It fills `Note` from the vendor's own `error` field. It sets no exit code and no
error, because nothing has failed at the process level and nothing has exited yet. It does not set
`EndsTurn`; no agy event does. `applyEvents` therefore reached its `KindError` tail, wrote
`PhaseFailed`, and retired the column only on `ev.ExitCode != 0 || ev.Err != nil`. This event
carries neither.
**One evidence boundary, because the phrase is easy to over-read.** "Exit code 0" here is the
EVENT's field, read off the adapter. It is not a measurement of what the `agy` process exits with
after a failed turn, and no such capture exists — `vendors/testdata/wire/README.md` records that
the one probe that could have produced an agy error frame does not produce one, because that CLI
answers a lost thread with success. Nothing here depends on the process's exit status. `KindDone`
and a failing `KindError` both retire the column, so either exit ends the turn.
**The room that produced.** The column read `failed`. The turn stayed live, because nothing
retired it. So the seat was neither `Busy()` nor `Settling()`, the spinner stopped, every column
on screen read terminal, and `q` was still refused with *"a turn is in flight"* — which was true.
The footer offers `q` on `InFlight()`, so the room named a key that answers with a notice. That is
§7.8's surprise, and it is the same defect the block above fixed for the seat that answers early,
reached through the other terminal phase.
**The fix is that block's split, applied to `failed`.** The column settles instead of retiring:
phase to `failed`, the vendor's sentence kept, `Settling` set, and the vendor left in
`m.turn.live`. The exit still arrives, still retires the column, and still clears the word. The
seat renders `failed 5s exiting`, which is the same honest sentence `done 5s exiting` makes about
a different outcome.
**Retiring the column was the alternative, and it is refused for §9.33's reason rather than for a
new one.** `turnColumnFinished` cancels the turn's context, and `runner.Start` kills the child on
that context. Retiring here would kill a process that is still winding down. Codex's linger was
measured before that argument was accepted; this vendor's linger was not measured at all when this
block was written, which made the case stronger rather than weaker. A room may not shorten a
process's life on a number nobody has.
> **Amended 2026-08-16 (§9.43).** That number now exists, for the SUCCEEDING turn only: agy's tail
> is 0.049s, 0.135s and 0.314s on `agy 1.1.13`, against codex's 4.06s and 4.25s. The settle above
> stands unchanged. The measurement makes the ruling cheaper rather than wrong — there is almost
> nothing left to cut short — and it does **not** cover this branch, because all three trials
> succeeded. The failing turn's tail is still a number nobody has.
**Moving the fact onto `State` was the second alternative, and it is refused too.** `State` cannot
see `Model.turn`, so "a turn is live" would have to be copied onto it and reset on every path that
ends a turn. `Settling`'s own doc comment already rules on that shape: a second home for a fact is
a second thing to forget to reset. `Settling` IS the state-side fact this needed, and the failure
path was simply not setting it.
**The guard is the same one review found for the answer case.** A killed process drains its
buffered stdout, so this line can land on a column that is already terminal, or after the turn
boundary entirely. The settle is therefore behind a phase test AND `m.turn.live`, read before the
phase is written. Without it a late failure line would hold `InFlight` true with nothing running —
the room wedged the other way, where the footer never offers `q` again. `failedturn_test.go` pins
all three states: the settle, the retirement on the exit, and the late line that must change
nothing.
**Scope, stated because the branch is vendor-neutral and the case is not.** Any spawn-per-turn
seat whose adapter reports a turn failure in-stream reaches this branch. agy is the only seat that
does so today: codex and grok have no structured error frame, so their failure IS the exit and the
`ExitCode != 0` leg already retires them; Claude runs persistent; the Cursor seat is ACP. The fix
is written on the shape rather than on the vendor id, and a seat that adopts the same shape later
is covered without an amendment.
*(Amended 2026-09-01: codex adopted the shape. codex-cli 0.151.0 puts a `turn.failed` frame on
stdout, and the adapter now parses it into exactly this event. The amendment is not free: a
spawn-per-turn failure now produces TWO events, the vendor's sentence and then the exit, and the
exit must not overwrite the sentence. [§9.58](#s9-58) records that rule.)*
### 9.34 the rebuttal stopped naming its authors
A `ctrl+r` turn used to quote each seat's answer under its vendor's name: *"quoted reply from
Codex"*. The fence's security framing was right and is untouched (quote.go); what was wrong is
subtler — **the receiving model was told who wrote what**, and models weigh an argument
differently when it arrives under a name they recognise. That is the self-preference /
identity-bias class that peer-review setups blind for, and the reason llm-council anonymises its
review stage before models rank each other. A rebuttal exists to test the argument, so the
argument is now what crosses: *"quoted reply from participant A"*.
The mechanics, because each carries a decision:
- **Labels are positional per receiver** — seat order with self skipped — and a seat with
nothing to quote this turn keeps its letter reserved, so a quiet seat does not shuffle every
neighbour's identity. The letters exist so a multi-turn argument stays attached to a
consistent speaker; letters that agreed BETWEEN receivers would need a shared assignment
written somewhere the models could correlate, and nothing downstream may join on them anyway.
- **The blinding is label-deep, and says so.** A reply whose content self-identifies ("as
Claude Code, I…") has identified itself, and editing another participant's words to hide it
would be the censorship the fence refuses — the room shows what was said. Best-effort
blinding, stated as such, over silent redaction.
- **The user is not blinded.** Columns stay labelled by vendor; the blind applies to what the
models read, never to what the person sees. Nothing in the room's rendering changed at all.
- **One test inverted, on purpose.** `TestQuotedMaterialIsFencedAsUntrusted` used to fail with
"quoted material is not attributed to its author"; attribution to the model is now the
defect. The name-absence checks in the nothing-quotable test were re-grounded on the fence
itself at the same time, because a vendor-name check passes vacuously against a prompt that
never contains vendor names — a guard that cannot fail is not a guard (the same one-level-up
rule the architecture repo's test policy states).
What this deliberately does not add: a ranking stage, a chairman, or any synthesis hop.
llm-council's stage 3 collapses the answers into one; §9.2's position is that independent
answers ARE the product, and a synthesis is available today as an explicit `/flow` hop the user
types. Blinding sharpens the comparison; it does not delegate the verdict.
### 9.35 a running chain can be told to stop after this hop
`/flow` shipped with one control over a chain already moving: ctrl+c, the turn-cancel key. That
is a hard abort — it interrupts the hop mid-sentence — and it was also a lie in three parts,
which is where this section starts, because the honest baseline had to be measured before a
gentler control could be designed against it (the flow-autoadvance plan named this gap and
nothing tracked it).
**What cancelling mid-chain actually did, measured 2026-08-08.** Ctrl+c during hop 1 of 3
killed the hop and the chain never advanced — that much was right, and stays. Everything around
it was wrong: the current step stayed `running` forever, the header went on claiming `hop 1/3`
over a room doing nothing, and the next brief the user typed was **eaten** — dispatch saw a
live chain, tried to continue it, failed with `flow start error: cannot start step in state
running`, and threw the user's words away. The second enter worked. And the corpse was not the
cancel path's alone: **a chain that COMPLETED left the same corpse**, and the first brief after
a successful flow died as `cannot start step in state returned`. The happy path was charging
the same tax.
**The fix under everything else: a chain ends whole, or it has not ended.** `endFlowChain`
retires the chain, its draft, its carry, its pending flags and its header marker in one move,
and every ending — finished, failed, cancelled, stopped, refused at the write gate, start
error — goes through it. Half of these paths used to clear only the marker and half only the
chain; each half-state was a room asserting something false about the other half. The
teardown in `turnColumnFinished` is the backstop for the endings `finishFlowHop` never sees
(a cancelled hop, a vendor failure, a hop that returned nothing), because teardown is the one
place every turn's death already passes through — and it says which death it was: `flow
cancelled at hop 1/3 — 2 later hops not run` against `flow stopped at hop 1/2 — 1 later hop
not run`. Cancelled and failed are different facts, and a stopped chain must never render as a
finished one (§4a.1).
**The control: `s`, stop after the hop that is running now.** The middle ground ctrl+c cannot
offer: the current hop finishes on its own terms — artifact saved, receipt verified, Returned
or Published exactly as if nothing had been pressed — and the chain ends there instead of
handing off. Pressed while the hop streams, because that is when the decision forms: you are
reading hop 2's output when you learn hops 3 and 4 are no longer worth their quota. Reading
the columns is part of deciding — §9.16's own argument for its gate — so the key lives in view
mode where the columns stay scrollable, and interrupts nothing.
The grammar decisions, each against a precedent:
- **A key, not a room command.** `c`'s reason (§9.17): a key takes no vocabulary from the
composer, and `/stop` is a word people address vendors with. Not `c`'s confirmation though —
`c` spends a `y` because a dropped thread is irreversible, and this destroys nothing: the
hop completes, every artifact lands, and the cost of a stray press is one keystroke to undo.
So `s` is a toggle, `a`'s shape, and pressing it again re-arms the handoff.
- **The armed state lives on the chrome, not in the notice.** The WRITE badge's argument: a
notice scrolls away and the promise persists. The hop cell reads `hop 2/4 @codex (stops
here)` — words, so it survives `--ascii`, and on the hop cell because it is a fact about the
hop: this one runs, its successors do not. The busy mode line offers `s stop after hop`
while a chain is live and flips to `s continue chain` when armed, the label-renames-itself
rule `a` set. `TestFlowStopIsAToggleAndTheArmedStateRenders` pins all four states.
- **The last hop refuses to arm.** The chain ends there whether or not `s` is pressed; a key
that "worked" would claim credit for an outcome it did not cause. The refusal says so —
`hop 2/2 is the last — the chain ends here anyway` — and a press with no chain running is
answered rather than swallowed, §9.12's attribution rule for a key that did nothing.
- **Not in helpKeys, deliberately.** The key exists only while a chain runs, and the mode line
names it on every frame of exactly those moments — the contextual-control surface `y`/`n`
and the card's `a` already use. The help panel's 17-row budget has no row for a key that is
dead in every room the panel is usually read in; the §9.17 sweep's rule was about controls
reachable from inside the room, and this one is announced there.
**What was declined.** Pausing — a stopped chain that could resume where it left off — is a
real feature and not this one: the carry artifact is consumed at dispatch, the draft would need
re-parsing against a chain whose position moved, and stop-then-retype is the honest v1. A stop
that also killed the current hop was declined because ctrl+c already is that, and two keys with
one meaning apart is how a keymap grows synonyms. And ctrl+c stays exactly as hard as it was:
the interrupt semantics did not move, only the lying state it left behind.
The general lesson, in this file's own terms: §9.16 built the chain's authority grammar and
§9.17 built the room's mid-session controls, and the seam between them — a chain that is
neither obeying nor gone — belonged to nobody, so nothing asserted on it. The corpse survived
every ending, including the successful one, because the tests all stopped at "did not
advance" and none typed the next brief. The regression tests here end by dispatching one.
### 9.36 the cursor seat re-founded on ACP: what the wider capture said, and what it cost
§9.33 ended with a build verdict, a verified seam and three named decisions. This is the build.
The ruling that shaped it was **wholesale**: ACP replaces the spawn-per-turn path rather than
sitting beside it, there is no fallback, and git history is the record of what went. A fallback
would be a second protocol to keep honest, and the numbers below are the reason nobody would
want to fall back to it.
**Version pinned, and it is the same one.** `cursor-agent` **2026.08.04-aaa8809**, the version
§9.33 measured, on Windows 11 — so nothing here is confounded by a bundle change. Instrument: a
throwaway JSON-RPC client driving `node.exe index.js acp`, every line stamped against launch,
across **thirteen arms**. Then the finished seat re-verified through the room's own code.
**One environment note, because it cost an arm and will cost the next reader's.** The first
capture had every tool call blocked by this machine's own `PreToolUse` credential guard, whose
wrapper fails closed when cursor-agent is launched from a Git Bash parent on Windows (a known
upstream wrapper bug, agent-ops ADR-012). Nothing was wrong with ACP. **Drive cursor-agent from
a PowerShell or cmd parent on Windows**, or every tool in the capture will read as failing.
#### Phase 1: what two trivial turns could not have told us
§9.6c's lesson is that a rule is only as general as the capture it came from, and §9.33 flagged
its own two-turn no-repeat finding as a hypothesis for exactly that reason. So the capture was
widened first, and the widening changed three conclusions.
| what was asked | trials | what came back |
|---|---|---|
| a turn that runs a TOOL | 3 | `tool_call` then `tool_call_update`(in_progress) then `tool_call_update`(completed). `title` always populated; `rawInput` **empty** for Read/Find/grep and populated for shell |
| several model calls in one turn | 2 | four tool calls and three message segments in one turn, interleaved, with no envelope around a "call" |
| a long streamed reply | 1 (300 words) | 24 `agent_message_chunk` in 2.6 s — ~95 chars each, ~9 a second |
| an interrupted turn | 1 | `session/cancel` (a notification) and the open `session/prompt` resolves `{"stopReason":"cancelled"}` **23 ms** later; the process took a further turn 1.1 s after that |
| resume in a NEW process | 2 | `session/load` works, and **replays the whole prior conversation** onto the update stream before it answers |
| a dead thread | 2 | `-32602 … Session "…" not found` in **0.45 s**, and the process survives — a fresh `session/new` answered 0.45 s later |
| cwd binding | 1 | **per SESSION, not per process.** One server ran a session in `ws1` reading ws1's file and another in `ws3` reading ws3's |
| workspace trust | 2 dirs | **does not apply.** Print mode refused the same directory with "⚠ Workspace Trust Required"; the ACP server wrote a file into it |
| a permission prompt | 2 | `session/request_permission` **blocks**; `allow-once` ran the command, `reject-once` did not and nothing was created |
| an edit | 2 | ran ungated — **no permission request at all**, in a never-trusted directory, under the user's own `approvalMode: allowlist` |
| plan mode | 1 | `session/set_mode {"modeId":"plan"}` accepted; asked to create a file the seat declined and **no file landed** |
| the dedup hypothesis | every arm | **no whole-message repeat anywhere, and no `model_call_id` field in ACP traffic at all** |
Timings across the twelve arms that ran a handshake: `initialize` **1.43–4.30 s**, `session/new`
a further **0.85–2.55 s**, `session/load` a further **0.89–1.37 s** — so resume is once again no
more expensive than a fresh conversation, which is the same shape §9.33 measured for `--resume`.
Warm turns: **1.12 s, 1.79 s, 1.82 s**, against §9.33's print-mode ~13 s.
**Three of these overturn something.**
**The dedup rule is not carried over, and it is now a measurement rather than a hypothesis.**
§9.6c's rule exists because print mode sent a model call's deltas and then that call's complete
message, so appending both rendered the passage twice. Across a turn with four tool calls and
several model segments, ACP repeated nothing and carries no `model_call_id` at all. §9.33 was
right to refuse to generalise from two turns; the wider capture is what earns the conclusion.
**The safety net that rule leaned on is gone with it.** §9.6c named the fallback explicitly —
"the failure mode is a column that fills at the end, never one that is wrong" — because print
mode's `result` carried the whole reply. An ACP turn resolves with `{"stopReason":…}` and
nothing else: no reply, and no token usage either. So a broken chunk parser here gives an
**empty** column, not a late one. There is no mitigation that would not be invented, so it is
stated instead — in `cursor.go`, in `dispatch.go` where the old special case was, and in
`STATE.md`.
**Workspace trust does not apply on this path.** This is the one finding that makes a claim
*worse*, and it is on the badge for that reason. The tightest form of it: the directory print
mode had just refused was written to over ACP, with no prompt, minutes later.
#### Phase 2: what was built
**runner grew a second protocol shape, and the stream-json path did not move.** `ParseFunc` sees
a line and has nowhere to reply to, which is enough for a monologue and cannot express ACP: a
turn cannot be *encoded* until `session/new` answers, the child asks questions that block it,
and ids come in two independent namespaces. So `runner.Protocol` is a stateful per-process
driver that owns both directions and returns lines rather than writing them — which keeps it
replay-testable exactly as a `ParseFunc` is. `StartSession` and `StartRPCSession` are two
wrappers over one body; the Claude adapter was not touched and its tests did not change.
`Session` grew `SendTurn` and `SendAside` in place of a bare `Send`, and the split is the turn
clock: an ACP protocol may **take** a turn it cannot yet encode, and the person who pressed
enter is waiting from that moment whether or not a byte has moved. An answer to a question the
vendor asked mid-turn goes the other way — it belongs to the turn it is holding up, so it must
not start a new one.
**The seat.** `vendors.Conversational` is a sibling of `Persistent`, not a subtype: `Open`
returns a spec plus the protocol. The invocation is now the single word `acp`. Posture arrives
as `session/set_mode`, the workspace as `session/new`'s `cwd`, and the brief as a JSON string —
so no prompt text can reach argv by any path, which retires the shell-shim question this seat
used to have to reason about.
**Re-measured, not inherited.**
| claim | verdict |
|---|---|
| granularity `tokens` | **re-earned, with a caveat.** ACP chunks are ~95 chars at ~9/s — coarser than print mode's real tokens ("P", "ONG") and *finer in time* than the Claude seat that already carries this word (§9.7: ~80 chars, ~3/s, flagged there as an overstatement). The word stays with its existing looseness and no new looseness; fixing it is one change to both seats at once, which is why §9.7 left it as a separate change to a separate surface |
| `ro:requested` | **level unchanged, reasons replaced.** Plan mode did better than print mode's ever did, and it is one trial of a mode the model obeys |
| `--sandbox enabled` | **gone.** ACP takes no sandbox parameter on any OS, so the badge is no longer split by platform. On Windows nothing was lost — the flag was measured killing the turn. On macOS and Linux what was lost is a *request* whose enforcement was never observed |
| `gated` | **withheld, deliberately.** `canGate` used to read "is this a live process", off the registry. That was right only while those two questions had one answer. This seat can ask *and does not ask about edits*, and `gated` promises that nothing which changes anything runs without a keystroke — so it keeps `WRITES`, and its detail says what the cards cover and what they do not |
| the §9.6c dedup rule | **retired**, above |
| the `result` fallback | **retired**, above |
**The gate, such as it is.** Council answers every `session/request_permission` — an unanswered
one blocks the vendor forever, which is a column that never finishes. In a write posture the
request becomes the room's ordinary approval card; in a read posture the adapter refuses it
itself and records the attempt in the trace, because a read-only seat asking to change something
is not a question for the user, it is already answered. `allow-always` is never selected in any
posture: it writes a permanent rule into the user's own `~/.cursor/cli-config.json`, which is
the line this adapter already declines to cross by never passing `--trust`.
Two shapes fall out of the capture and both are in the code beside the lines that produced them.
A **rejected** call arrives as `completed` with no output at all — indistinguishable on the wire
from a completion that said nothing — which is §9.8's `ActDenied` argument in a sharper form: the
room records the refusal from its own keystroke and `recordAct` refuses to let the echo overwrite
it. And ACP's rejection carries **no message field**, so unlike the Claude seat this one cannot
ask the model not to retry; it was measured saying "DONE" afterwards as though nothing had
happened.
**`session/load` replays history, and dropping it is load-bearing.** A loaded session streams the
entire prior conversation back — old prompts, old tool calls with their real output, old replies
— *before* it answers. A parser that appended it would refill a reattached column with the whole
previous room. The gate is the pending response rather than the `replay-` prefix those ids happen
to carry: a prefix is a spelling, the pending request is the protocol.
**Two hazards this protocol has that a one-way stream does not, both found in review and both
ending in a room nobody can quit.** They are recorded because neither is visible from the wire
format alone.
- **A turn's end is a RESPONSE, so a turn that was never sent can never end.** Anything that
holds a brief — the handshake, the `session/set_mode` round trip — is a window in which there
is no outstanding `session/prompt` for the vendor to answer or for a cancel to abandon. So the
protocol refuses an interrupt in that window rather than reporting a quiet success, which is
what makes the room fall through to killing the seat; the alternative was a turn that never
ends, a room that then refuses every further brief, and a `q` that will not quit.
- **A failed handshake is TERMINAL, because the server does not exit on one.** An ACP server that
refuses `initialize` answers and stays up — and a live process is exactly what §9.8's stale-exit
guard correctly reads as a healthy seat. Without a terminal state the room would keep handing
that process briefs forever. The protocol refuses instead, the seat is killed and forgotten, and
one retry inside the same dispatch gives the user a working column rather than an error naming a
handshake they cannot see. The likeliest trigger is an unauthenticated CLI: somebody's first run.
The same class of care applies to the `set_mode` window in the other direction: a turn *taken*
there must wait too, or it would go out under the server's default `agent` mode while the badge
said `ro:requested` — invisibly, because a reply from the wrong mode looks exactly like a reply
from the right one.
**A refused thread now costs two round trips instead of a process.** The one-attempt rule the
ninth amendment established is unchanged and is simply cheaper here: the id is spent, the same
process opens a new conversation 0.45 s later, and the brief still runs. Reattachment therefore
never fakes a restored thread — if the load is refused the seat honestly starts fresh and
`settleRestoredThread` says so, exactly as it does for the other three seats.
#### The forks, and the one that was decided rather than measured
§9.33 named three. The first (Persistent cannot express ACP) and the third (every claim was
measured against the wrong surface) are settled above. The second is a **choice**, and it is
called out here because a reader would otherwise find a measurement in the code and wonder why
it was not acted on:
**`cwd` and posture are no longer argv-bound, and the seat is respawned anyway.** Measured: one
process really did run two sessions in two directories. So a `/cd` *could* cost a new session
(~1 s) instead of a new process (~3 s). It costs a process — because what a move actually costs
the user is a new conversation either way; because one rule across four seats is worth more than
three seconds; and because re-opening a session inside a live process has failure modes (a
half-moved session, a queued turn addressed to the old one) that nothing has measured. The
argument is on `seatProc`, the behaviour is pinned by `TestAMovedRoomReplacesTheCursorSeatToo`,
and it is revisitable with a measurement rather than with a preference.
The **stale-exit guard** (eleventh amendment) applies to this seat unchanged and is re-asserted
for it: a terminal event names a vendor, not a process, and acting on a predecessor's exit would
fail the live turn and leave a real process running and invisible.
#### Verification
Fixture replay in the #62 style over synthesized shapes
(`vendors/testdata/cursor-acp-turn.jsonl`), lifecycle pinned by **process counts** rather than by
anything the adapter says about itself, and one live multi-turn conversation through the merged
seat (`-tags=live`):
```
turn 1 phase=done elapsed=9.744s act "Read File" → ok body: github.com/sanlee-ys/telltale
turn 2 phase=done elapsed=1.120s same process body: github.com/sanlee-ys/telltale
```
Turn one read a file it could only have read by running a tool in the workspace; turn two
answered a question only turn one's history could answer, from the same process, in 1.12 s. That
is the whole of §9.33's prize, measured through the room rather than through an instrument
standing beside it.
**Not verified here: macOS.** Every arm ran on Windows 11, and the Mac's ACP seat is unmeasured.
That belongs in `PARITY.md` rather than in this section.
### 9.37 /arena: the seats race in worktrees, and the human picks the winner
`/arena ` is one brief raced across every seated vendor, each attempt in its own git
worktree, compared by diff instead of by prose. It is §9.2's thesis — independent answers ARE
the product — applied to code, where "independent" stops being free: four writers in one shared
tree are not four answers, they are one trampled tree. The isolation the manager lane built its
whole category on (Crystal's same-prompt sessions, claudexor's best-of-N envelopes,
parallel-code's AI Arena) is what makes four *write* attempts comparable at all.
Ruled 2026-08-08, four decisions and their reasons:
- **Per-turn, typed at the room** — not a launch posture. §9.17's rule; a race is something you
want *about a brief*, not about a session.
- **Every attempt is a FRESH session.** All three comparable products race fresh, a continued
thread would anchor each seat on its own prior answers, and whether resume even survives a cwd
change is measured for none of the spawn-per-turn seats — so fresh is also the only option
that costs no new vendor measurements. Mechanically: every seat goes through the `FirstTurn`
one-shot it already implements, the persistent seat included. The room's live process, saved
ids and conversations are untouched — dispatch guards the session-id capture so a race's
throwaway ids can never replace the room's saved threads (the reattach-swap bug, killed in a
test before it could exist).
- **Worktrees are KEPT until the user deletes them**, named `-arena-t-` as
SIBLINGS of the workspace with branches `arena/t/`. Siblings, not a state
directory: kept-until-deleted means the user must SEE what is kept, it matches the README's
own worktree convention, and /cd's sibling resolution makes `/cd repo-arena-t7-codex` work
with zero new code.
- **Comparison lands in-column** — `git diff --stat` against a base SHA recorded once before any
seat spawned, rendered in the transcript's boundary grammar; `y` yanks the full diff (capped
at 1 MB, truncation stated). Three outcomes, three renders: a diff, a measured "no changes
against ", and "diff unavailable: " — zero, absent and degraded stay three
different facts (§4a.1).
Two mechanics carried in from the deep-read of claude-squad's `session/git/diff.go`, because
they are the difference between a diff surface and a lying one: the diff anchors on the
**recorded base SHA**, never HEAD, so an attempt that commits mid-turn cannot show an empty
diff; and `git add -N .` runs before diffing so an attempt whose whole answer is a NEW file
cannot read as "no changes" — the false zero, again.
Posture is `PostureWrite` for every racing seat, stated rather than hidden: a one-shot process
has no channel to be asked on, so the gate structurally cannot exist here, and the containment
is the worktree — which is the whole reason the worktree exists. A read room refuses `/arena`
with the in-room remedy named (`/write lets it`), per §9.17's tell.
**What council deliberately does not do, having read the competition:** claudexor AUTO-ADOPTS
the winning patch into the live tree. This room offers the diffs and the human picks — adoption
is a git command the user runs against a kept branch, never an action taken for them. And a
race is not routable in v1 (`@codex`-only arenas): the value is the comparison, and a one-seat
race is an ordinary turn in a worktree, which `/cd` already provides.
Deliberately deferred, each its own change judged against this section — and every one has now
landed, each as its own change (the 2026-08-09 amendments below): ~~commit-per-turn inside arena
worktrees~~ (with its undo), ~~`.worktreeinclude` seeding (the first real arena run on a repo
needing `.env` will surface it)~~ (landed on exactly that argument, ahead of that repo showing
up), and ~~a deletion guard stronger than git's own refusal to remove a dirty worktree~~ (landed
as `/arena drop`'s counted refusals, beside `/adopt`). This paragraph briefly existed as two
half-struck copies of itself — two same-day changes each struck their own item and a text merge
kept both variants — collapsed back to one on the same day.
Verification note: the git mechanics (worktree creation from one base, add -N, the three
collection outcomes, the session-id guard, the renders, the yank) are all pinned by offline
tests against a real temp repository. ~~No live vendor has raced yet.~~ **The first live race
ran 2026-08-09** — turn 4 of a real room on the Windows box, four seats dispatched, `/arena`
against this repo at 422b1c3 — and it paid the debt this note carried while measuring exactly
where the predicted risk lived:
- **The core is verified live.** Worktrees created as named siblings, three seats raced fresh,
ranks rendered in host-observed order (agy 1st · 7s, codex 2nd · 15s, claude 3rd · 19s), and
the zero-render said "no changes against 422b1c3." on every finisher — honest zeros, since
the brief was a harness check that asked for no changes. The room's threads survived intact.
- **The cursor seat cannot race, by its own design.** The ACP refounding (§9.36) gives
`Cursor.FirstTurn` a deliberate refusal — "driven as a live ACP process, not as one child per
turn" — which arena's uniform one-shot path duly surfaced on the column. The fix is a
follow-up with its own shape: an EPHEMERAL ACP session per race (spawn in the worktree, one
turn, kill), which is §9.36's machinery pointed at a throwaway session. ~~On the deferred list
below until someone builds it; until then a race is honestly 3-of-4 on Windows.~~ **Built
2026-08-09, in exactly that shape — second amendment below.**
- **The write seat hit the allowlist-prefix trap.** claude's one probe — `git -C
status --short --branch` — met `autoAllowedTools`' `Bash(git status:*)` rule and failed the
prefix match, so an ungated print-mode seat had an approval request and nobody to ask (act
rendered ✗, correctly). The fix is NOT a blind `Bash(git -C:*)` — that constant also serves
`--auto` in real workspaces, where pre-approving every `-C` form is a wider grant than the
verbs it lists — and the vendor file already warns its rule grammar has not been driven.
~~What this needs first is one measured probe of whether the rule syntax can scope a verb
behind `-C` at all; the finding is filed, the measurement is the next step.~~ **The probe
ran, 2026-08-09, and closed this.** A four-arm probe on the reference box measured the
matcher as prefix-only, with no rule spelling that scopes a verb behind `-C`, so
`Bash(git -C:*)` stays rejected and a seat runs plain `git` — its cwd is already the
workspace. The record is `STATE.md`'s "Closed without code" entry with its 2026-08-10
amendment, plus `autoAllowedTools` in `internal/council/vendors/claude.go`. Do not re-open
it without a new measurement.
**Amendment, 2026-08-08: the finish line and the `d` key.** Two deferrals came off the list:
- **Every racer's arena block now carries a finish line** — *"2nd of 4 · done · 25.0s"* — and
each part keeps its own epistemics. The rank is the order the ROOM saw seats land
(finishColumn call order, host-stamped; event batching bounds the resolution, which is the
honest limit of what was measured — a vendor's own claim about when it finished is an
inferred value wearing measured clothes and is not consulted). The phase word is welded to
the rank on purpose: "2nd · failed" and "2nd · done" are different facts, and a bare number
would let a fast crash read as a podium. A DNF ranks — it landed, just not well. The elapsed
is the column's own measured clock. parallel-code's results screen is the pattern source,
minus its star rating, which is a judgment no gauge here is allowed to render.
- **`d` flips the focused seat's arena block from stat to the full patch** and back. Per
column, because reading A's stat against B's whole diff is a legitimate way to compare. The
frame renders at most 400 patch lines (`arenaDiffScreenLines`) — a render cost bound, not a
data bound — and the cutoff names how many lines it dropped and both routes to the rest
(`y`, and the worktree itself). Three refusals with three sentences: no race this turn, a
measured nothing-to-show, and an unreadable diff carrying its reason. ~~Plain text, no diff
colouring yet: +/- prefixes are the first signal and survive `--ascii`; colour through the
existing palette is a later, separate change under style.go's no-new-hues rule.~~ **Coloured
2026-08-09, through the existing palette and nothing else** (`Styles.ForDiffLine`): added
lines wear `SevOK`, removed lines `SevCrit`, headers (`diff --git`, `index `, `---`/`+++`,
`@@`) the muted chrome style — no new hue, per style.go's rule. Classification reads the raw
prefix with headers matched first, so `+++` never wears the addition's green. The `+`/`-`
prefixes stay the first signal: `PlainStyles` renders the same bytes as before, which is why
no golden moved, and `--ascii`/`NO_COLOR` see exactly the frame they always did. The stat
blocks (interim and final) stay unstyled — a stat is a summary, not a patch line.
**Amendment, 2026-08-09: the live stat — the race shows the diff growing.** Until now a
racer's stat appeared only when its column finished; the audience of a 20-second race
watched three spinners and then a scoreboard. The pattern is the manager lane's
event-triggered diff refresh (codeg's), rebuilt under this room's honesty rules
(`internal/council/arenalive.go`):
- **Event-triggered, throttled, off the loop.** Stream activity on a racing column (text or a
tool call — a session id arriving is not evidence the tree moved) ARMS a re-read of that
seat's worktree — `git add -N . && git diff --stat`, the same two claude-squad
mechanics the finish-time read carries, stat only (the full patch stays a finish-line
deliverable). An armed seat is read at most once per `arenaRefreshInterval` (2 s: the read
is a subprocess pair, the audience is human, and the first live podium ran 7 s/15 s/19 s —
faster buys frames nobody can distinguish), timed off the tick-stamped `State.Now` so the
throttle is testable and Render stays pure. The read runs as a Bubble Tea command
(goroutine → `arenaStatMsg`), never inline in Update, never in Render; one read in flight
per seat, a due refresh that finds one running SKIPS rather than queues. An idle seat never
arms, so an idle seat is never read; a seat whose worktree failed setup has no refresh slot
at all, by construction.
- **The interim marker ruling.** A mid-race read is a measured value at a moment that is
already past, so the block's label is `arena · so far` — the "so far" is the whole marker,
the same honesty spend as an estimate's `~` — and it withholds the finish line's receipt
(branch, tree path, rank), which would dress an interim block in the final's clothes. Three
states stay three renders (§4a.1, mid-race edition): no read yet is the nil pointer and
renders NOTHING (absence, not a zero); a read that returned empty says "no changes yet
against " (the "yet" is what separates a running seat's measured zero from the
final's settled one); a failed read carries git's first stderr line, never dressed as
no-changes. A failed read degrades only the live stat — the race runs on — and
`arenaRefreshMaxFails` (3) consecutive failures end the seat's live stat WITH THE STOP
NAMED on the column ("stopped re-reading … the finish-time diff still runs"), because a
gauge that quietly freezes goes on reading as live. A success resets the count: the
likeliest failure is the refresh contending with the vendor for the worktree's own index,
and one contended read is not evidence the tree is unreadable.
- **The finish-time `collectArena` read stays the authoritative final and REPLACES the last
interim — cleared, never merged.** The refresh state lives on the turn, so teardown ends
all refreshing with no cleanup path to forget; a read that outlives its turn or its seat
arrives as a stale message and is dropped by comparison (turn number, final-already-landed),
not by hoping the timing worked out. The one collision the feature introduces is named in
the code: an interim `add -N` holds `index.lock` for milliseconds, so a final read that
fails while a refresh is in flight is retried once — reporting "diff unavailable" for a
lock this feature itself held would be the refresh degrading the read it exists to
complement.
Verification note: the mechanics — arming, the throttle, single-flight, the three interim
renders, replacement by the final, stop-on-turn-end and stop-on-repeated-failure, the
`add -N` false-zero property of the interim read — are pinned by offline tests
(`arenalive_test.go`), the git ones against a real temp repository. ~~No live race has
watched the stat move yet~~ **Half paid, 2026-08-09, and the halves are worth keeping
apart.** The block APPEARING mid-race and reading honestly is live-verified by the give-up
amendment's own race below: the stuck cursor racer's `arena · so far` read "no changes yet
against \" for 26m40s, which is the interim empty state observed live for longer
than anyone wanted. What is still owed is the other half — a "so far" block that GROWS and
then swaps for the settled block at the finish — because no live race is recorded as having
watched a NON-empty interim stat change. One `/arena` against a brief that changes files
pays it, and it is stated here rather than implied paid.
**Paid, 2026-08-15/16, race t9 — the first 5-of-5, and the growing half both.** The owner
raced a file-changing brief across all five seats. The Antigravity and Cursor columns both
drew a non-empty `arena · so far` that CHANGED on a later refresh (one file, then two) and
was replaced by the settled block at landing — the growing half, watched twice over. The
race's full record: three clean finishes with ranks and receipts (claude `1st of 5 · 50s ·
committed 2770c0c`, grok `2nd · 1m8s · 4aba168`, codex `3rd · 1m10s · 6874ff2`), and two
seats given up with `x` after ~11 minutes (agy `4th · cancelled · committed 91fc53e`,
cursor `5th · cancelled · committed 6b94b55`) — the give-up's second and third live
exercises, and both cut seats kept their commit receipts exactly as the finish-line design
promised. The cause of the two stalls was measured from OUTSIDE the room before the cuts:
both racer processes were alive with ~zero CPU over an 8-second sample and no go toolchain
process existed anywhere, so the work was done and the vendors' own turns had stalled —
agy inside a `manage_task`/`schedule` poll loop that stopped polling, cursor after its
final tool step. `/adopt claude` then exercised the dirty-room refusal live (the probe
turn's two throwaway edits held the tree; the card named them; the operator restored and
re-ran) before cutting `adopt/t9-claude` and landing the `--no-ff` merge cleanly — the
refusal path's first live run. **One new gap, found by the same race:** `/trace` was armed
before the turn and recorded NOTHING for it — the trace file holds only the preceding
ordinary turn's line — so an arena turn produces no per-seat spawn/wait/stream split, and
grok's timing on that axis stays unmeasured. Recorded as an unowned gap in STATE.md.
**Amendment, 2026-08-09: the cursor seat races too, on a throwaway session.** The deferred
follow-up the verification note filed is built, in the shape it predicted. For an arena turn —
and only there — dispatch recognises the Conversational seat and, instead of the `FirstTurn`
one-shot it deliberately refuses, launches a throwaway `cursor-agent acp` server rooted in that
racer's worktree, runs exactly one `session/new` and one `session/prompt` through §9.36's own
protocol driver, and kills the process when the column lands. It is `startEphemeralRacer`
(persistent.go), a sibling of `spawnSeat` reusing `Cursor.Open`, `acpProtocol` and the counted
`startRPCSession` spawn — no second ACP implementation exists to drift, and where the client
was welded to the room seat's lifecycle the seam extracted was placement, not protocol.
- **The room's conversation is untouchable by construction, not by discipline.** The race
session opens with an empty resume id (never persisted, never resumed), registers on the
TURN (`turnState.arenaEphemeral`) rather than in the seat-process registry — so a live
persistent cursor seat and its racer coexist without either being mistaken for the other —
and the throwaway session id it reports is refused by the existing arena guard before it
can reach the saved threads or room.json.
- **Kill, never wait.** §9.33 measured this vendor's process lingering ~2.5 s after answering,
so the racer is killed at its own finish line — before the diff is read, making the receipt
a snapshot of a stopped attempt — and on a protocol-reported failure (an ACP server survives
its own refusals, so no exit event would ever have come), on ctrl+c, and at room teardown.
Its context is the turn's rather than the room's, which is the backstop on every one of
those paths; a seat cannot be cleared mid-race at all, because `askClearSeat` refuses while
a turn is in flight.
- **Two processes now wear one vendor id during a race**, so the eleventh amendment's
stale-exit guard grew an attribution rule on the same liveness test it already trusts: an
exit that arrives while the racer is alive can only be the room's idle seat dying in the
background (forgotten, race untouched); one that arrives after the racer is dead is the
racer's own, and must not be eaten by the guard reading a live room process as "this seat
is fine" — that would hang the race column forever.
- **The exits keep their epistemics.** A racer that dies without its end-of-turn response
FAILED — on this seat the turn's end is a response, so a bare exit means no answer arrived,
and rendering it done would be the empty-success this seat's missing result line makes
possible. A turn that ends cleanly having streamed nothing lands done with a note naming
the ambiguity, because a silently-working racer and a broken chunk parser are identical on
this wire (§9.36's stated loss). Token usage stays what ACP makes it: absent, never zero.
And the containment phrase every racer carries — write posture, contained by the worktree —
is stated at its weakest here: §9.36 measured workspace trust not applying over ACP, and an
arena worktree is a freshly created, never-trusted directory, so nothing but the session's
cwd scopes the attempt. The posture detail already says so; the worktree gives it more
force, not less.
Verified offline only: fixture-driven tests (arena_cursor_test.go) pin the spawn choice, the
untouched room thread, the kill on finish / protocol failure / cancel / teardown, both
degraded exits, and the exit-echo not re-ranking the race. cursor-agent was not installed
where this was built, so ~~this amendment owes a live race on the Windows box~~ **the
amendment owed a live race; two have now run it, and what they paid is narrower than the
word "verified" would suggest.** The throwaway racer has been spawned live and rooted in
its own worktree — it streamed for 26m40s on the race the give-up amendment below records,
and was cut loose mid-race with `x` on race t9 — so the spawn, the live interim read
against its tree, and the kill path are measured. **A clean completion is still owed**: no
live race is recorded in which this seat's own `session/prompt` resolved and the racer was
killed at its own finish line with its diff read. Until then the *finishing* half of this
amendment stands on `arena_cursor_test.go` alone, and that is the honest split. *(The
2026-08-15/16 5-of-5 race cut this racer a second time — alive at ~zero CPU with its edits
complete and committed on the cut, ~11 minutes in — so the debt stands and gained a second
data point: two live races, two stalls, zero self-finishes.)*
**Amendment, 2026-08-09: every attempt survives as a commit, and `u` takes one back.** The
commit-per-turn deferral came off the list, and it brought the rollback it makes possible
(mechanics stolen with attribution: Crystal's commit-per-turn checkpoint, cc-haha's turn-level
undo). Two halves that stack:
- **Commit-per-turn.** The moment a racer lands and `collectArena` has read its diff, the
worktree's whole state is staged and committed onto `arena/t/` — subject
`arena t: ` (64-byte cap) — so every attempt is durable on its
branch: diffable, adoptable, rollbackable, and the worktree itself becomes deletable without
losing anything. Staging everything is correct *there and only there*: the tree contains
nothing but the racer's own output, so the reason blanket staging is wrong in a real
workspace does not exist in that one. The sha the column renders is exactly what
`git rev-parse HEAD` returned, shortened for display only. On a machine with no git identity
anywhere (CI runners, fresh boxes) the commit carries a fallback identity via per-command
`-c` flags — never a config write, which on a worktree would land in the shared repo config,
i.e. in the room's repo. A commit that cannot land (a stale ref lock, a signer that cannot
run) degrades **that seat's receipt only**, as `not committed: ` —
the diff was still read, the race and the other racers are untouched. A racer that committed
for itself mid-turn already parked its attempt; its own tip is reported rather than papered
over with an empty commit — and the diff still answers against the recorded base, so the
mid-turn commit cannot hide the work.
- **The empty-commit ruling.** A zero-diff attempt commits **nothing**, and that is a ruling,
not an omission: an empty commit would be a receipt claiming work that did not happen —
§4a.1's false zero, mirrored into the write path. The seat renders no commit line and no
failure either (nothing was owed); "no changes against \." stays the whole story, and
the branch tip staying at base is the machine-readable form of the same fact.
- **Undo-the-whole-turn.** `u` on a focused arena seat, y/n-gated exactly like `c` (a stray
keystroke must cost a `y` before it costs an attempt), runs `git reset --hard ` inside
the racer worktree **only**. Branch and tree agree by construction rather than by a second
command: the worktree has its arena branch checked out, so `--hard` moves that ref itself.
The safety argument is an explicit path guard, not trust in recorded state: the reset runs
only on a path equal to the arena-tree name recomputed from the room's *current* workspace,
turn and vendor — a name that structurally cannot be the workspace itself — and anything
else refuses before git runs. Refusals are four sentences for four facts: no race this turn;
the attempt changed nothing (nothing to take back); already undone (pressing again is not
more undone); and the reset itself failed, carrying git's own first stderr line. After an
undo the stat stays on the column — the measured record of what the attempt changed — under
an "undone" line saying the tree and branch no longer hold it.
- `u` landed on the help panel's room-controls row at its exact 114-cell budget by trading the
word "worktrees" for it: the arena block prints the worktree path on every race, so that
clause restated something the screen already teaches, while an undo key documented nowhere
is a control nobody finds.
Verification note, on the same terms as the section's own: the git mechanics — the commit
landing on the branch with the racer's files, the per-seat degrade, the zero-diff skip, the
self-committed tip, the undo round-trip (files restored, created files gone, branch back at
base), the path guard, and every refusal sentence — are pinned by offline tests against real
temp repositories. ~~No live race has exercised either half yet.~~ **Commit-per-turn is
paid; the undo is not, and it is now the last unpaid item in this section.** Race t9
(Windows box, 2026-08-09) landed the claude racer's attempt as `cf69634` — subject
`arena t9: write table-driven tests for the small render helpers …`, parent `e1bf983`,
which is the base the race recorded — onto `arena/t9/claude`; that commit outlived
`/adopt`, `/arena drop` of its worktree, and a merged PR (#164), with the branch
`adopt/t9-claude-helpers` left as the receipt. **`u` has still never run against a real
racer commit.** The debt is one `/arena` against a brief that changes files, then `u` on
the finisher between turns.
**Amendment, 2026-08-09: `.worktreeinclude` — a race carries the files git ignores, when the
repo names them.** The seeding deferral came off the list, on the schedule the original note
predicted (a real race on a repo needing `.env` fails falsely on every seat at once, so the fix
is worth landing before that repo shows up). What was built, and the rulings inside it:
- **The file and the copy.** A `.worktreeinclude` at the room repo's root — gitignore-style
patterns, one per line, `#` comments and blank lines ignored — and during `arenaSetup`, after
each racer's worktree is added, every matching file is copied from the room repo into that
tree, relative paths preserved, parent directories created. The grammar is a documented
subset of gitignore's (bare names match at any depth, anchored patterns from the root, `*`
within a segment, `**` across segments, a directory pattern takes its subtree; no negation).
- **Copy only, never execute — the half of agent-deck deliberately not taken.** agent-deck (the
pattern source) pairs seeding with repo-carried setup scripts that run after the copy. That
half crosses a trust boundary this project has explicitly parked: byte-level trust gating is
on the parked list pending an audit, and a repo that can run code on the machine by merely
containing a file is a different product with a different threat model. Copying bytes into a
tree the room already owns is containable; execution is not.
- **Candidates are untracked files only** (`git ls-files --others`, ignored files included —
exactly the set a fresh worktree lacks). A tracked file already arrives with the checkout,
and seeding the room's possibly-dirty copy of one would plant the room's own edits in every
seat's diff — a lying diff, §4a.1's class. Known limit, stated: a seeded file that is
untracked but *not* git-ignored still surfaces through collection's `git add -N .`; name
git-ignored files and it cannot.
- **Containment.** Patterns resolve from the repo root and matches come from git's own
enumeration, so they structurally cannot leave it; absolute and `..`-carrying patterns are
refused by name anyway, per pattern, so one bad line disables only itself. Symlinks are never
followed — Windows is primary and symlink semantics differ per platform — a symlink match
copies nothing and says so.
- **The budget: 64 MiB per seat** (`seedBudgetBytes`). Exists for the node_modules pattern — an
over-broad line must fail loud and named, not hang the room copying a dependency tree into
four worktrees. Over-budget is refused wholesale (copying *some* of the file's list would
hand every seat a tree that half-works), with the measured total in the sentence, and the
budget is enforced again on actual bytes during the copy, because files grow between stat and
copy.
- **Honesty in the column.** "seeded 3 files" is the count actually copied into that seat's
tree; no `.worktreeinclude` means no line at all (zero and absent stay two facts). A pattern
that matches nothing is a named notice, not silence — an allowlist-shaped file fails both
ways — and not a failure either. A copy error degrades that one seat with the path and the
error's first line, through the same per-seat lane a failed worktree add uses; the race runs
on, and the half-seeded worktree stays on disk (kept-until-deleted receipts include the
broken ones).
Verification note, same shape as this section's original one: the mechanics — the copy into
each racer tree, nested parents, every named refusal, the budget, the per-seat degrade channel,
zero-vs-absent on the seed line — are pinned by offline tests against real temp repositories.
**No live race on a repo that actually needs a `.env` has run yet; that run is the debt this
amendment carries**, and it is the same debt the original note carried for the core, paid the
same way.
**Amendment, 2026-08-09: the end of a worktree's life — `/adopt` and `/arena drop`.** The
deferred deletion guard came off the list, and adoption moved from a printed suggestion to a
typed room command; both are §9.17 verbs, reachable mid-session with no flag and no relaunch
(`lifecycle.go`). The shapes borrow from the two products that had already worked this seam —
claude-squad's adopt, Pane's deletion guard — with one deliberate fork each:
- **`/adopt ` merges; it does not check ~~out~~ the racer's branch out.** claude-squad's
adopt is a checkout of the
attempt's branch over the user's — which moves HEAD, rewrites the tree wholesale, and leaves
the user's own branch behind. That is more state than the act requires, and §9.37's whole
posture is offer-never-take, so council does the least-magic git operation that lands the
work: `git merge --no-ff arena/t/` in the room's repo, ~~on the branch the user is
already standing on~~ **on a fresh branch council cuts for it — ruled 2026-08-11, the last
block in this section; the merge itself is unchanged**. `--no-ff` keeps the adoption a visible event in history — the merge
commit is the receipt saying where the work came from. Because arena seats leave their work
uncommitted (commit-per-turn is still deferred), a dirty attempt is first committed in its
OWN worktree, on its OWN arena branch, under the user's own git config — and the y/n card
says so, naming the exact command(s) y will run, the flow write gate's contract. Hard
precondition, refused by name with the path count: the room tree must be CLEAN
(`git status --porcelain` empty) — a merge writes into that tree, and adopt must never eat
the user's uncommitted work. A racer that changed nothing refuses (an empty merge commit
would claim work that does not exist); a merge that conflicts is `git merge --abort`ed —
tree restored, attempt intact on its branch — and the notice hands the merge to a human. A
merge that failed before starting reports the tree as untouched instead: the two endings
are different facts (§4a.1). Posture is not consulted: read/write governs the seats, and an
adopt runs on the user's own y, the same footing as /cd.
- **`/arena drop ` (or `all`) deletes tree + branch, guarded; the force is a spelling,
not a keystroke.** Two guards, each refusing with exactly what would be lost and the way
forward: a worktree holding uncommitted changes (counted), and an arena branch holding
commits the room's HEAD cannot reach (counted, with `/adopt ` offered beside the
force). The force form is a trailing bang — `/arena drop codex!` — re-run by the user,
chosen over a second y/n on purpose: y is one keystroke answered against a notice half-read,
while the bang travels in the command, records that destruction was asked for, and cannot be
produced by a stray key. (`/adopt` keeps y/n because its act is additive and revertible;
drop orphans work.) Mechanics are `git worktree remove` (git's own `--force` only when the
user spelled it) then `branch -D`, argv via gitOut — and the path check is mechanical: a
tree is only ever removed if it re-derives, from the recorded race's own workspace/turn/seat
through the same `arenaTree` that minted it, to exactly the recorded path. A receipt entry
that fails that check is refused even under force. `drop all` degrades per seat rather than
refusing wholesale: clean trees go, survivors are named with their reasons.
- **The target is the RACE'S receipt, not the column.** `Column.Arena` is a per-turn fact the
next dispatch clears; the worktrees are kept until deleted. So dispatch records the race —
workspace, turn, base, each racer's tree — on the model (`arenaRace`), in memory only:
room.json stays keys-and-numbers, and a room reopened after a quit finishes the lifecycle by
hand with the same git commands, against worktrees that are visible siblings precisely so no
session state is needed to find them. Grammar note: only the exact two-word form
`/arena drop [!]` is the verb; anything longer after `/arena` is a brief and races as
prose, the roomcmd vocabulary rule applied inside the one command that takes free text.
The help panel's room-commands row is at its width budget and does not name the two verbs;
they are taught by the slash refusal (which lists `/adopt` in the live table), by bare
`/adopt` and bare `/arena drop` answering with usage, and by every guard refusal naming its
remedy. Verification, ~~owed~~ **paid 2026-08-09**: the git
mechanics — merge, commit-then-merge, conflict abort, both guards, the force, the path check,
`drop all`'s partial degrade — are pinned by offline tests against real temp repositories
(`lifecycle_test.go`), and ~~no live adopt has run on the Windows box~~ **race t9's winner
went through `/adopt` and then `/arena drop` on the Windows box**, which is why no
`arena/t9/*` branch and no `telltale-arena-t9-*` sibling exist on it while every earlier
race's leftovers sit exactly where they were left. What that adoption then cost in hand-run
git is the open question at the end of this section, filed rather than ruled. Two guard
paths a SUCCESSFUL adopt cannot reach still rest on `lifecycle_test.go` alone: a merge that
conflicts, and a drop refused for unmerged commits.
**Amendment, 2026-08-09: the race numbers itself off the refs, and a failed race says why.**
A live `/arena` at turn 3 (Windows box, real room) failed on all four seats, and each column's
whole explanation was `arena: Preparing worktree (new branch 'arena/t3/')`. Two
measured defects, one incident — the second is what made the first expensive:
- **The collision.** Kept-until-deleted cuts both ways: arena branches and worktrees outlive
the room, but the turn counter — and the in-memory race receipt `/arena drop` needs
(`Model.lastRace`) — reset with every launch. So a fresh room's turn 3 minted the exact
names an older room's turn 3 had already parked, `git worktree add -b` refused every seat,
and drop could not reach the old trees because their receipt had died with the old room. The
only remedy was hand-run git. The fix reads instead of guessing: at setup the race number is
`arenaRaceNumber` — one past the highest N among the repo's existing `arena/t/...`
branches (`git for-each-ref` over `refs/heads/arena/`, argv via gitOut), floored at the turn
number, so a repo with no leftovers keeps racing as `t`. The refs are the one record
that shares the leftovers' lifetime, which is what qualifies them to number the race; a scan
that cannot run degrades to the turn-number floor with the race running, because a broken
for-each-ref must not brick `/arena`. The number is recorded once
(`turnState.arenaRaceN`, `arenaRace.raceN`, `ArenaResult.RaceN`) and EVERYTHING that mints
or re-derives a name reads it — the branch on the receipt, the `arena t:` commit subject,
`/adopt` and `/arena drop`'s re-derivations, undo's path guard — because the turn and the
race now legitimately disagree, and one call site still reading `Column.TurnN` would aim a
verb at names the race never created. The arena block's render is untouched: it already
shows the branch name, which carries the (now honest) `t`.
- **The lie about the collision.** gitOut surfaced the FIRST stderr line of a failed command,
and `worktree add` prints progress chatter ("Preparing worktree ...") before its
`fatal: a branch named '...' already exists` — so the column showed the narration and
swallowed the diagnosis. gitOut now prefers the first line git itself marks as the problem
(`fatal:` / `error:`), falling back to the first non-empty line only when no marked line
exists (some refusals print bare prose). git's own prefixes are the measured marker of
which line is the error; the old rule displayed the nearest string to the failure instead of
the failure. If a residual collision still happens despite the scan — a sibling directory an
old room left with no branch to be scanned, a ref minted between scan and add — the seat's
error now carries the fatal line plus the named remedy (`git worktree remove` /
`git branch -D`), since those leftovers are exactly the state no receipt can reach.
Both fixes are pinned by offline tests against real temp repositories: the two-line-stderr
collision fixture (the live transcript, replayed), renumbering past stale `t3` branches with
every seat racing clean, the turn-number floor on a failed scan, the residual-collision
sentence, and adopt/drop/undo driven end to end against a race whose number outran its turn —
leftovers untouched throughout. ~~The live re-race is owed~~ **The live re-race ran, and has
kept running ever since**: every race after the fix has been numbered over a growing pile of
leftovers — 27 `arena/t/` branches and 28 sibling worktrees, t2 through t8, are
still on the reference box — and race t9 raced all four seats clean over them, which is the
claim. One consequence worth knowing before the next race: t9's branches were dropped, so
the highest surviving `arena/t` is t8 and the scan will mint `t9` again.
**Amendment, 2026-08-09: `x` gives up on one racing seat, and the race runs on.** The second
live `/arena` (same day, Windows box) measured the gap: four seats raced, three landed
(7m51s / 26m28s / 5m07s), and the fourth — the cursor throwaway ACP racer — streamed for
**26m40s** with the live stat honestly reading "no changes yet against \" the whole
time. The operator sat ~20 minutes after the race was effectively decided, because one stuck
racer holds the WHOLE turn hostage: ctrl+c is the only exit and it cancels everything. The
room displayed the truth and offered no per-seat act on it. The act built:
- **`x` on a focused, still-racing arena seat, mid-turn** — the one per-seat key that runs
while a turn is in flight, because mid-flight is the only time it means anything.
y/n-gated exactly like `c` and `u` (a stray keystroke must cost a y before it costs a
process), and the question names the vendor and what y does. On y, the room kills THAT
racer only — the ephemeral ACP session when one is racing, else that vendor's one-shot
process through `turnState.arenaHandles`, the per-vendor record dispatch's arena branch now
keeps beside the flat `handles` list (the flat list stays: cancel and teardown are
all-or-nothing acts and never address a single process; the give-up is the first act that
does). The kill lands on the racer's side of the two-processes-one-vendor-id split — the
room's idle seat behind the same id survives, per applyEvents' existing attribution rule.
- **A given-up seat lands like any other finisher, wearing the honest phase.** The stream
tail is flushed, the elapsed stamped, the note says "given up after \ — anything
it wrote is in the diff", and the column retires through `finishColumn` with the CANCELLED
phase — the same phase and render ctrl+c's cancel produces ("cancelled — the output above
is partial" is that path's wording; this one names the give-up instead). Everything the
finish line already does happens unchanged: the racer dies before the diff is read (the
receipt is a snapshot of a stopped attempt), a dirty tree commits its receipt onto the
arena branch, a clean tree stays a measured zero with no commit, the rank is stamped in
host-observed landing order — a DNF finished too, and the render welds the rank to the
phase word so "4th · cancelled" cannot read as a result — and the interim stat clears. The
seat leaves the turn's live set through the same drain every landing uses, so **the turn
ends when the remaining seats land** — which is the whole point.
- **Three refusals, three sentences** (the undo key's rule): no turn in flight; ~~an ordinary
turn — its seats share one fate by design, this key is arena-only and **ctrl+c remains the
whole-turn act**, said in the refusal~~ **(REVERSED 2026-08-17 — the block below; the key
now runs on an ordinary turn, and the refusal that took this one's place is a turn ctrl+c
is already cancelling. ctrl+c is unchanged and is still the whole-turn act)**; and a seat
that already landed (its result is
settled — a y arriving after the seat lands under the question refuses the same way,
killing nothing and re-ranking nothing). The help panel's room-controls row is at its
exact 114-cell budget and does not name the key; it is taught by these refusals and by
this amendment, the way `/adopt` and `/arena drop` are taught by theirs.
Verified offline only, the section's standing debt shape: real-temp-repo plus fake-session
tests (giveup_test.go) pin the ephemeral kill and the cancelled landing with rank and
committed receipt, the keyed one-shot kill with every other racer's handle surviving, the
room process surviving its racer's give-up and the exit echo landing inert, the turn ending
when the remaining seats land, all three refusals, the y/n/stray gate, and compose leaving
`x` a letter. ~~A live give-up on the Windows box is owed~~ **Paid on race t9** (Windows box,
2026-08-09): the cursor throwaway racer stalled, was cut loose mid-race with `x`, and the
turn ended when the remaining seats landed — which is the whole claim, measured. The three
refusals and the y/n/stray gate stay offline-pinned, and always will be: a live race has no
way to exercise a refusal it never trips.
**Amendment, 2026-08-17: `x` gives up on one seat of an ORDINARY turn too — the one-fate
line is reversed.** The amendment above refused the key outside a race on a stated design
position: *an ordinary turn's seats share one fate by design*. The owner reversed that
position on 2026-08-17, and the reason is that the position was **written for the four-seat
room**. It said, in effect, that a brief and its answers are one act, so a seat that is
still working is the turn still working. The five-seat room (§9.39) supersedes it: with
five vendors on an `@all` turn, **one stalled vendor while the other four have answered is
the most probable live failure this room has** — it is the failure two live races already
produced inside `/arena`, and nothing about it is a property of worktrees. The hostage
argument that built the key does not change when the brief is prose instead of a race. Only
the cost of the cut changes, and that is what the room now says per seat kind.
- **How a seat is stopped is per seat kind, and the three arms are not interchangeable.** A
batch seat (codex, agy, grok) is KILLED, through `turnState.seatHandles` — new plumbing
that mirrors `arenaHandles`, not a widening of it, because `arenaHandles` also answers
`arenaRacing`'s question about whose exit a `KindDone` is while two processes wear one
vendor id, and an ordinary handle in that map would send every ordinary exit down the
racer's branch. The persistent claude seat is **INTERRUPTED** (`interruptSeat`), never
killed: killing it would work and would also throw away the conversation and the
session-init cost that bought it, so cutting one turn would silently make the next one
expensive — `cancelTurn`'s own argument, applied per seat. The next brief resumes it. The
ACP cursor seat needs no third arm: on an ordinary turn it IS a persistent seat and takes
the interrupt, and the throwaway racer the arena kills only exists during a race. The
flat `handles` list is untouched, for its own reason: cancel and teardown are
all-or-nothing acts that never address a single process.
- **Four endings, four sentences.** The cut column lands CANCELLED through `finishColumn`,
keeping everything it streamed, and its note has to stay distinguishable from the three
other ways a column ends with no answer: *not addressed* ("not addressed in turn N", with
`Column.Skipped` set), *ctrl+c* ("cancelled — the output above is partial", still the
whole-turn act), and *a measured empty answer* (a body reading "[Turn completed with 0
text chunks streamed]" under `done`). The give-up's own note names the elapsed, whether
anything had arrived **when it was cut**, and what became of the seat — past tense on
purpose, so a killed child's last buffered chunk landing after the column retires cannot
make the sentence false. **A seat that streamed nothing must never acquire the
placeholder**: that would be §4a.1's false zero in its sharpest form, a seat the operator
stopped claiming to have measured nothing. `testdata/golden/given-up-vs-zero.txt` pins the
two side by side. Holding that took one fix outside the give-up itself, and it is the same
defect the eleventh amendment's end-of-turn branch was already fixed for: the placeholder
is a claim that THIS turn completed, so only a column still in a live phase may acquire
it. The phase write on the `KindDone` exit path had always been guarded and the BODY write
had not, so an exit landing on a column that had already ended overwrote it. The give-up
makes that reachable — a cut seat's child can exit after the turn boundary, past
`turnState.givenUp`'s lifetime — and the guard is now on the body write too.
- **The cut seat's own late events land inert, by guard rather than by luck.**
`turnState.givenUp` is recorded BEFORE anything is stopped, because both stops provoke
one more event: a killed child drains its buffered stdout, and an interrupted persistent
seat answers with its own failed `result` (measured is_error true, terminal_reason
"aborted_tools"). Unguarded, that error would overwrite "given up after 4m12s …" with the
vendor's abort text and record a vendor failure against a seat the user stopped. The
guard still does the PROCESS bookkeeping — an interrupted seat whose process later dies
for real is forgotten, so the next brief does not write into a closed pipe.
- **The refusal set changed by one.** "This turn is not a race" is gone. Its place is taken
by a turn ctrl+c is already cancelling: every seat is going anyway, so a per-seat act
would only re-label one of them. The other two are unchanged — no turn in flight, and a
seat that already landed — and the re-check on `y` is unchanged too, because events drain
between the card arming and the answer and the seat can land while the question is up.
Verification, stated honestly and in two halves. **Offline**: `giveup_test.go` pins both
seat kinds on a real `@all` turn — the batch kill reaching exactly one process with every
other seat still working, the persistent seat interrupted rather than killed and still
registered, the interrupted seat's own abort error not overwriting the give-up, the cut
seat that streamed nothing never acquiring the placeholder, the turn ending when the
remaining seats land, the four endings reading as four different sentences, the new refusal,
and `--ascii`/`NO_COLOR` parity on the cut column. No test here spawns a vendor
(`countSpawns`, per the council-test rule). **Live: a LIVE ordinary-turn give-up on the
Windows reference box is OWED**, as a dated payment before 2026-09-30. Offline tests cannot
exercise it: what is unmeasured is whether a real vendor's interrupt lands on a real
persistent seat mid-turn and whether that seat's NEXT brief actually resumes the
conversation, which is the whole claim the persistent arm makes and the one thing a fake
session cannot witness. Until that date and that run, the interrupt arm stands on
`giveup_test.go` and on `cancelTurn`'s already-measured interrupt, and this paragraph is the
record that it does.
**Amendment, 2026-08-09: the brief carries the conduct line — the one place the room adds
words.** A write-posture racer's confinement is its worktree, but the machine's git and gh
credentials are ambient, so a racer can reach GitHub — and one did: the codex seat took the
t5 gofmt brief, pushed its arena branch, opened a PR, waited out CI, and merged it into main,
then announced the same plan on the very next race. This section's founding ruling — the
room offers the diffs and the human adopts, never an auto-adoption — binds this codebase and
cannot bind a vendor that runs `gh pr merge` on its own initiative. The operator ruled the
same day: every `/arena` dispatch now prepends `arenaConduct` (arena.go) to the brief —
*"This tree is a race attempt in its own git worktree. Do not push, open pull requests, or
merge — the operator compares the attempts and adopts the winner. Do the work, verify it
locally, and stop."*
The bend to the brief-verbatim promise is bounded three ways, each pinned by test: races
only (an ordinary turn's prompt is byte-verbatim), a constant — the same published line for
every racer on every race, prepended so a long brief cannot bury it, leaving the cross-seat
comparison undisturbed — and recorded here rather than discoverable only in a wire capture.
Stated honestly for what it is: an instruction, not a control. A vendor can ignore it, and
the mechanical version of this boundary — credentials a racer cannot reach — is a different,
harder change that this amendment deliberately does not claim. ~~The live measurement owed:
the next race showing the codex seat stopping at its commit.~~ **Measured on race t9**: the
codex seat raced under the preamble and ended with "nothing pushed." One race is one race —
an instruction a vendor obeyed once is still an instruction, so the sentence above stands
exactly as written and this measurement does not promote it to a control.
**Open question, 2026-08-09 — RULED 2026-08-11, option (b): where should an adoption land?**
Filed from the first live
`/adopt` rather than decided, because the answer depends on a convention the room cannot see.
`/adopt` merges into the room repo's current branch — for most workspaces, local `main` — and
that is the smallest honest act the verb can perform. But an operator whose repos are run
branch→PR (this project's own convention, and this operator's standing rule across every
machine) then holds a merge commit on a local `main` that must never be pushed as-is; the
first live adoption ended with four hand-run git commands turning the merge into a PR branch
and resetting `main` back to origin. Options, ~~none ruled on~~ **(b) ruled, 2026-08-11 — the
block below**: (a) keep the current shape and
document the hand-off (an adoption is a local act; publishing is the operator's, as it
already is for every other commit); (b) `/adopt` onto a NEW branch cut from the room's HEAD
(`adopt/t-`?), never touching the current branch — closer to branch→PR, but the
room minting branch names in the operator's repo is a bigger footprint than one merge; (c) a
flag or second verb for each. The founding posture — offer, never take — leans (b) no further
than it leans (a): both are one revertible act on the operator's own y. Whoever picks this up
starts from the measured friction: four commands, once per adoption, only on branch→PR repos.
**Ruling, 2026-08-11: an adoption lands on a fresh branch, never on the branch the workspace
has checked out.** The owner ruled option (b), and the reason is the convention the room
could not see: the owner's workflow is branch-then-PR on every repo and every machine, and a
commit straight to `main` is never made. The room's own posture does not decide this — both
options are one revertible act on the operator's y — so the measured friction decides it, and
the friction is four hand-run git commands per adoption on every branch→PR repo. A fresh
branch turns that hand-off into one `gh pr create`. What was built:
- **The name is `adopt/t-`, cut from the room's current HEAD and checked out.** The
race number, never the turn, like every other arena name (the renumbering amendment above).
The seat is joined by a dash rather than a slash so `arena/t9/claude` and `adopt/t9-claude`
differ in the last segment, which is where a reader of `git branch` looks. `--no-ff` and the
commit-then-merge order are untouched: what moved is WHERE the merge lands, not what it does.
- **A collision takes the next free suffix** (`-2`, `-3`, … to 50, then a named refusal). The
collision is ordinary rather than exotic: race numbers repeat once a race's branches are
dropped, since the scan that numbers a race reads `refs/heads/arena/` alone, and an operator
can adopt, revert and adopt again. Reusing the name would either fail the checkout or land
the merge on an older adoption's branch, where the PR would carry work nobody asked about.
The free name is resolved WHEN THE CARD ARMS and carried to the `y`, because a card that
named `adopt/t9-claude` and then cut `adopt/t9-claude-2` would be the card describing
something other than what it ran — the one contract this gate exists to keep. One
`for-each-ref` over the adopt namespace answers every candidate at once and separates "no
such branch" from "git could not answer"; a scan that cannot run degrades to the plain name
with the adoption still running, and `git checkout -b` reports the collision with git's own
fatal line.
- **A failed adoption leaves nothing behind.** The branch is cut for a merge, so a merge that
conflicts or refuses ends with the branch deleted and the room back on the branch it came
from — an empty branch handed over as the receipt of a failure would be the verb charging
for its own failure. The two failure endings stay two facts (§4a.1): a conflict aborts and
says a human merge is needed, a merge that never started says the tree is untouched, and
both now also say where the room stands. A restore that cannot finish says THAT instead,
naming the branch the room is left on and the command that puts it back.
- **The alternative reading, recorded rather than taken**: leave the room standing on the
fresh branch after a conflict, so the human merge happens there. It is defensible — that is
where the resolution belongs — and it was not chosen, because the ruling covers where a
SUCCESSFUL adoption lands and a failure quietly moving the operator to a new branch is a
state change nobody asked for. If a live conflict makes the restore feel wrong, this is the
line to amend.
- **Existing refusals are untouched**, and so are their tests: the dirty-room gate with its
untracked-bystander rule, the zero-change refusal, the mid-turn refusal, and `/arena drop`'s
unmerged-commit guard, which reads the room's HEAD and therefore counts zero as soon as the
adopt branch holds the merge.
Verification, on this section's own terms: the mechanics — the branch cut and checked out, the
room's own branch not moving, the suffix on a taken name, the conflict restore, and the notice
naming the branch and `gh pr create` — are pinned by offline tests against real temp
repositories (`lifecycle_test.go`). **Paid 2026-08-14** by race t14 on the reference box: the
card named the branch and the merge before the `y`, `adopt/t14-claude` was cut and checked
out, `git merge --no-ff arena/t14/claude` landed as `91c5f3e`, `main` did not move, and the
notice named `gh pr create`. The push half was deliberately not run — the racer's payload was
a throwaway date comment, and opening a PR to prove a verb works would put noise in the
repository to record that the repository is fine.
#### A warm seat's racer could not finish, 2026-08-13 — measured, then fixed
**A race against a seat the room had already used never ended.** Race t10 on the reference
box: one Claude racer, a one-line edit, and the room rendered `streaming` for **21 minutes**
after the racer had exited. The vendor was not slow and did not fail. Its transcript ends at
52 seconds with a complete reply, and by then no process wearing that vendor id was left alive
except the room's own persistent seat. Nothing downstream ran — no diff, no commit, no rank,
no seed receipt — because all of them live inside `finishColumn`, and `finishColumn` was never
called.
**Both of the column's exits were closed at once, which is why nothing caught it.** A one-shot
racer ends its turn by exiting, so `KindDone` is its only retirement signal; the two earlier
paths do not apply to it, and each declines for its own correct reason. `arenaEphemeral` is
populated only for a `Conversational` seat, so `ephemeralRacer` is nil for this vendor. The
arena spawn never sets `turnState.persistent`, correctly — the racer *is* a spawn — so
`isPersistent` is false. That leaves `KindDone`, and `KindDone` reaches the stale-exit guard
first: a terminal event names a **vendor**, not a process, so a live entry in `m.procs` reads
as "this seat is fine" and the exit is discarded as a predecessor's. The room's persistent
Claude seat is exactly such an entry.
**The failure mode was already written down, one path over.** `KindDone`'s own attribution
comment names it for the ACP racer — a guard "reading a live ROOM process as *this seat is
fine* would leave the race column streaming forever and the turn unable to end" — and fixes it
with the `ephemeralRacer` check. `giveUpSeat`'s comment then observes that the guard eats a
racer's exit when a room process wears the id, and judges it harmless. For `giveUpSeat` that
judgement is right: a given-up column is already terminal when the exit lands. On the ordinary
path the column is not, and the same swallow is the difference between a race that finishes
and one that cannot. Two processes wear one vendor id for every racing seat, not only for the
one whose racer happens to be a live session.
**The trigger is a warm seat, and it is why this survived several clean races.** `m.procs`
must already hold a live process for that vendor when the race dispatches. Race before the
room's first ordinary brief and the guard never fires, which is what t9-on-2026-08-09 did when
all four seats raced and a winner was adopted. The drive that found this sent two ordinary
briefs first, so the seat was warm. It also means codex, agy and grok racers were never
affected: none of those vendors holds a persistent process, so their exits pass the guard
untouched. The bug reached exactly one seat, and it is the seat the room dispatches to by
default.
**The fix is attribution, not a hole in the guard.** `arenaRacing` is the one-shot sibling of
`ephemeralRacer` — keyed presence in `arenaHandles`, because a handle is not a session and
cannot be asked whether it is alive, which is precisely the case at hand: the process has
already exited and the map is the only record of whose exit it was. A racing vendor's
`KindDone` retires its column; a vendor that is not racing keeps the stale-exit guard exactly
as it was. `dropProcess` is deliberately not called on that path — the exit belongs to the
racer, the room's own seat is still running, and forgetting a live process would leave it
running and invisible, which is the state this product refuses. `TestAWarmSeatsRacerRetires-
OnItsOwnExit` was verified to fail without the change, reproducing the hang as
`phase = streaming`; its sibling pins that a non-racing vendor's predecessor exit is still
discarded.
**Verified live the same day.** Race t13, one Claude seat, first dispatch of a cold room:
the column retired in **18 seconds** and drew the whole settled block — `arena
arena/t13/claude`, `seeded 1 file`, `no untracked file matches ".env"`, `1st of 1 · done ·
18s`, `committed 1ef0f00.` and the diff stat. Eighteen seconds against twenty-one minutes is
the measurement. Race t14 then landed the same way with a warm seat behind it, which is the
arm that matters: it is the state the bug needed.
**And it paid the three debts that had been stuck behind it.** `u` on t13 reset the branch to
`ba2d00b` with the stat kept above an `undone` line, and a second press refused as already
undone; the reset was confirmed against git rather than against the room's own claim.
`/adopt claude` on t14 cut `adopt/t14-claude`, ran the `--no-ff` merge, left `main` at
`ba2d00b` and named `gh pr create` — the first live adoption under the shape ruled
2026-08-11, and the debt this section states three paragraphs above its own dated block is
now paid.
**Amendment, 2026-08-16: a racer's turn says which race it was.** `STATE.md` recorded the gap
on 2026-08-15/16: an operator armed `/trace` before a race, and the file held only the
preceding ordinary turn's line. The report named the consequence correctly — grok's
spawn/wait/stream split stayed unmeasured, because the arena is the one turn shape that
dispatches grok.
**The mechanism is a seam, not a dropped record.** A probe reproduced the arena's exact spawn
shape against the real runner — `runner.Start`, a one-shot child, its own `Dir` — and the
record came back: `grok spawn=5ms wait=43ms stream=2ms total=51ms`. So the clock does run for a
racer, on both arena spawn paths. What the record could not say is that it was a racer.
`runner/clock.go` states its own rule at the top: a `TurnClock` is keyed by the seat and the
moment, because those are the only facts that package holds. The race number is not one of
them. It is read off the repo's own arena refs at setup (`arenaRaceNumber`) and it lives on
`turnState`, which is gone by the time the record is written — the runner emits at process
exit, on its own goroutine. So four racers wrote four lines that named neither the race nor
the worktree, and each line was byte-identical in SHAPE to an ordinary turn's line for the
same seat. The trace held the race and could not point at it.
**The fix carries the label the room already has.** `runner.Spec` gains `Race`, the one field
in that struct the runner does not use and only carries. Council stamps it on both arena spawn
paths: the batch seats through `FirstTurn` in dispatch's arena branch, and the merged cursor
seat inside `startEphemeralRacer`, which builds its own spec through `cv.Open`. Both arms are
stamped because a label applied at the obvious call site alone would leave that seat as the one
racer nobody could find. `newClock` takes the race and fixes it for the life of the process,
which is correct by §9.37's founding ruling: every attempt is a FRESH one-shot session, so no
second turn on that process could belong to a different race.
**Nothing is derived, and the ordinary line does not move.** The tag is `arena/t`
(`arenaRaceTag`), minted from the same race number as the branch, so a trace line and the
worktree its attempt is parked on cannot disagree — the vendor is already a column on the line,
and the two together spell `arenaBranch` exactly. The field is APPENDED after `total=`, so
every reader that already parses a trace line keeps its field order; a test pins that position
rather than trusting whoever edits `String` next. An ordinary turn appends nothing at all. That
is §4a.1 rather than terseness: an unmeasured `Span` prints `-` because the stretch existed and
was not measured, but an ordinary turn is not a race whose id went missing — it is not a race,
so there is no field to mark absent, and `race=-` on every ordinary line would invent a
category for the room's normal case.
**What is verified, and what is not.** The council half is pinned at the spec, and that split
is forced rather than chosen: a council test never spawns a vendor (`CLAUDE.md`), so the clock
cannot run there and the spec is the whole of what that package contributes to the record. The
runner half spawns for real and asserts the race survives onto the emitted line, with the
spawn/wait split still measured beside it. `TestArenaSpecsCarryTheRaceIntoTheTrace` was
verified failing before the change, on all four racers and both spawn paths. **The live half is
owed.** No race has been run against this build, so the claim that a real `/trace` now holds an
attributable racer line rests on the probe and the tests, not on a race. One thing this
amendment deliberately does not claim: it does not explain the operator's empty file. The
records are emitted, so the reported absence has some other cause, and the same drive recorded
two candidates beside it — a room that opened on workspace `~`, and `/trace` resolving a
relative path against it. That stays open and belongs to whoever runs the next live race.
**Amendment, 2026-08-17: the worktrees are cut off the render loop, and the room stays a room
while they are.** Until now `arenaSetup` ran inline in `dispatch`, which runs inside `Update` —
so for as long as git took, council drew no frame, read no key and answered nothing. The
operator measured the failure the way these things are always measured, by living in it:
parallel sessions against one repository, a `index.lock` held by another of them, and a room
frozen with **ctrl+c unread** — the one act that could have ended the wait, sitting in a queue
whose only drainer was blocked inside `git worktree add`. A room that cannot be stopped is
worse than a slow one, and this is the same class of defect as §9.37's own give-up amendment:
the room displayed a true thing and offered no act on it.
**The setup is a command now, and the seat order is unchanged.** `arenaSetup` runs on a
goroutine and reports back through a channel (`internal/council/arenasetup.go`); `dispatch`
stops at the point of preparing and returns, and the turn is born later in `applyArenaSetup`,
which calls the extracted `sendTurn` with what the setup measured. **The per-seat
`git worktree add` calls stay SERIAL, and that is a ruling rather than an unfinished
optimisation** — those adds write the repository's own refs and administrative files, so N at
once contend for exactly the lock this change exists to survive, and the parallel version would
turn one slow setup into N racing ones each able to fail the others. Everything stamped when
the turn starts — its clock, its context, the snapshot of the previous replies — is stamped at
the SPAWN rather than at the keypress, which is the honest reading of every duration the turn
then renders.
**The frame names the step and refuses to name the progress.** Each stage reports the words for
what it is about to do — `reading the base commit`, `numbering the race`, `reading
.worktreeinclude`, then `preparing worktree for codex` and `seeding worktree for codex` per
seat — and the footer draws that sentence with the spinner beside it. There is no percentage,
no "2 of 4" and no elapsed figure, because council cannot measure how long a checkout takes and
a number it did not measure is a number it may not draw (§4a.1). The spinner is the second
signal and is liveness, not progress: a step that takes a minute prints the same sentence
throughout, so without a moving cell a working room and a dead one render identically — which
was precisely the old lie. `TestSetupStepsNameTheWorkAndNeverTheProgress` fails on any step
carrying a digit or a `%`, which is the rule stated as a test rather than as an intention.
**The deadline, and the measurement behind it.** The setup carries one context with a **90
second** deadline over the WHOLE of it, enforced through `gitOutCtx` — a context-carrying
sibling of `gitOut` used by the setup path and nowhere else. Every other git call council makes
stays un-deadlined by construction, because `gitOut` takes no context to hand one: a diff read,
a config probe or a commit killed by somebody's guess at a timeout is a worse outcome than a
slow one everywhere the room is not blocking on it. One deadline over the whole setup rather
than one per call, because the number an operator experiences is how long the room was
unusable, and a per-call bound times five seats is a total nobody chose.
The 90 is measured against, not guessed. On the reference Intel Mac (macOS 26.5.2, 2026-08-17),
a five-seat setup against a synthetic repository built to this repo's own shape — 540 files in
60 directories, ~8 MB of content, against telltale's 526 tracked files and 8 MB — ran end to
end in **2.3s cold and 1.3s / 1.4s warm**, worktree adds included. The deadline is therefore
~40x the measured case, and the margin is the decision rather than the number: the failure it
exists for is a lock another session holds, which is unbounded by nature and says nothing about
how large the repository is. What the deadline must never be is tight enough to kill a setup
that would have finished, since a `git worktree add` killed mid-checkout leaves a half-created
tree the operator clears by hand.
**Every ending hands the room back.** A deadline hit or a git refusal ends the setup
WHOLESALE — a clock is a fact about the room's patience, not about a seat, so recording it as
four per-seat skips would blame four vendors for one timer and then race whatever survived as
if the operator had asked for a 1-of-4 race. It lands on the room's existing arena notice,
opened the way a refused race always was, and the sentence leads with the STEP before quoting
git verbatim (`arena: preparing worktree for codex: fatal: …`): the git line names a lock and
not which of eight calls met it, and the step is the half the operator cannot reconstruct. A
process the context killed is never quoted as if git had refused — `cmd.Run` reports the signal
there, and dressing "signal: killed" up as git's own sentence would be §4a.1's bug pointed at a
failure. The brief returns to the composer and the room composes again, so the same enter
retries it. ctrl+c stops the setup in every mode and does NOT quit, which is the keystroke the
freeze ate; the trees already added are **kept and named** rather than swept up, per this
section's founding ruling that worktrees live until the user deletes them, and the next race
numbers itself past them anyway (`arenaRaceNumber`).
Verification note: the mechanics are pinned offline against real temp repositories — the room
drawing and reading keys mid-setup, the step vocabulary, the serial adds (measured, not
assumed: when seat N's add is announced, seat N-1's tree already exists on disk), the
deadline's wholesale stop, the failure handing the room back, ctrl+c, a stopped setup's
messages being dropped by comparison, and the rendered frame. **The live half is owed**: no
race has been run against this build, so the claim that a real held `index.lock` now ends in a
notice instead of a freeze rests on the tests and on an expired-deadline fixture, not on the
lock that started this.
**Amendment, 2026-08-29: the brief also arrives as `AGENTS.md`, for the seats that were
measured reading one.** The candidate (competitor sweep 2026-08-18) proposed AGENTS.md as the
one cross-vendor context channel needing no per-vendor prompt plumbing, on the strength of
agents.md's own "read natively by 20+ tools". That is a docs claim, and ADR-001 does not accept
docs claims about vendor behavior. **The measurement came first, and the build was conditional
on it**: one headless probe per vendor CLI on this box, from a scratch directory whose only
content was an `AGENTS.md` naming a codename nothing else on the machine knew, asked for the
codename and nothing else.
| seat | version | result |
| --- | --- | --- |
| codex | codex-cli 0.149.1 | **answered `ZEPHYR-9`, no tool call** — the file reached the model as context |
| grok | grok 1.0.5 | **answered `ZEPHYR-9`, no tool call**, and named its source on the wire: *"From the always_applied_workspace_rules, the Agents.md file says"* |
| claude | Claude Code 2.1.251 | **answered `ZEPHYR-9` by going to look** — both trials ran `ls -la` then `cat`, recorded in the probe sessions' own transcripts |
| agy | 1.1.25 | **reads it by going to look**, once `--add-dir` names the tree ([§9.59](#s9-59), 2026-09-03): asked only to create a file, it ran `list_dir` then `view_file AGENTS.md` before writing, and then wrote the file the brief in AGENTS.md asked for. The claude row's fact, not the codex row's. The codename probe itself was not run |
| cursor | — | **unmeasured** — this seat races over ACP on a throwaway session, and no probe of that path ran |
Two seats demonstrably ingest the file unprompted, which is the bar the sweep set, so the
feature is built. The claude row is deliberately NOT counted as the same fact: it is a real read
of a real file in the cwd, in a directory holding exactly one file — the easiest possible
discovery — and nothing here claims that seat auto-loads AGENTS.md. Whether the answer was
shaped by the operator's global `CLAUDE.md` load order cannot be separated out by this probe
either; what the transcripts DO show is that the words came from the file on disk, because the
model went and read it before answering.
The ruling that follows from that table is what shapes the feature: **council writes the file
for every racer and claims it for none.** No column, no notice and no snapshot field says a seat
was briefed via `AGENTS.md`, because the room cannot tell per race which seats ingested it — and
two of five are unmeasured. Writing it costs a seat nothing; claiming it would be the room
narrating a fact nobody measured (§4a.1). The file is offered exactly the way the worktree is.
- **Identical for every seat, by construction.** Marker, `arenaConduct`, then the brief — the
same bytes in all five trees. `arenaBriefText` takes no seat parameter at all, so the
per-seat constraint text the candidate pitched is unrepresentable rather than merely
discouraged: it collides with `arenaConduct`'s standing position that the room's added words
are a CONSTANT so the cross-seat comparison stays undisturbed, and a later change that wants
divergence has to argue for it here.
- **The attempt's receipt stays the racer's.** Council's file would otherwise land in the stat
through `git add -N .` — §9.37's own "lying diff" known limit, arriving from the other side.
The three reads that could pick it up (the finish-time `collectArena`, the live
`collectArenaStat`, and `commitArena`'s stage plus its dirty check) append one pathspec,
`:(exclude)AGENTS.md`, measured on git 2.55.0.windows.3: the file stays untracked, stays out
of the commit, and a tree holding nothing but it still reports clean — which is what keeps the
empty-commit ruling working. **`/adopt` therefore merges a branch that never held the file**,
and the operator's repo cannot acquire a stray `AGENTS.md` from a race.
- **The lifecycle verbs read the racer's tree too, and all three reads were wrong until they
carried the same pathspec.** This was found by building the feature, not by reasoning about
it, and each one is a different bug: `/adopt`'s arming read (`lifecycle.go`) would have offered
to adopt a seat that changed nothing, because council's file made a clean tree look dirty;
`/adopt`'s OWN commit — the one it makes for a racer whose work never reached `commitArena`,
which is the give-up path race t9 exercised twice — would have staged the file with `add -A`
and merged it into the operator's repo; and `/arena drop`'s refusal would have named council's
write as the operator's uncommitted work and demanded the `!` spelling on every clean attempt.
- **`/arena drop` takes council's file back before git sees the tree.** `git worktree remove`
counts an untracked file as a dirty worktree and refuses, so the pathspec alone was not
enough: an ordinary drop failed at git. `removeArenaBrief` deletes the file only while the
marker still stands, so a racer's own `AGENTS.md` keeps the refusal it has earned. That is the
ONLY deletion — the worktree is kept until the user drops it, and until then the file is the
visible record of what that seat was told.
- **The exclusion is conditional on the marker, re-read per call.** `arenaBriefArgs` opens the
file and checks it still starts with council's marker. A racer that REPLACED it authored a
file, and it appears in the stat like any other; a file council never wrote is never excluded.
A stale flag recorded at setup would have hidden that authorship.
- **Council never overwrites an `AGENTS.md` the checkout or `.worktreeinclude` seeding already
put in the tree.** In a repo that ships one, no racer gets council's copy, every seat reads
the repository's own instructions identically, and the comparison is as uniform as it was
before this existed. The pathspec is off there too, so a racer's edit to the repo's own file
is in the diff.
- **A write that fails skips that seat**, named on its column through the existing `seatErr`
channel — the `.worktreeinclude` rule, applied for the `.worktreeinclude` reason: a tree the
room KNOWS holds a different brief from its siblings races a different question, and that is
not the comparison the operator opened. Nothing new is rendered for it.
Known limit, stated rather than hidden: while the marker stands, a racer that APPENDS to
council's `AGENTS.md` is excluded from its own diff on that path. A racer editing the room's
brief file is not an answer to the brief, and the alternative — an exclusion that lapses on the
first stray edit — would drop council's own file into every stat instead.
Not built, and not by omission: `telltale doctor` reporting whether a repo carries an
`AGENTS.md` was part of the same candidate. It is a different surface with a different reader
and it is left for its own change.
Verification note, on this section's own terms: the mechanics — the identical file in every
tree, the file never reaching the stat, the patch or the commit, the zero-diff attempt staying a
measured zero with council's file in its tree, the racer-authored file NOT being hidden, the
repository's own file being left alone, the skip on a failed write, the ended-context stop, and
all three lifecycle reads (`/adopt` still refusing a brief-only racer, `/adopt` not merging the
file, `/arena drop` needing no force) — are pinned by offline tests against real temp
repositories (`arenabrief_test.go`), and no test spawns a vendor. **The live half is owed**: the per-vendor probes above were run headlessly in a
scratch directory, not inside a racer worktree during a real `/arena`, so no live race has yet
watched a seat act on this file.
**Amendment, 2026-08-29: `/adopt` says what it is about to merge INTO, before you say y.** The
card named the act — the branch it cuts and the exact `git merge --no-ff` it runs — and named
nothing about the room the merge lands in. Everything the operator needed in order to weigh the
`y` was in a second terminal: how far the racer's branch had drifted from the room, what had
landed in the room since the race was cut, and whether the two had written the same files. So
the answer was "yes because I trust it" or "no because I don't", which is §9.41's finding about
the room's *other* gate, arriving a second time at the one gate that merges.
**The card now leads with measured git state and then names the act.**
```
adopt codex? vs main: 1 ahead, 1 behind · 1 overlapping path (a.txt) · y cuts adopt/t4-codex
and runs git merge --no-ff arena/t4/codex · n cancels
```
- **Every count carries its baseline, and the baseline is the room's own HEAD** — because that
is the commit `/adopt` cuts the adopt branch from, so it is genuinely what the merge lands in.
It is named as the branch when one is checked out (`vs main`), as the short commit on a
detached HEAD, and as `vs the room's HEAD` when git could not answer at all. The clause is
never dropped: a bare `1 ahead` is a number with no question attached. `behind` is the half
the operator had no other way to see, and it is the whole point of the line — it is
everything that landed in the room while the race sat there, including an earlier adoption
from the same race.
- **One `rev-list --left-right --count HEAD...` answers both counts**, and `ahead` is
the same figure `unadoptedCount` was already reading for the zero-change refusal, so the
preview costs the card one git call rather than two. Measured at git 2.55.0.windows.3: the
left count is what only HEAD holds and the right is what only the branch holds.
- **"Overlap" is a read; "conflict" would be a claim.** The overlapping paths are the
intersection of two `diff --name-only` reads over the same merge base — `HEAD...` is
the incoming half git actually applies, `...HEAD` is the room's own half — so the
card states that both sides wrote a path and stops there. A repository can overlap on a path
and merge cleanly. The word "conflict" belongs to a merge that ran, and the reactive path
below still owns it.
- **Three overlap states, kept apart (§4a.1).** A read that returned nothing renders `no
overlapping path`; a read that returned paths renders the count and names the first; a read
that failed renders `the overlap check could not run:` with git's own line. An unreadable ref
never renders as a clean one. The counts and the overlap fail differently on purpose: the
counts are load-bearing, so a failed read refuses the whole command by name, exactly as the
older `unadoptedCount` call did; the overlap is advisory, so a failed read degrades to its own
sentence and the card still arms. A broken preview must not brick a verb (`arenaRaceNumber`'s
rule).
- **The preview states its own limit rather than leaving it to be discovered.** Every figure is
read off COMMITTED state, so a racer whose worktree is still dirty has work none of the
figures cover — and that card adds `these counts exclude 1 uncommitted path` beside the
clause that already says `y commits its worktree`. `TestAdoptConflictAbortsCleanly` is exactly
that case: an uncommitted racer edit conflicts against a room commit while the overlap read
correctly reports nothing shared. Without the clause, `no overlapping path` would be read as a
promise about the merge.
**Two shapes recorded rather than taken.**
- **`git merge-tree --write-tree`, which computes a REAL merge result.** It would let the card
say "conflict" honestly. It is not here because the claim would need a live measurement at a
pinned version on this box before it could ship (this section's own rule), it needs git ≥2.38,
and it writes objects into the repository — which puts a preview on the write side of a room
whose posture is offer, never take. The reactive abort already owns the real merge result, and
it owns it after the operator asked for one.
- **Folding the racer's uncommitted paths into the overlap set**, by parsing
`git status --porcelain`. It would close the limit named above, and it was declined because
those paths are a prediction of a commit nobody has made yet — council reading a tree to guess
what a future commit will contain, where §4a.1 asks it to read what exists. The exclusion
clause states the gap instead.
**The preview leads the line, and the cost of that is stated.** The notice truncates from the
right at a narrow width, so leading with the measured state can cost the action clause its tail
— and the action clause is the older contract. It leads anyway: an operator who can read only
the first clause can still press `n`, and the preview is what makes that `n` a decision rather
than a mood. The alternative, recorded and not taken, is a second sheddable cell on the status
line (the mechanism `st.ArenaSetup` already uses), which would drop the preview whole instead of
slicing it — a new render surface for one notice, in a file this change otherwise does not
touch.
Verification, on this section's own terms: the mechanics are pinned by offline tests against
real temp repositories (`lifecycle_test.go`) — the counts against an unmoved room and against
one that moved, the named overlapping path, the uncommitted exclusion, the two overlap failure
states held apart, and the baseline on a named branch, on a detached HEAD and on no answer at
all. No golden moved, because the card is a notice string and no golden renders one. **The live
half is owed**: no real `/adopt` has been armed against this build, so every sentence above
rests on the fixtures rather than on a race.
**Amendment, 2026-08-29: `/adopt` can take the winner plus the parts of the runner-up you point
at.** A race ends with four attempts and one decision, and the decision the room offered was
all-or-nothing: adopt one seat whole, and retype by hand whatever the runner-up got right. The ask
is not ours — it is the one users put to Cursor's own multi-agent judging thread, in those words:
synthesize a best-of-both instead of picking one wholesale. No surveyed tool ships it. Council
already owns the substrate — per-attempt worktrees, one base SHA, commit receipts, and a y/n card
that names exact git commands — so this is a grammar and four refusals, not new machinery.
**The grammar is one more argument, and the fork is PER-PATH.**
```
/adopt claude +codex internal/council/helper.go docs/council.md
```
- **Per-path, not per-hunk.** The sweep's own evidence is users asking to mix at path OR hunk
level, so the choice was open. Per-hunk needs an interactive picker inside the room — a
full-frame body with its own scroll, its own keys and its own mode word — which is a new render
surface for a v1 whose value is that the operator can take one file from the runner-up. A path
is also the unit already in front of them: the column's `git diff --stat`, and this card's own
overlap clause, both speak in paths. **Per-hunk is deferred, not rejected**, and this shape does
not block it: a hunk picker would narrow what `+` contributes and leave the grammar alone.
- **`+` glued to the donor seat.** A bare `+` as its own word would make `/adopt claude + codex`
legal, and that reads as a request for two whole attempts — which this verb cannot do and must
not appear to offer. `/adopt` already takes its whole argument (roomcmd's `parseCommand`), so the
longer form needs none of the vocabulary handling `/arena drop` needed.
- **User-typed, never computed.** §9.34 rejected a synthesis hop, and that ruling binds here:
council applies the paths the operator named and chooses nothing. The refusals below are how that
promise is kept mechanically rather than by intention.
**Four refusals, and not one of them resolves anything.** Each names the path and a way forward
(§9.17's tell), and each fires before the card arms, so a `y` is always one that can be honored:
- **A path BOTH racers wrote.** This is the founding refusal. `git checkout -- `
would discard the base attempt's answer with no merge and no conflict marker, so council refuses
by name and the operator decides — drop the path, or merge it by hand afterwards.
- **A path the ROOM wrote since the race was cut.** The same silent clobber one level out: the
merge machinery never sees a path taken by checkout, so the room's own work there would vanish.
- **A path the donor did not write.** Taking it would land the base attempt's own content under a
receipt saying it came from the donor.
- **A path the donor deleted.** A hybrid takes files a racer wrote, never a deletion — a stated v1
limit rather than a `git checkout` pathspec error discovered after the branch was already cut.
**The card says exactly what will be merged from where, composed with the divergence preview
above.** The preview still leads, for that amendment's reason; the leading question gains the
hybrid's own scope, and the action clause gains its second half:
```
adopt claude + 1 path from codex? vs main: 1 ahead, 0 behind · no overlapping path · y commits
both worktrees, cuts adopt/t4-claude+codex and runs git merge --no-ff arena/t4/claude, then
takes helper.go from arena/t4/codex · n cancels
```
Every path is named rather than counted-with-an-example. The count-plus-first grammar the overlap
clause uses is right for a measurement the room took; these paths are the SCOPE the `y`
authorizes, and a card that authorized "2 paths (a.txt)" would leave the second one unread. The
operator typed them, so the list is short by construction.
**The branch name carries both seats: `adopt/t-+`.** This is a naming decision and
the arena record (§9.47) forced it, because that page derives everything it knows from these refs.
The alternative — keep `adopt/t-` and let the commit message carry the donor — would leave
one seat's name alone on a branch holding another seat's work, in the one place `git branch` shows
a reader, and the record would then count the base seat as having won the race outright. `+` is the
joiner because it is legal in a ref name, because `-` is already the collision suffix and because
`/` is already the arena namespace. `freeAdoptBranch` suffixes a collider identically, from the
same single scan, so the two spellings cannot disagree about what "taken" means.
**The receipt names both sources.** The base arrives as the unchanged `git merge --no-ff`, and the
paths arrive in a second commit whose message names both arena branches, lists every path, and says
what council refused to do:
```
adopt race t4: arena/t4/claude whole, plus 1 path from arena/t4/codex
the merge below this commit carries arena/t4/claude whole. this commit
adds the paths that came from arena/t4/codex, and it adds nothing else:
helper.go
telltale council took no path that both seats wrote. a shared path is
refused by name, and the operator merges it.
```
The notice says it a third time, because that is the last moment the operator is still looking:
`adopted claude onto adopt/t4-claude+codex, with 1 path from codex (helper.go)`.
**The arena record renders a hybrid as its OWN state, and credits nobody.** The refs can say a race
was decided and which two seats the adoption was cut from; they cannot say which paths came from
where, because that lives in a commit message the page does not read. So a hybrid raises a fourth
per-seat count and moves no rate at all: the race counts as decided, both seats count as having
entered it, and neither seat's `adopted of decided` moves. Crediting the base seat would score it
for work the donor wrote; counting it against both would score two seats down for a race the
operator resolved in both their favour. A seat whose only decided races were hybrids reads
`no attempt adopted whole part of 2 hybrid adopts`, which is the true statement — and the window
sentence carries the difference a reader adding the seats up would otherwise not find:
`3 decided by you (2 by a hybrid adopt, counted for no seat)`. A whole adoption of the same race
still outranks a hybrid of it, because an operator who adopted whole, reverted and then took a
hybrid did adopt it whole once, and both receipts survive.
**One fork from the divergence-preview ruling above, recorded because it is a fork.** That ruling
declined to fold a racer's uncommitted paths into the OVERLAP set, on the grounds that they are a
prediction of a commit nobody has made. The hybrid's path checks DO read them, and the difference is
what the read is for. There it was a preview of a merge RESULT; here it decides a refusal about
paths the operator named, and `y` commits both worktrees in the same act with `git add -A` — so the
set is `tracked changes ∪ untracked-and-not-ignored`, which is what that commit will contain by
definition rather than by forecast. Refusing to read them would refuse every hybrid on an ordinary
race, because arena seats leave their work uncommitted (commit-per-turn is deferred). The card still
states the older ruling's limit, in the same clause it already used.
**A conflicted hybrid restores exactly like the conservative whole adopt.** The base merge is
unchanged, so a conflict aborts, the branch is deleted, the room goes back to the branch it came
from, and the donor's paths are never written. A failure in the second half restores the same way,
with one extra step named rather than hidden: a `git reset --hard` before the checkout back. It is
bounded to a branch council cut, at a commit council made, over a room tree measured clean before
any of it — the only content it can discard is a half-checked-out copy of files that exist whole on
the donor's own branch.
**Two shapes recorded rather than taken.** A `+` with no paths, meaning "take everything of
the donor's that does not collide" — declined because it makes the scope a thing council computed
rather than a thing the operator read, which is the whole contract of the card. And renaming the
donor's paths on the way in — declined as a second grammar to learn, for a case `git mv` already
handles after the adoption.
Verification, on this section's own terms: the mechanics are pinned by offline tests against real
temp repositories (`hybrid_test.go`) — the merge plus the named path landing while the donor's other
file does not, the receipt naming both branches, all four refusals, the grammar's own refusals, the
collision suffix, a committed donor read off its branch instead of its tree, and a conflicted hybrid
restoring the room. The record half is pinned in `record_test.go` against ref lists, with its own
golden in both glyph sets. **The live half is owed**: no real hybrid has been armed against a race on
this box, and it is owed on the same keystroke as the divergence preview's live debt above.
### 9.38 paste lands whole, and never sends (2026-08-09)
The ask, in the operator's words: *"how i can paste things into the area i can type in."* The
answer required measuring what a paste even was in this room, because the two obvious guesses —
it works, or it fires a send per pasted line — were both wrong.
**What was measured about today's behaviour.** All of it read off the pinned module source, not
vendor docs. bubbletea v2.0.8 enables bracketed paste unless a view opts out
(`cursed_renderer.go` writes `SetModeBracketedPaste`; council's `View()` never sets
`DisableBracketedPasteMode`), and ultraviolet's terminal reader buffers everything between the
paste markers into ONE `PasteEvent` — a newline inside the paste lands as `\n` in its content,
never as an Enter keypress, and the win32-input-mode encoding Windows Terminal uses is decoded
into the same buffer (`terminal_reader.go`, ultraviolet pinned at v0.0.0-20260703014108). So in
a bracketed-paste terminal a paste could never have fired a send. What it did instead was
NOTHING: council's `Update` had no `PasteMsg` case, the message fell through the type switch,
and the clipboard's offer was silently discarded. The composer never learned a paste happened.
The fires-sends failure is real on exactly one path: a terminal with NO bracketed paste replays
a paste as keystrokes, each pasted newline arrives as an Enter keypress, and compose mode's
enter dispatches — a five-line paste is up to five turns, each to live vendor CLIs. Council
cannot distinguish that replay from typing without a timing heuristic, which would be inferred
behaviour, and this product does not ship inferred behaviour (§4a.1). So that path is left as
it is and named here instead: the text chunks are flattened safely by `sanitizeKeepingSpace`,
the enters are enters, and the fix on such a terminal is the terminal. Windows Terminal — the
reference environment — brackets its pastes.
**What was built** (`paste.go`): the room's half of the contract the runtime already offers.
- **One `PasteMsg` case in `Update`.** The content goes into the draft and nowhere else — a
paste never dispatches, never answers a gate, never quits. Enter, a keystroke from a person,
remains the only way a brief leaves the room. Pasted control characters cannot act: a pasted
`\x03` is not ctrl+c, a pasted `q` is the letter q; controls without width are dropped.
- **The multiline ruling: newlines are PRESERVED, raw.** The composer has been a block since
ctrl+j existed (`State.Draft` may hold newlines; `wrap()` honours them; the compose area
grows to `maxComposerRows` and elides with "N more above"), so there is no single-line
prompt to protect and no need for a `⏎` display glyph — a pasted paragraph renders as the
rows it is, in both glyph sets, and dispatch hands the vendors the draft with its real
newlines. The string on screen is the string sent (§7.14). CRLF collapses to `\n` (the
Windows clipboard's line ending; splitting it would gift every line a trailing space). The
one lossy rewrite is stated rather than hidden: a tab becomes one space, because a cell grid
cannot budget a tab and a guessed tab stop would be fidelity theatre.
- **A cap with a named refusal.** `maxPasteRunes` (8,192, over draft-plus-paste) refuses
atomically — nothing lands, not a truncated prefix — and the notice carries both numbers and
the remedy: *"paste refused: 20481 chars against the composer's 8192 — put long text in a
file and name the path in the brief."* The number is anchored to the narrowest pipe a brief
must fit through (the Antigravity seat's prompt rides argv; Windows caps a command line at
32,767 UTF-16 units; 8,192 runes is at most half that even all-surrogate-pair) and to the
point where a footer composer that deletes rune-by-rune stops being an editor.
- **A paste from view mode inserts and opens compose.** A paste is not a keystroke — view
mode's letters are commands because they are keys; pasted text can only be material, and the
only place material goes is the draft. The mode line states the switch on the next frame.
The exception is a pending y/n (tool gate, `c`, `/write`, a flow write hop): the paste is
refused by name and the question stays exactly where it was — nothing about a pending
request happens implicitly.
No golden changed and none was added: a pasted draft produces the same `State` shape ctrl+j
already produces, and a golden that did not change is the claim that the room's appearance did
not either. The tests (`paste_test.go`) drive `Update` with the real message shape and assert
the observables the flow security tests trust — spawn count, draft, pending flags — end to end
through enter, which must deliver the pasted newlines to the seat intact.
**The live verification owed.** No test in this container can observe Windows Terminal
bracketing a paste — that is the terminal's half of the contract. The check, one minute at the
real machine: open `telltale council` in Windows Terminal, copy a three-line snippet, paste
into the room. Expected: one insertion, three rows in the composer, zero dispatches; then enter
sends it as one brief. If the paste instead lands as separate turns, the terminal did not
bracket it — record the terminal build in PARITY.md, because that is a measured vendor fact,
not a council bug.
**Paid, 2026-08-13, on Windows Terminal 1.24.11911.0.** A three-line snippet pasted into a
live room landed as ONE insertion, three composer rows, and **zero dispatches**; `ctrl+u` then
cleared it and reported the count, which is the second half of the same gesture and is why the
clear is evidence too — a paste that had dispatched would have left nothing to clear. The
terminal build is named because the bracketing is the terminal's half of the contract and a
version is the only thing that claim can be pinned to; PARITY.md stays out of it, since that
file records a machine BEHAVING DIFFERENTLY and this machine behaved as specified. **One half
of the check is still unexercised**, stated rather than rounded up: the draft was cleared
instead of sent, so "enter sends it as one brief" has not been observed on a pasted multi-line
draft. The property that was owed — a paste never sends — is the one that was measured.
**The other half paid, 2026-08-15/16, and it found a render bug on the way.** A three-line
brief was pasted and SENT with enter in a live 5/5 room: ONE dispatch, the turn counter
moved by one, and the newlines reached the seat as bytes — verified in the vendor's own
session transcript (`now.\nThen add a second one…`), not read off the wrapped column echo,
because the echo's row breaks are ambiguous between newlines and word wrap. **What did not
hold was the composer render: the three-line draft drew as ONE row before enter.** The
2026-08-13 check drew three rows in a one-seat room on the same terminal build
(1.24.11911.0), so the suspect is the composer's height behavior under a full seat strip,
not the paste path — the wire is proven right and the drawing is proven wrong, which is
the exact split this section exists to keep. The render defect is recorded as an unowned
gap in STATE.md; this section's claim is amended to say the paste property holds ON THE
WIRE, with the row rendering owned separately.
**Amendment, 2026-08-16 — the render was measured, and the seat-count suspect is REFUTED.**
The bullet above named a suspect: the composer's height behavior under a full seat strip. That
suspect is wrong. `TestAMultilineDraftNeverCollapsesSilently` sweeps the drawn geometry with a
three-line draft and asserts that every line is on screen, or that the frame says how many are
not. The sweep covers one, three and five seats, widths from `MinWidth` to 240, heights from
`MinHeight` to 40, both glyph sets, and the expanded projection. All 588 combinations draw all
three rows. `composer-multirow-five-seats.txt` is the 5/5 room at 200x24 as bytes. **The seat
count changes nothing about the compose area**, and the reason is structural: the composer is
full-width chrome, so `composerRows` wraps against `promptWidth(st.Width)` and never against a
column width. Five columns narrow the COLUMNS. They do not narrow the composer.
Two further things were measured, because a refutation is only useful when it also closes the
paths it rules out.
- **The one-row compose area is the only silent collapse in the render, and it is unreachable.**
`composerLines` flattens newlines to spaces when `lay.Prompt == 1`, with no marker. Its comment
calls that unreachable for a real draft. The claim holds. `resolveLayoutIn` clamps `Prompt` to
1 only below a height of 9, and `Render` refuses to draw a room under 60x10 at all — it prints
`council needs 60x10 (have WxH)` and no compose area. So no drawn frame reaches the flatten.
The branch stays as it is, and this paragraph is the measurement that says why.
**This bullet is WRONG and the 2026-08-17 amendment below replaces it.** It is kept as
written because the amendment is about how a sweep can prove the wrong thing, and the
sentence that did it is the evidence.
- **The paste transport carries the newlines, on this platform's own path.** ultraviolet, pinned
at v0.0.0-20260703014108, buffers a bracketed paste into ONE `PasteEvent`. The win32-input-mode
path Windows Terminal uses converts a pasted Enter record to a literal `\n` inside that same
buffer (`terminal_reader.go`, the `isWin32 && event.Code == KeyEnter` case) rather than
emitting a key event. bubbletea v2 maps that event to `tea.PasteMsg`, which `Update` routes to
`paste()`. `sanitizePaste` keeps the newlines and `setDraft` stores them.
`TestAWindowsPasteKeepsItsLinesAndLosesItsCRs` already drives `Update` with LF, CRLF and
bare-CR content and pins the draft that results, so this half needed no new test.
**What is left, stated rather than rounded up.** The live observation is not explained. The draft
held newlines on the wire, the render draws newlines at every geometry it will draw, and the two
cannot both be true of one frame. The leading remaining hypothesis is an observation error at the
composer, and §9.38 already names the trap that produces one: the echo's row breaks are ambiguous
between a newline and a word wrap. That warning was applied to the column echo and the wire was
checked against the transcript because of it. The composer's own row count was still read by eye.
**The decisive next measurement is cheap: reproduce the paste in a live 5/5 room and record the
terminal's rows and columns with it.** A geometry at or above 60x10 makes the render correct by
the sweep above, which moves the defect off this section entirely. A geometry below it means the
room drew the floor refusal, and the report is about a different frame than the one assumed.
**Amendment, 2026-08-17 — "correct at 60x10 by the sweep" was not true, and the sweep is why.**
The sentence above says a geometry at or above 60x10 makes the render correct. It does not, and
neither does the bullet two paragraphs up that calls the one-row flatten unreachable. Both rest
on the same sweep, and the sweep could not see the cell they were claiming.
**What the sweep pinned false.** `TestAMultilineDraftNeverCollapsesSilently` built every one of
its 588 cells from `fiveSeats()`, and that fixture carries no pending gate and no collapsed seat.
Those are the room's two chrome rows, and `resolveLayoutIn` spends both out of the SAME budget
the compose area is measured from — the needs-you strip (§9.40) first, because it does not yield,
then the collapsed-seat notice. Pinning both absent held the compose area two rows taller than a
real room's, in all 588 cells at once. The sweep was not measuring a geometry the frame reaches.
It was measuring a fixture, and the fixture was generous in exactly the dimension under test.
**The cell it could not see, and it is an ordinary room.** At the 60x10 floor with a tab bar, both
chrome rows on screen, `resolveLayoutIn` leaves the compose area one row: `Prompt` is clamped to
`Height - rows - 1`, which is 1, while the draft still holds three. Nothing exotic builds that
room — a five-seat machine with one vendor not installed, at the smallest terminal council agrees
to draw, the moment a gate goes up. **Forty cells of the widened sweep land there**: heights at
`MinHeight`, three and five seats, every width from 60 to 240 once `Expanded` forces the tabs
tier, both glyph sets. In every one of them the old code flattened the three typed rows into one
row of prose joined by spaces, with no marker and, at that width, nothing even elided.
**And a flatten is worse than a clip, which is why no count could have fixed it alone.** Clipping
drops text and the marker vocabulary describes exactly that. Flattening dropped nothing — it
RESTATED the draft, and §7.14's promise is that the string on screen is the string sent. Three
typed lines and one long line are different briefs; the room drew the second and the wire sent the
first. So the widened sweep asserts two things now, not one: no draft line is dropped without the
frame saying how many, and **no two typed rows are ever drawn welded into one**. The weld check is
what fails on the old code; the drop check passes there, because at 60 cells the flattened draft
still fit.
**What the marker shows.** `composerLines` no longer flattens. When the compose area is one row and
the draft wants more, the row carries the marker and the TAIL, in that order and at two
intensities: `↑ 2 more above third line_`. The words and the count are `moreAbove`, which is the
column overflow marker's own spelling (`overflowMarker`, §9.10) and the one the multi-row composer
already spent a whole row on — a reader who learned `↑ 36 more above` on a column is not taught a
second vocabulary here. The count is rows not drawn, so it agrees with the multi-row path's own
arithmetic. The tail stays because that is where the cursor is, elided from the left when the
remaining width will not hold it, which is the rule this branch was already built on. The
separator is three cells rather than the room's `│`: the composer is a box and its sides are that
same glyph (§9.44), so a bar inside it would read as a column rail through the frame's own edge.
Three cells is `needsYouGap`'s answer to the same question.
Nothing about the frame's HEIGHT changed, deliberately. A floor of two rows on the compose area
would have closed the same gap, and it would have moved every golden in the package and taken a
body row from a room already at its floor — a room you can type in but not read is the trade
§9.44's budget refuses. The marker costs nothing but the row it was already drawing.
The goldens are `composer-clipped-to-one-row.txt` and its `-ascii` partner, at 60x10, so the claim
is bytes in both glyph sets; they render `PlainStyles`, which is the NO_COLOR half. No existing
golden moved — a draft that fits its compose area reaches none of this, and that is every room
above the floor.
**Amendment, 2026-08-17 — the decisive measurement was taken, and the composer drew THREE
rows.** "What is left, stated rather than rounded up" above named one cheap next step:
reproduce the paste in a live 5/5 room and record the terminal's rows and columns with it.
The operator ran it. A three-line brief went into a live 5-of-5 room in Windows Terminal at
a recorded geometry of **170 columns x 54 rows**, and the composer drew **three rows**. The
geometry was recorded this time rather than read back off the frame, which is the whole
reason the step was specified that way.
**So the live render agrees with the sweep (#253), and the t9 one-row observation stands
UNREPRODUCED.** The render is now verified twice by two different instruments — 588
synthesized geometries and one live room — and the wire was never in doubt. Against that,
one unrepeated eyeball reading. The finding is recorded as a **probable observation error**,
and this section already documents the trap that produces one: the echo's row breaks are
ambiguous between a newline and a word wrap. That warning was applied to the column echo at
the time and the wire was checked against the vendor's transcript because of it. The
composer's own row count was the one thing still read by eye, and it is the one thing that
did not reproduce.
**What this run does NOT establish, because the geometry is the generous kind.** 170x54 sits
inside the swept band on width (60 to 240) and ABOVE it on height (`MinHeight` to 40), and
above means MORE compose room, never less. It is nowhere near the 60x10 floor where the
amendment above found forty cells that used to flatten. At 170x54 the pre-marker code and
the marker code draw the same three rows, so this run cannot tell them apart and is not
evidence about either. It closes the reported defect and it closes nothing else. The
one-row branch is still exercised only by `composer-clipped-to-one-row.txt` and its ascii
partner, which is where that claim belongs.
**One independent corroboration of the column figure, and only the column figure.** The agy
statusline capture taken on the same machine the same evening (§3.8's re-capture block)
carried `terminal_width: 170` on all fifteen fires. That is a second instrument reading the
same terminal width, which is worth one sentence because the geometry is the evidence here.
It says nothing about the row count, which no payload carries.
**Amendment, 2026-08-09 — the ergonomic other half: ctrl+u clears the draft.** Paste changed
the arithmetic on regret. A draft used to cost at most a typed sentence, so backspace's
rune-at-a-time delete was proportionate to any mistake the composer could hold; one wrong
paste is now up to 8,192 runes in a single gesture, and 8,192 backspaces is not an editor.
`ctrl+u` — readline's own kill-line — empties the composer in one keystroke, in compose mode
only. A chord on purpose: `sanitizePaste` drops every control character, so no paste can
carry the key into the room, and no stray letter can fire it — the one gesture that can empty
a draft is a deliberate hand on ctrl, the same argument that keeps a pasted `\x03` from
cancelling. No y/n gate, unlike `c` and `u`, whose confirms price drops nothing can reverse:
a cleared draft's ways back are ordinary (the clipboard still holds a paste; a sentence
re-types), so the loss is *stated* instead of priced — the notice carries the measured rune
count of the string just dropped ("draft cleared — 1204 chars"), in the paste refusal's own
unit and spelling, never an estimate (§4a.1). An empty draft clears silently — backspace's
own empty-draft behaviour applied at size: after the press the state the key promises is
already on screen, and an every-press "nothing to clear" would put noise where dispatch
answers land. The key sits below every pending gate in `key()`'s routing, so a stray ctrl+u
under a y/n gets that gate's standing stray-key answer (cancel, or the question restated) and
never reaches the draft; in view mode it does nothing at all, because esc parked the draft
there under the promise "keeping the draft", and a chord that revoked it from the other mode
would make esc unsafe in hindsight. Nothing but the draft moves — the routing indicator falls
with it only because it describes it. Taught on the help panel's compose-keys row
(`ctrl+j/u/esc`, landed inside that row's 114-cell budget; the words "compose" and "the
draft" paid), deliberately not on the compose mode line, which is at its own width budget.
Tests: `draftclear_test.go` drives `Update` with the real chord through every gate and both
modes.
### 9.39 a fifth seat, and the first that reports money (2026-08-09)
The ask, in the operator's words: *"grok 30 dollar subscription paid for. create a seat for
grok in council."* The seat is built and the invocation is verified end to end; what follows is
what had to be measured to earn each claim on the column, including two flags that were refuted
and one hazard the seat ships with because nothing in the CLI can close it.
**The vendor.** grok 1.0.0 (3cd0d0cbce), signed in against grok.com — the subscription, not an
API key — with `grok models` reporting one model, `grok-4.5`. It is a spawn-per-turn batch
program like Codex and Antigravity, not a live process like the Cursor seat: `-p/--single`
takes a prompt, answers, and exits. So it implements `Vendor` and neither `Persistent` nor
`Conversational`, and its column takes the same spawn-per-turn shape those two already have.
**The invocation, and what it deliberately does not carry.** `--output-format streaming-json`
and nothing else, with the prompt as the value of a trailing `-p`. argv rather than stdin, and
unlike Codex that is not a preference — there is no `-` sentinel and no stdin channel for a
prompt at all. It is safe for the reason the Antigravity seat's argv transport is safe: the
installer drops a native `grok.exe`, not a `.cmd`, so `classify()` never reaches the shim
refusal and no `cmd.exe` ever sees a brief.
Two flags whose names promise containment were probed, and NEITHER is passed:
- **`--permission-mode plan` was refuted, with the write landing.** Asked to create a file under
it, the seat called its `write` tool, reported the call `completed`, said so in prose, and the
file was on disk afterwards. The control run without the flag wrote its file too, via
`search_replace`. The only difference observed between the arms was which write tool the model
picked. This is the Antigravity ledger a second time (ADR-008, seventeenth amendment) and it
lands the same way.
- **`--sandbox` is worse than refuted — it is unobservable.** `grok --sandbox bogus-profile-xyz
-p "hi"` does not error, does not warn, and answers normally with exit 0. A flag that silently
accepts a profile name that cannot exist gives council no way to tell a real profile from a
typo, so asking for one would put a word in the badge backed by a value the CLI may never have
read. Feeding a vendor a deliberately INVALID value, rather than a plausible one, is what
turned "unverified" into "unobservable" here, and it is the cheapest probe in this document.
So the badge is `unsandboxed`, and its detail says both of those things in the vendor's own
terms. Also deliberately absent: `--always-approve` and `--permission-mode dontAsk /
bypassPermissions`. The default headless mode already writes without asking — measured, above —
so an approve-everything flag would buy nothing in exchange for the badge cost ADR-008's fifth
and seventh amendments attach to that whole class.
**The first seat that reports money, and why that is the honesty rule working rather than
bending.** grok's `end` event carries a `total_cost_usd` it computed itself (`0.0407676` on the
first captured turn, with a `modelUsage` breakdown beside it). Codex and Antigravity report
token counts and no dollar figure, which is why both adapters leave `CostUSD` nil forever — a
cost derived from tokens and a remembered price is exactly the invented number section 4a.1
forbids. The constraint was never "council does not show cost"; it was "council does not INVENT
cost". This figure is read, so it passes through untouched, and
`TestGrokEndCarriesThreadAndTheVendorsOwnCost` asserts the exact captured value so that any
future rounding or unit conversion fails there.
The pointer matters and is not defensive: a captured turn ends with no `usage` and no cost keys
at all (see the slash hazard below), so "reported nothing" and "reported zero" are both real
states of this field on this vendor. `TestGrokAbsentCostStaysAbsent` pins it.
**Streaming, and the one judgement call.** The deltas are genuinely token-level — `"I'll"`,
`" read"`, `"notes"`, `".txt"`. That is finer than the ~80-character chunks section 9.7 flagged
as overstating "tokens" on the Claude seat, and finer than the ~95-character ACP chunks section
9.36 measured on the Cursor seat, so this column carries `GranTokens` on the strongest evidence
in the room.
The judgement call is `thought`, which is dropped. It is the model reasoning rather than
answering — 46 lines against 14 of `text` on the first capture, opening "The user wants me to
read notes.txt" — and routing it to the column would put private deliberation where the answer
goes, in a room built to compare answers. It is the line `codex.go` already draws when it
excludes `reasoning` items. The cost is stated rather than hidden: a turn that thinks for a long
time before speaking shows an empty column while it thinks.
**A hazard that ships, because no invocation closes it.** A brief whose first non-space
character is `/` is eaten by grok's own slash-command parser and never reaches the model. The
turn is not an error — `available_commands`, then an `end` with no usage, no cost and no text,
exit 0. On screen: a column that finishes instantly with nothing in it.
Three channels were tried and all three were eaten: `-p "/context"`, `--verbatim -p "/context"`
(whose help text reads "Send the prompt exactly as given"), and `--prompt-json` with the text as
a content block. The third had a CONTROL — the same `--prompt-json` invocation with a non-slash
prompt answered normally — which is what makes this a property of the parser rather than a guess
about a flag that might not work.
The room mostly protects this seat already, and by accident rather than by design: section 9.31
refuses to dispatch any draft whose first character is a slash, so nothing spawns and nothing is
billed. What does NOT hold is that refusal's documented escape hatch. A user who genuinely means
a leading slash types one leading SPACE, and the space survives the composer, the parse and the
dispatch untouched — and then grok trims it and eats the slash anyway (measured). So the escape
hatch reaches four seats and not the fifth.
Nothing is rewritten to compensate, and that is the decision rather than an omission. Editing a
brief on the way to ONE vendor would make five columns answer different questions while the room
claims they answered one, and the room's whole premise is that the seats got the same brief. A
blank column is the lesser failure, and it is documented in the file whose column shows the
symptom.
**The hue argument section 9.28 said would be owed.** That section closed by saying a fifth
vendor would have to argue for its colour, and here is the argument. Grok is `14`, bright cyan —
the twin of Codex's `6`, so section 9.28's honest weakness is now TWO twinned pairs rather than
one. That is forced, not chosen: after 4/5/6/12 the legal set holds only 13 and 14, both twins
of a seat already seated. The only real decision was which seat to pair with, and it went to
Codex because silence routes to Claude alone — the Claude column is on screen in nearly every
room, so keeping magenta unshared protects the seat a reader sees most. `CX` versus `GR` carries
the distinction when a scheme renders 6 and 14 close, which is what the two-letter tags are for.
Worth stating plainly for whoever adds a sixth: **the legal set is now full.** 13 is the last
free index, and after it a new seat cannot have a hue of its own without taking a severity
(making a seat read as failed) or abandoning 4-bit indices (council asserting a colour over the
user's own scheme). `TestSeatHuesAreExhaustive` fails with that sentence in it. The tag is what
scales; the hue was always going to run out.
**What is verified, and what is not.** The parser is pinned against captured lines, and the
INVOCATION is verified separately by `grok_live_test.go` (`-tags=live`) — because unit tests
over captures would still pass if the argv this adapter builds were rejected outright by the
CLI, which is the ADR-008 failure mode in its purest form. That test ran: the first turn
returned the exact expected text, a session id and a positive cost; the resume turn came back on
the SAME session id and recalled its own first answer, which is what distinguishes a real resume
from a re-send. Not verified: anything on macOS — this is a Windows measurement, and the Mac's
grok is untouched (`PARITY.md`). ~~Not built: an `internal/adapter/grok`, so no HUD row carries
this vendor. Council can drive this seat; the gauges cannot yet see it.~~
**Amended 2026-08-11: it was built, and this paragraph went on saying otherwise.**
`internal/adapter/grok` landed in PR #183, on the live survey §3.9a records (2026-08-09),
so grok sessions render as
HUD rows with name, model, workspace, a vendor-REPORTED context percentage and last
activity. The gauges now observe this vendor as well as drive it, and the seat's parser
and the adapter read the same wire from two sides. One thing stays dropped and it is not
an oversight: grok writes a per-turn dollar figure and no session total anywhere, so the
last turn's cost reaches the detail pane as a labeled Extra and never the `COST` column
(§3.9a).
**Fleet guard wiring for Grok under ADR-012.** agent-ops ADR-012 rules that guard
wiring, not lane shape, is the control on every vendor — and grok is the fifth vendor seat in the room.
~~To fulfill the ADR-012 guard obligation for Grok, Grok's `PreToolUse` fleet guard is configured by setting
`[compat.claude] hooks = true` in `~/.grok/config.toml` (which routes Grok tool calls through the fleet's
Claude-compatible `PreToolUse` credential guard wrapper `pre_tool_use_credential_guard.py` / `hooks.json`), or by
registering the `PreToolUse` hook script via `grok hooks-add`. This ensures secret stores, published history,
and dangerous mutations are screened by the `PreToolUse` guard across all five seated vendors (Claude, Codex,
Antigravity, Cursor, and Grok) without restricting lane capabilities.~~
**Corrected 2026-08-11: the struck sentences were instructions, and none of it is wired.**
Two things were wrong at once. First the reading: `[compat.claude] hooks` is `false` on
this box, and grok's native hook system (`grok hooks-add` / `grok hooks-trust`) has
nothing installed into it, so **the seat has no fleet guards wired today** — the struck
text described a configuration that does not exist and stated the screening as a fact.
`STATE.md` carries the same measurement. Second the shape: this file records what was
measured about a vendor, and those sentences told a reader how to configure a machine. A
design record is not a runbook, and a runbook here would go stale silently on a box
nobody re-measured.
What replaces them is the obligation, recorded and routed rather than acted on. Under
agent-ops ADR-012 an unwired vendor is an **open obligation on the FLEET**, never a
reason to avoid the seat or to route work away from it — the gap closes by building the
guard. **That work belongs to agent-ops and not to this repository.** telltale seats the
vendor and states each seat's own posture on screen; it does not own the fleet's guard
layer, and nothing here wires one. This paragraph exists so the next reader of §9.39
learns the obligation is open, and learns where it is owned.
**Amended 2026-08-09, same day, from a live room: the seat could not take a briefed turn at
all.** It was merged green and failed on its first real dispatch — `✗ failed 0s`, every seat
answering but this one. The whole of it:
```
error: unexpected argument '--- operating context ---
You are in a room...' found
tip: to pass '...' as a value, use '-- ...'
```
`-p` was passed SEPARATED from its value. `Brief.Apply` prepends a fence to every first turn,
so council's real prompt begins with `---`, and clap will not accept a hyphen-leading token as
a flag's value unless that flag opts into `allow_hyphen_values` — which this one does not. So
grok read the entire brief as an unknown flag and exited 2 before emitting a single event.
Exit 2 with an empty stdout is the failure shape this adapter already documents for a bad
resume id, so the room reported it correctly; there was simply nothing to report but a dead
turn.
The fix is one token: `--single=`, attached. Everything after the first `=` is the
value, hyphens and newlines included. Verified against the exact failing shape on the first
turn, and composed with `--resume`, where the resumed turn recalled a codeword only the first
turn carried. The long spelling is deliberate — clap's attached form for a SHORT flag is
`-pVALUE`, not `-p=VALUE`, so `--single=` is the unambiguous one. `--resume` keeps the
separated form, because a session id is a UUID and cannot begin with a hyphen.
**The lesson is about the live test, not about clap.** This seat shipped WITH an end-to-end
live test that ran the real argv, and that test passed — because its prompt was
`"Reply with exactly: LIVEOK"`, which begins with a letter. The test exercised the transport
and never the shape the product actually sends. That is a narrower version of the same
mistake §9.39 was written to avoid: a claim verified against a case nobody ships is not
verified. `grok_live_test.go` now sends a fenced prompt on both turns, and
`TestGrokAttachesThePromptToItsFlag` pins the property offline — no argv element may be a bare
`-p`/`--single`, and none may be the naked prompt.
Worth stating for the next adapter, because it generalises past this vendor: **a probe prompt
should be shaped like a brief, not like a greeting.** Three of this file's captures used
friendly one-liners, and the one hazard they could never have surfaced is the one that took
the seat down on its first real turn.
**Amended 2026-08-14: the drift alarm fired, and the seat was re-measured rather than re-dated.**
grok reached **1.0.4 (d846eb93d9)**, four patch versions past the pin every claim above rests on,
and nothing in the repository had noticed. PR #174 added version-pinned wire fixtures for exactly
this, so this is the mechanism working, not a surprise. The re-measure cost three billed turns and
one free one, and it is written up as a measurement because a version bump is not evidence that
anything changed — nor evidence that nothing did.
**The wire is unchanged, and that is a checked claim rather than an impression.** The 1.0.4
capture and the 1.0.0 one were diffed by SHAPE — frame types, key names, nesting, value types —
and they are identical. The single difference in the two files is the KEY of the `modelUsage`
map, `grok-4.5-build` → `grok-4.6-build`, which is a model id rather than a schema key, and one
this seat's parser never reads. `testdata/wire/grok-1.0.4-turn.jsonl` replaces the 1.0.0 file.
**Both containment flags are still dead.**
- `--permission-mode plan` is refuted again, with the write landing again: the `write` tool was
called, the update reported `completed`, the process exited 0, and `probe-plan.txt` held
`WROTE` on disk. `--help` still offers `plan` among six permission modes, which is the whole
point — the help text has said the same thing across five builds while the flag has never once
been observed to stop a write.
- `--sandbox` is still unobservable on Windows: `bogus-profile-xyz` drew no error, no warning and
exit 0. **This one was re-probed for free, and the technique generalises.** It was passed
alongside `--single=/context` — a prompt this vendor is already known to eat — because profile
validation happens at startup, so a refusal surfaces before any model turn. A turn that is
never billed still answers the question. macOS still diverges and fails closed (`PARITY.md`).
**The flag that earned its re-run is `--resume`.** 1.0.4's help spells it
`-r, --resume []` — an OPTIONAL value that also matches session titles,
where the pinned build took a required id. An optional-value flag is precisely the clap shape
whose SEPARATED form can stop binding, and this seat passes the id separated. Had it stopped
binding, every follow-up turn would have quietly become a fresh conversation — the room's worst
failure mode, because it looks like success. It did not: the resumed turn echoed the same
`sessionId`, recalled the first turn's own word, and reported `input_tokens` 454 against
`cache_read_input_tokens` 21504. The conversation was on the vendor's side.
**The slash hazard survives, and so does the absent-cost shape.** `--single=/context` still
produces `available_commands`, then an `end` with no usage, no cost and no text, at exit 0. So
`TotalCostUSD`'s pointer-ness is measured on the CURRENT build rather than inherited from the
pinned one.
**What grok gained, and what it did not.** `grok models` now offers `grok-4.6` (default) and
`grok-4.5`, where the 2026-08-09 survey found one model. The CLI has a `grok trace` subcommand
that exports or uploads a session's trace data. **Neither changes a verdict here, and nothing is
built on either.** Quota is still structurally absent (§7.16a, #195): a rate/limit/quota sweep
over a session directory 1.0.4 itself wrote matches nothing account-level, and `grok trace` moves
a transcript rather than reporting a ceiling. The telemetry seam §7.16a already spends is the
OTLP push, and it is untouched.
**One shape was measured that nobody had looked at before, and it is recorded as new rather than
as drift.** A WRITE tool call's `content` element is a diff —
`{"type":"diff","path":…,"oldText":"","newText":"WROTE\n"}` — with no nested `content` object, so
`grokDetail` reads nothing from it. No write call was captured at 1.0.0, so this is a gap in the
original measurement rather than a change under us, and `grok.go` says so at the function. It is
left alone: a detail renders only on a FAILED outcome, and composing a sentence out of
`oldText`/`newText` would be this package writing the vendor's line for it (§9.6a).
**Not re-verified at 1.0.4:** the bad-`--resume`-id error shape, the `--verbatim` and
`--prompt-json` slash channels, and anything on macOS.
### 9.40 the room said something was stopped and never said which seat (2026-08-09)
The gate (§9.8) blocks a vendor until a key is pressed, and the room already announced that in
two places. Neither of them names the seat.
- The **card** sits in the blocked column and does not have to name it: its position *is* the
seat, which is exactly why `gateCard` passes an empty subject there.
- The **mode line** cannot. `gateLabel` prints the oldest request's own text —
`GATE Write: internal/council/gate.go (+2 queued)` — which says what is blocked and never who.
So in a five-seat room the footer says something is stopped, and the reader finds out which
column by going through them one at a time while it stays stopped. On a projector, driving four
seats live, that is the stall. **The needs-you strip is one line of room chrome that answers the
question the other two cannot:** `⚠ NEEDS YOU 2 Codex 3 Antigravity`.
**It is driven by the gate queue and by nothing else, and that is the whole safety property.**
`State.Gates` is a structured record of vendors that asked for permission and have not been
answered. Every name on this line comes from one of those entries. A seat that has gone quiet,
a seat streaming nothing, a seat whose prose happens to end in a question mark — none of them
reach it, because none of them is a *measurement* that anyone is blocked (§4a.1). "Needs you"
is a claim about a vendor waiting on a keystroke, and the queue is the only thing in this room
that knows. `TestTheStripSaysNothingWithoutAPendingGate` builds all three of those look-alikes
at once and asserts no strip is drawn.
**A seat leaves when the reader goes to it — and that is the only thing besides answering that
takes it off.** Derived from `State.Focus` rather than stored as an acknowledged set, on
`Gating()`'s own argument: a stored set is a second place for the same fact to live, and the two
drift the first time a seat's gate is answered and a *new* one arrives while the old
acknowledgement is still in the map. At that point the anti-stall silently omits a seat that is
waiting, which is the one failure it exists to prevent. Derived, the worst case is the strip
re-listing a seat the reader visited and left — and that is true: it is still stopped and nobody
is looking at it any more. `TestGoingToASeatIsWhatClearsIt` pins the three clearings that are
refused (the seat going quiet, the seat producing output, time passing) alongside the one that
is allowed.
**The default focus is a hole in that rule, and it is deliberately left open.** `NewState` seats
the keys on column 0 without anyone pressing anything, so a gate on whichever seat happens to be
focused never appears here at all. That is the correct outcome rather than a gap: the reader is
looking at the column whose card is already spelling the question out in full, and a room-level
line naming the seat under their own cursor is the duplication §9.30 spent a section removing.
#### Where it sits, and what it does not yield to
Directly under the frame's heavy rule — **above** the collapsed-seat notice and above the band.
The other two are ordered by subject size (the room outranks the turn); this one is ordered by
**urgency**, on `modeLine`'s own precedent: a gate is the only state in this room where
something is STOPPED until a key is pressed, which is why `GATE` outranks every other mode word
on the footer. Seats that are not on screen and a brief that was just sent are facts a reader
can come back to. A blocked vendor is not.
It costs a row and, unlike the band, **it does not yield to a short terminal.** The band's whole
value is removing a duplication, so retiring it falls back to the frame as it was and loses
nothing that was not already on screen. This line is the only place the room says *which* seat
is stopped, so a row spent on it is the last row the height budget should reclaim. It is spent
in **every tier** for the same reason — at the tabs tier the blocked seat may be the one column
not on screen, which is precisely where a reader has no other way to learn it exists.
#### The ladder, and one ordering that had to be reasoned about again
Longest-first, widest-that-fits-wins — `stripHeader`'s idiom, so this package has one shedding
shape rather than three. Three rungs:
1. every seat, by name;
2. every seat, by the two-letter tag §9.25 made permanent;
3. as many tagged seats as fit, and a count of the rest (`+2 more`).
**Identity yields before a SEAT does**, which is §9.18's order and the rung boundary worth
writing down. A four-seat strip at sixty columns can hold `2 Codex 3 Antigravity +2 more` or
it can hold `2 CX 3 AG 4 CU 5 GR`, and the second is better by the only measure this line
has: the reader is asking who is stopped, four abbreviations they already learned from the
column headers answer it completely, and two names plus a count answer half of it. So the tag
rung is tried at *full roster* before any seat is dropped.
Nothing is ever clipped at any rung — an entry survives whole or leaves and is counted, because
`Ant` is not a shortened `Antigravity`, it is a seat this room does not have (§9.18). The count
is never traded away either, on `overflowMarker`'s rule: "there is more" without a number is the
marker §9.10 shipped and got reported as a room that could not scroll. Below the width where
even one tagged seat fits, `⚠ NEEDS YOU` survives alone — still true, still the signal, and
honest about being unable to say who. That floor is unreachable in a real frame (MinWidth leaves
the strip 56 cells and the lead costs eleven) and is written out anyway, because the last time a
floor was assumed rather than enforced it was wrong by four.
#### Two smaller rulings
**The seat number is printed only where the key is live.** Digits focus a seat through
`viewKey`, and `gateKey` falls through to it, so while the room is gating the number works in
both modes — except on a turn page, where `focusSeat` refuses outright because there are no
columns to move between (§9.22). A number printed there would be the room naming a key that does
nothing, which is §7.8's surprise. The names stay, because *who is stopped* is still true on a
page.
**A seat folded out of the grid is still named.** Its vendor is blocked whether or not the room
drew it a column, and a blocked vendor with no card, no column and no line anywhere is the
disappearance §4a.1 forbids. It carries no number, because no key in this room reaches it, and
it is therefore the one entry that focus cannot clear — which is honest: nothing the reader can
press from here will unblock it.
**No new hue, and no new site for the old one.** The line is `Alert` — SevWarn at weight, the
gate card's own title style — with the seat numbers `Muted`, which is the chrome/anchor split
the column header and the tab bar already use for those same two things. §9.28's list of three
places a seat hue is spent stays closed; it was ratified as closed, and a fourth site is a
decision for whoever wants to reopen it rather than a side effect of this line. Under `--ascii`
and `NO_COLOR` the strip reads exactly the same, which is the property every distinction this UI
makes has to have: `NEEDS YOU` is the signal, and the mark and the weight only make it findable.
#### 2026-08-16: the Notification hook, measured on three runtime surfaces
The strip above reads council's own gate queue and nothing else. A 2026-08-15 research candidate
proposed a second source for it: Claude Code's `Notification` hook, with the matcher names
`agent_needs_input`, `permission_prompt` and `agent_completed`. Those three names were a vendor
claim. No measurement in this repository supported them. §7.21's trap 1 measured that hook firing
changes with the runtime surface, so a claim about one surface says nothing about another. This
subsection records the survey. It builds nothing.
**The rig.** A throwaway workspace holds a `.claude/settings.json` that registers one recorder
command against every hook event the survey can name, including three `Notification` entries: one
with no matcher, one with `matcher: "permission_prompt"`, and one with
`matcher: "agent_needs_input"`. The recorder appends its raw stdin plus an arrival timestamp to a file named
for its entry, and exits 0 on every path. **The filesystem is the observable, never the stream**,
which is §9.8's rule. A file that does not exist means the hook did not run. The operator's own
`~/.claude/settings.json` was never written to, and no credential store was copied anywhere.
Claude Code **2.1.233**, Windows 11, `claude-haiku-4-5` on every turn, two trials per arm. The
turn asks for one shell command, `install -d probe-marker`, which is the shape §9.8 already
measured that no allow rule on this box covers.
**The rig lied once, and the record says so.** The recorder took an output directory argument and
ignored it, so the first four arms reported "no hook files" when the files were landing in the
workspace directory instead. That reading survived three arms and produced a false conclusion,
that project settings were not loaded at all. It was caught by running the recorder by hand. The
zero a probe reports is a claim about the probe until the probe itself is checked, and this one
was wrong.
**The source read, at the pinned version.** CLAUDE.md permits a source read at a pinned version as
evidence, and the shipped 2.1.233 binary carries four facts the live arms then tested. First,
`notification_type` is an enum of eleven values, not three: `permission_prompt`, `idle_prompt`,
`auth_success`, `elicitation_dialog`, `agent_needs_input`, `agent_completed`,
`elicitation_url_dialog`, `worker_permission_prompt`, `push_notification`, `computer_use_enter`
and `computer_use_exit`. Second, the payload builder sets `hook_event_name` to `"Notification"`,
copies the type into a `notification_type` field, and passes that same type as the hook's match
query. The matcher therefore does select on the notification type, and the three claimed names are
real matcher values. Third, the permission notification is armed by a `setTimeout` of **6000 ms**
that returns a cancel function, and it is armed immediately before the `can_use_tool` control
request is sent. The environment variable `CLAUDE_CODE_DISABLE_PERMISSION_PROMPT_NOTIFY_HOOKS`
turns it off. Fourth, `agent_needs_input` and `agent_completed` are emitted from a React effect
that watches **background agent** sessions, beside a telemetry event carrying a `jobSessionId`.
Their subject is a background job, not the session the hook is installed in.
**What fired, per surface.** Every arm denied the tool call, so the `PostToolUse` column is
uninformative and is left out: a call that never ran has nothing to report, and this rig therefore
neither confirms nor contradicts §7.21's trap 1.
| surface | `Notification` (no matcher) | `permission_prompt` | `agent_needs_input` | controls that did fire | trials |
|---|---|---|---|---|---|
| `claude -p` | no | no | no | `PreToolUse`, `SessionStart`, `Stop`, `UserPromptSubmit` | 2/2 |
| `claude -p --output-format stream-json --verbose` | no | no | no | the same four | 2/2 |
| control protocol, request answered after 12 s | **yes** | **yes** | no | the same four | 2/2 |
| control protocol, request answered at once | no | no | no | the same four | 2/2 |
The control-protocol rows replicate council's own gated seat: `baseArgs` plus `gateArgs` plus
`--input-format stream-json`, with the recorder supplied through `--settings`. The last two rows
are the same rig, and they change one thing, which is how long the probe waits before it answers
the `can_use_tool` request.
**The six seconds are the finding, and the fast arm is what makes it one.** A request held for
12 s produced the hook on both trials. The notification arrived 6931 ms and 9727 ms after the
request appeared on stdout, which is the 6000 ms timer plus the cost of starting the hook process.
The same rig answering the same request immediately produced **no `Notification` at all**, on both
trials, while all four control hooks ran in every arm and prove the settings file was loaded. So
the hook does not report that a vendor is blocked. It reports that a vendor **stayed** blocked for
six seconds. Every prompt an operator answers faster than that is invisible to it.
**The payload, re-typed with synthesized identifiers.** This is the whole record. The session id,
the prompt id and the paths are fake, per the fixtures rule.
```json
{
"session_id": "11111111-2222-3333-4444-555555555555",
"transcript_path": "C:\\Users\\example\\.claude\\projects\\C--probe-ws\\11111111-2222-3333-4444-555555555555.jsonl",
"cwd": "C:\\probe-ws",
"prompt_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee",
"hook_event_name": "Notification",
"message": "Claude needs your permission to use Bash",
"notification_type": "permission_prompt"
}
```
**There is no correlation handle in it.** The payload carries no `tool_use_id`, no control-request
id, and no tool-name field. The `title` key the builder allows was absent on every trial. The only
statement of *what* is blocked is an English sentence, so a reader that wanted the tool name would
have to parse `message`. §4a.1 does not permit that to become a displayed value.
**The verdict for this strip: refuse the seam, and the refusal is `CapNone`-shaped.** For a seat
council spawned, the gate queue already holds the fact, holds it immediately, and holds it
structurally, with the vendor, the tool and the request id all present. The `Notification` hook
offers the same fact six seconds later, without the tool identity, and only when the operator was
slow. It is strictly worse on every axis the strip cares about, and adding it would break the
safety property §9.40 is built on, which is that every name on the line comes from a measured
pending gate. A second source that is late and lossy would put a seat on the line after the
operator already answered it, or leave one off entirely.
**Two of the three claimed names never appeared, and the source read says why.** `permission_prompt`
exists and fires. `agent_needs_input` and `agent_completed` exist in the enum and fired on no
headless surface in any of the ten runs. Their emitter watches background agent sessions from a
render effect, so their subject is a background job rather than the current session, and the
effect has no render tree to run in outside the interactive terminal. A reader that took the three
names as equivalent would have wired two signals that cannot arrive and one that arrives late.
**What this does not close.** The interactive surface is **OWED**, and only an operator can drive
it. The prepared steps are: start `claude` in a throwaway directory whose `.claude/settings.json`
carries the recorder, ask for `install -d probe-marker`, leave the approval prompt untouched for
more than ten seconds, then answer it; repeat and answer within two seconds. The open question
there is whether the interactive surface adds `idle_prompt`, and whether a real background agent
makes `agent_needs_input` fire at all. Three smaller gaps stay open beside it. The operator's own
`~/.claude/settings.json` could not be read, because the credential guard refused it and the
refusal was not worked around, so the number of user-level hooks in the control column is unknown.
`worker_permission_prompt` was never registered and never seen. And the six-second timer was
measured only from the request appearing on stdout, which is a few milliseconds after the CLI arms
it, so 6931 ms is an upper bound on the delay rather than the delay itself.
**A different question this survey does not answer.** Everything above concerns seats council
spawns, where the gate queue is the better source. Claude Code sessions telltale merely observes,
through the HUD and the statusline, have no gate queue behind them. Whether the `Notification`
hook could carry a needs-input signal for **those** is untested here, and it inherits the same six
second delay and the same missing tool identity, so it would start from a weak position.
### 9.41 the gate asked about an edit and would not show it (2026-08-09)
The approval card (§9.8) names the call — `⚠ waiting on you: Edit: internal/council/gate.go` —
and that is the whole of what a user has to decide on. For a `Bash` it is enough: the command
*is* the action. For an edit it is not even close. The path says which file is about to change
and says nothing about *what changes*, so the only two answers available are "yes, because I
trust it" and "no, because I don't" — which is the gate reduced to a mood. **The card now shows
the edit itself, as a red/green before/after, when the vendor's payload carried one.**
```
⚠ waiting on you: Edit:
internal/council/gate.go
- func gateCard() {
- return nil
- }
+ func gateCard() []string {
+ return lines
+ }
y approve n deny a stop asking
```
#### What it renders is what was measured, and nothing else
This is §4a.1 on the one card in the product that guards a write, so the rule is stricter here
than anywhere: **council never opens the file, never reconstructs a before from an after, and
never shows one half as if it were two.** A preview is drawn only when the vendor's own
permission request carried *both* halves. Everything else — every `Bash`, every `Read`, every
`Write`, every request from the Cursor seat — renders the card exactly as it rendered before
this section existed, and `TestAPayloadWithoutABeforeShowsNoPreview` pins that by comparing
against the *existing* `gate-card` golden rather than a new one of its own.
Which payloads carry both was measured on 2026-08-09 against **Claude Code 2.1.226** on Windows,
driving the gated invocation (`--permission-prompt-tool stdio`, `--permission-mode manual`,
`--setting-sources ""`) in a throwaway directory. Two requests from that session, quoted whole
on `runner.Gate` and replayed verbatim as tests:
| tool | what the payload carries | what the card draws |
|---|---|---|
| **Edit** | `old_string` **and** `new_string` (plus a `replace_all` the room has no use for) | the before/after |
| **Write** | `file_path` and `content` — the after, and no before at all | nothing |
| **Cursor / ACP** (`session/request_permission`) | a title, a kind, and an options list (§9.36's capture) | nothing |
`content` is deliberately *not* read as a new half. A Write knows what the file will say and
says nothing about what it says now, so treating it as an addition would paint a green block
against a before council never saw — a plausible value where §4a.1 requires an absent one. The
Cursor seat is the same refusal one level up, and a smaller hole than it sounds: that seat does
not ask about edits at all (§9.36), so the card it does raise is about a command, where the
command is already the whole decision.
**The two halves are a pair or neither.** `editHalves` fills both or returns nothing, which
makes the renderer's test simply *do they differ* — and that one question folds in three "show
nothing" cases honestly: no halves at all, an edit that changes nothing, and an edit whose only
difference was a redacted secret. An empty `new_string` beside a non-empty `old_string` is a
legal, measured **deletion** and draws all removals and no added-side count; "0 more added
lines" would be the card filling a slot rather than answering a question.
#### The one thing that crosses the Input boundary, and why it is two strings
`runner.Gate.Input` is the vendor's whole argument blob, held only to be echoed back on an
approval, and it has always been kept **off** `State` on purpose: for a Write it is the entire
file content, one careless line away from the screen. That rule is not repealed here. What
crosses is a **projection** — two named strings, `PendingGate.Old` and `.New` — because two
fields with one purpose cannot be reached for by accident the way a map can. They are read by
the *adapter*, not by council, for the same reason the card's `Text` is composed there: the key
names are that vendor's, and the room renders a preview without knowing whose spelling produced
it.
They are redacted on the way through, and that matters more here than on the argument line the
existing redaction was written for. A command is a *likely* place for a token to appear; the
body of a file being edited is where one actually lives — and this lands in chrome that does not
scroll away. Each half is redacted separately because each is rendered separately. Only trailing
newlines are trimmed, never leading whitespace: an indent is content, and a preview that
silently unindented the code it is asking about would be showing an edit nobody requested.
#### The prefixes carry it; the colour only seconds it
`-` and `+` are the entire signal, exactly as they are on §9.37's raw patch lines — which is why
the preview reads identically under `--ascii` and `NO_COLOR`, and why the goldens, which render
`PlainStyles`, are the *proof* of that rather than an approximation. The styling reuses
`Styles.ForDiffLine`, the same classifier those patch lines already go through, so **council adds
no hue for this**: green-for-added and red-for-removed is one convention spent twice, not a
second vocabulary, and §9.28's closed list is untouched.
The marks are patch punctuation and **not** entries in the `Glyphs` alphabet, which is worth
saying because the ASCII set already spends `+` on `ActOK` and `-` on the light rule. They do not
collide, on `Glyphs.Range`'s own slot argument: those are marks that stand alone in a slot, and
these only ever open a line inside this block.
#### Bounded, per half, and it says what it dropped
The card is **chrome** — it costs body lines, `MaxScroll` derives the ceiling from it, and every
row spent here is a row of the reply the user cannot read while deciding. So each half shows at
most **three** lines and then counts: `2 more removed lines not shown`.
The count is **per half**, not one total, because a long removal would otherwise spend the whole
budget and take the additions with it, with no line admitting the additions had ever been there.
It carries no glyph of its own: `…` has an ASCII partner of `>`, and `> 2 more removed lines`
reads as a comparison rather than as a truncation — a marker that can be misread as a number is
worse than a longer sentence. Long *lines* are cut with the ellipsis glyph on the plain text
**before** the style is applied — classify, truncate, style, and only then let `fit` pad — because
`fit` alone clips silently and would be clipping a string that now genuinely carries ANSI, which
is §9.5's trap read backwards.
#### What is deliberately not here
**No intra-line diff.** The unit is the line, because that is the unit the payload arrives in;
highlighting *which words* changed inside a line would be council computing a diff rather than
displaying one, and the first thing it would get wrong is the case it was added for.
**No scrolling the preview.** A card with its own viewport is a second scroll surface in a room
that already has one per column plus a turn page, and the keys to move it would have to be taken
from a mode whose whole contract is that `y` and `n` mean one thing each (§7.8). The bound plus
an honest count is the trade; the whole edit is in the file the moment it is approved.
**No preview on a denial or after the fact.** The trace still says `✗ denied by you` and nothing
more. What the seat *would* have written is not what happened, and a room that showed it
afterwards would be displaying a file state that never existed.
### 9.42 `telltale doctor`: the one moment probing is allowed, and the three answers it may give (2026-08-09)
Council's detection has never run a vendor. `detect.go` says why in its own doc — "council
never runs a vendor to find out whether it works: a probe turn costs real quota, and 'is it
authenticated?' is a question the first real dispatch answers for free" (ADR-008 §6) — and that
rule is correct and stays. What it is a rule *about* is a **turn**: the room may not spend the
user's money to answer a question it does not have to ask, and it may not spend it silently,
mid-conversation, on a schedule nobody asked for.
So the boundary is drawn at **cost and side effect**, not at "running a vendor is forbidden".
`telltale doctor` runs ` --version` — a flag that parses argv, prints a string and
exits, with no model, no session and no billing anywhere in it. It starts no turn, reads no
credential store, makes no network call, and writes nothing at all: it adds no fourth exception
to the three writes `CLAUDE.md` lists, because it has none. §9.17 is the frame that makes this
the right place for it — *a fact that is true at launch and stays true belongs at launch*, and
what binary is on this disk and what it calls itself is exactly that shape. The inverse of the
same rule is why there is no `/doctor`: nothing this reports changes while the room is open.
**Three states, and the whole design is keeping them three.**
| | what it means | what it may carry |
|---|---|---|
| `ok` | the check ran and passed | the measured text — the path that was stat'd, the line the vendor printed |
| `FAILED` | the check ran and did not pass | the reason |
| `not checked` | the check did not run | **why** it did not, and never a value |
**Auth and network are `not checked` on every seat, always.** A binary that exists and answers
`--version` establishes nothing whatever about a login or about reachability, and a report that
let those two ride along on the good news would be believed on the one day it was wrong — the
§4a.1 collapse wearing a preflight's clothes. The reason is printed rather than shrugged: the
only thing that establishes a login is a real turn, and a turn costs quota. A seat that is
installed and signed out reports its own auth failure on its column the first time you dispatch
to it, which is where that fact was always going to come from.
**`NotChecked` is the zero value**, for §9.17's `GateOff` reason exactly: a safety property
whose default is the reassuring answer is the wrong way round however carefully the constructor
sets it. Every `Check` a test types by hand, and every field a future change forgets to fill in,
reads as an honest blank instead of a silent pass. `TestNotCheckedIsTheZeroStatus` pins it.
**The measurement that would have been a plausible lie.** One seat's resolved binary is not the
program it names. Detection steps over `cursor-agent.cmd` to the bundled `node.exe` its launcher
would have run (§9.33), so the obvious `--version` answers **`v24.5.0`** — node's version,
printed under a row labelled `cursor`, real and about the wrong program. Handed the bundle first,
`node.exe index.js --version`, the same install answers **`2026.08.04-aaa8809`**. Both measured
here, one after the other. The version argv is therefore per-seat data taken through
`vendors.CursorNodeBundle`, not a `--version` constant and not a second `filepath.Join` — that
function's own doc names two copies of one join as the agreement that silently stops holding.
The other four take the bare flag, verified by running them: claude `2.1.226 (Claude Code)`,
codex `codex-cli 0.147.0`, agy `1.1.11`, grok `grok 1.0.0 (3cd0d0cbce) [stable]`.
**Found and undrivable is a third thing, and it stays one.** `detect.go` already refuses to
collapse "not installed" with "installed somewhere council will not drive from", because the fix
differs. The preflight carries that through as two checks rather than one verdict: `binary`
passes and `drivable` fails, with the shim note attached. The version probe still runs on such a
seat — what is installed is worth knowing where the room cannot use it, and a fixed `--version`
carries none of the prompt text that made it undrivable.
**A fifth mode, not a council flag, and no colour in it.** What it prints goes somewhere else:
council's output is a full-screen room, so a preflight rendered inside one is unreadable at
exactly the moment it is wanted — before the room opens, piped into a file, pasted into an issue.
Every distinction it makes is carried by a **word** (`ok`, `FAILED`, `not checked`), so
`--ascii` and `NO_COLOR` have nothing to switch off and neither is a flag here; a flag that does
nothing is a promise that something was configurable. `Render` is pure over its `Report` for
council's reason, with the probe durations measured in `Run` and arriving as data. Each seat gets
its own `--timeout` (15s default) rather than the run sharing one: a wedged vendor must cost its
own deadline and a failed check, never the report. And the command exits **0** whatever it finds
— a failed check is this mode working, and exiting non-zero would make "I have four of five
seats" indistinguishable from "doctor itself broke".
**Capabilities are printed and are explicitly not checks.** How a seat streams and whether it can
be asked to ask first were measured once, against live runs (§9.7, §9.33, §9.36, §9.39, and
`canGate`), and written down; nothing re-measures them on this machine now. They render on their
own labelled line — *"council declares, and did not check here"* — outside the status column,
because putting them in it would give them a fourth state and imply a check that never happened.
The words are read off `granularityFor` and `canGate` rather than restated, so the preflight
cannot drift from the room's own badges. Worth the line at all because "installed" and "will
stream to you" are different promises, and the seat a user is most likely to think is broken is
the one that is working and silent until the end of the turn (§9.14).
#### 2026-08-16: the survey pin gets its maintainer loop, and it lives here
**The gap.** Every adapter is pinned to the vendor build its field map was surveyed at, and
§3.10 is the inventory of those pins. `internal/adapter/drift` watches the on-disk SHAPE and
tells the USER when a row goes quiet. Nothing told the MAINTAINER that the survey behind a row
is now older than the vendor on this disk. `agy` and `grok` self-update, so that pin goes stale
in silence: no check fails, no row blanks, and a set of displayed fields rests on a survey of a
build nobody runs any more.
**CI can never close this.** CI installs no vendors, so every version probe there resolves to
nothing and every comparison is vacuous. The loop has to live where the vendors live. The owner
ruled on 2026-08-16 that the place is this mode, and it is the obvious host: `doctor` already
resolves each seat's binary and already asks it its version. The comparison costs one string
compare on a probe that was happening anyway.
**It is a staleness fact, and it is not one of the three states.** A drifted pin says nothing is
wrong with the reader's machine. The seat works, the binary answered, and every check that ran
passed. So the notice renders on its own labelled line — *"telltale's field map for this vendor
was measured at"* — beside the capability line and outside the status column, for the
neighbouring reason: both are claims this repository measured once and wrote down, and neither
was re-measured here. Making it `FAILED` would redden a working vendor, move the tally, drop
`Ready`, and put a red cross in front of a user whose install is perfect. Making it `Passed`
would claim a check ran on their machine when what actually happened is that this repository
compared its own homework. The tally, `Ready` and the exit code are all untouched, and
`TestDriftIsNotAFailedCheck` pins that by running the same seat with and without a pin and
requiring the three counts to be identical.
**Four outcomes, and the two silent ones are silent on purpose.**
| | what the line says |
|---|---|
| the versions differ | the pinned build, the installed build, and the §3 section to re-measure |
| they match | the pinned build, and that it is what is installed |
| the version was not read | the pinned build, and that nothing is claimed in either direction |
| the two cannot be compared | the pinned build, and why no comparison is possible |
Rows three and four are the honesty of the feature. A seat whose `--version` probe failed hands
the comparison an empty string, and an empty string gets **no verdict** — claiming a match there
would be as invented as claiming drift, and both would rest on nothing measured. Row four is the
Cursor seat specifically: its pin names the Cursor **application** the store was surveyed inside
(§3.9, `Cursor 3.14.7`), and what this mode probes is **cursor-agent**, which answers a
date-stamped build of its own (`2026.08.04-aaa8809`, §9.42's own version table). Comparing
`3.14.7` with `2026.08.04` would manufacture a permanent drift notice out of two unrelated
numbering schemes — a notice that fired forever on a correct install, which is exactly the
report `internal/adapter/drift` rules out as one nobody reads.
**The comparison is equality, never ordering.** It extracts the first dotted-numeric run from
each string, because the pin and the probe agree on a number and on nothing else: the adapter
wrote `Claude Code 2.1.219` and the binary answered `2.1.226 (Claude Code)` when this was
measured (that pin has since moved to `2.1.233`, §3.1's re-measure); the adapter writes
`grok 1.0.4 (d846eb93d9)` and the binary answers `grok 1.0.0 (3cd0d0cbce) [stable]`. A commit
hash carries digits but no dot, so it cannot be mistaken for the version beside it. No line
claims "newer" or "older": a direction needs per-vendor precedence rules this program has no
business inventing, and a downgrade means the same thing an upgrade does — the survey was
measured somewhere else. `TestTheComparisonNeverClaimsADirection` holds that.
**This does not make a version comparison a canary, and nothing on the read path changed.**
`internal/adapter/drift`'s doc rules a version out as a TRIGGER, because a report that fires on
every release is a report nobody reads. That ruling stands untouched: no adapter degrades a
field, writes a diagnostic, or reads anything new. The version comparison is admissible only
because it is asked once, by hand, before the room opens, by an operator who ran a preflight —
and it is answered to a different reader, about this repository rather than about their machine.
The two are separated in the report itself: the closing note says a format that actually moved is
reported on the HUD row and does not wait for this.
**One source for the pins, enforced by the compiler.** `internal/adapter/pins` is the table, and
every `VerifiedAgainst` in it **is the adapter's own constant** — which is why those six
constants are now exported. No pin string is copied. `Section` and `DocLabel` are the only facts
the table adds, and they have no other home: an adapter knows what it was verified against, not
which prose section carries the evidence, and that pointer is what turns "this is stale" into an
instruction. `internal/doctor` stays stdlib-only and holds no inventory; `council.DoctorSeats`
attaches the pin, exactly as it already attaches the capability declaration, and for the same
stated reason.
**The doc guard reads the cell, not the document.** §3.8 records that
`TestTheCanaryInventoryMatchesThisAdapter` substring-matches a pin *anywhere* in this file, so a
dated paragraph quoting a new pin turns the guard green while the table it guards stays wrong —
which is what let the Antigravity cell read `agy 1.1.9` for a release after the adapter moved to
`1.1.13`. §3.8 named the fix, "scoping that assertion to the table", and left it unowned.
`pins_doc_test.go` parses §3.10's table and compares it cell by cell, in both directions: a pin
quoted in prose elsewhere cannot satisfy it, and a doc row no adapter claims fails it too. This
was verified by mutation — restoring the stale `agy 1.1.9` cell fails the new guard while
`internal/adapter/antigravity`'s own test stays green, reproducing the 2026-08-03 miss exactly.
**The six adapter tests are deliberately left as they are.** They assert a different thing (that
a pin is documented at all), nothing here weakens them, and rewriting them is a separate concern.
**Measured, not assumed.** Run on this Windows box on 2026-08-16, the report drifts `claude`
(pinned `2.1.219`, installed `2.1.233`) and `codex` (pinned `0.146.0`, installed `0.147.0`),
matches `agy 1.1.13` and `grok 1.0.4`, and declines to compare `cursor`. That is the feature
finding two genuinely stale surveys on its first live run. **Re-measuring those two surveys is
not part of this change** — the loop reports staleness; a re-survey is its own work, with its own
live corpus.
#### 2026-08-17: the preflight states each seat's posture, from the room's own claim
**The gap.** Every council column carries a sandbox badge, and the help panel's posture page
carries the measured argument under each one (§9.13, §9.2's ruling that a claim you cannot see
is not a claim). Both of those are read **inside the room**, which is after the decision they
inform. A user picks a workspace and a posture *before* the room opens, and the only surface
that runs before the room opens said nothing about either. §9.17 settles that the fact belongs
here: what a vendor's own flags buy on this machine is true at launch and stays true, and it is
a property of the vendor and the OS rather than of a turn.
**One source, two surfaces.** Nothing in the block is written in `internal/doctor`.
`council.DoctorSeats` builds it from `postureClaim` — the same function the room's own columns
are built from — and hands over the badge word off `SandboxClaim.Badge()` and an evidence class
off the claim's `Level`. That routing is the whole design: a preflight with a per-vendor posture
table of its own would agree with the badges on the day it was written and diverge the day a
level moved, and a reader looking at two disagreeing surfaces has no way to tell which is
lying. The capability declaration and the survey pin are attached at the same seam, for the same
stated reason. `TestThePreflightPostureIsTheRoomsOwnBadge` pins it through *different*
construction paths on each side — `DoctorSeats` against the columns `stateWith` builds — because
comparing `doctorPosture` with `postureClaim` would be comparing a call with itself.
**The badge says what the posture IS; the evidence class says what it RESTS ON.** They are
different questions, and `unsandboxed` is the case that proves it: two seats reach that badge
because a live run **refuted** the flags and because **no flag was ever passed**, and a reader
deciding whether to point council at a worktree needs the second sentence. §4a.1's rule that two
kinds of nothing must not render alike is the same rule one level up.
| badge | evidence class |
|---|---|
| `ro:tools` | enforced by **construction** — the write and shell tools are absent from the session |
| `ro:enforced` | enforced by an **operating system** — the vendor's own sandbox |
| `ro:requested` | **asked for**, and never observed on this machine — weaker than either above, and says so |
| `unsandboxed` | **measured** not to restrict — refuted by a live run, not merely unestablished |
| `WRITES` | nothing was asked for at all |
| `gated` | **your keystroke** — the seat asks before every tool call that changes anything |
`evidenceClass` is a table keyed by level, so `TestEveryPostureLevelHasAnEvidenceClass` can walk
the type and fail the build the day a sixth level renders a badge with nothing to classify it —
the guard `helpBadgeGloss` already carries inside the room. `TestNoEvidenceClassSoftensItsBadge`
holds the other half on `TestThePostureLegendDoesNotSoftenAnyClaim`'s terms exactly: these
sentences classify evidence and never weaken it, and none of them may call a posture read-only,
safe or unable to write.
**The rows are the `--read` room, and the argv is on the block.** The room **WRITES by default**
and `--read` is the opt-out (`cmd/telltale`, and the legend inside the room was already corrected
once for crediting the retired `--write` flag). The rows report the `--read` posture because that
is the only one that is a fact about the *machine* — the default room's badge is a property of an
argv the reader has not typed yet, and five cells all reading `WRITES` would carry nothing per
seat. So the header names `telltale council --read` in its first clause, and **one closing
declaration** states the default: the room writes, *n* of *m* seats can be asked to ask first,
and what contains a writing room is the workspace, not any of these words. The gating half is
**counted off `canGate`**, never written down — that measurement has already moved once, when the
Cursor seat became a live process that can be asked and still does not ask about edits.
**It is a claim, and it is not one of the three states.** A posture was measured once against a
live run and written into this repository; nothing re-measures it on the reader's machine. So it
renders outside the status column, beside the capability line and the survey pin, and it is wrong
in *both* directions as a check: a `FAILED` would redden a working install over a vendor's own
design decision, and an `ok` would claim this preflight established a containment property it
never probed and could not probe without spending a turn. `TestAPostureIsNotACheck` pins that the
way `TestDriftIsNotAFailedCheck` does — the same seat with and without the data, the three counts
required to be identical — and `TestThePostureBlockCostsNoProbe` pins the other half: the block
adds **no probe, no network call and no login check**, because every string in it arrives with
the seat. The exit code is untouched.
**A seat council states no posture for gets `no claim`, not a missing row.** A seat absent from a
posture table reads as a seat with nothing to declare; this one has an unanswered question. The
word deliberately is not shaped like `not checked` — the three state words are spoken for, and a
fourth column borrowing one would put this block back inside the block it is outside of.
**Measured on the reference Mac, 2026-08-17.** The five rows read `claude ro:tools`,
`codex ro:enforced`, `agy unsandboxed`, `cursor ro:requested`, `grok unsandboxed`, and the
declaration reads *1 of the 5 seats above can be asked to ask first: claude*. The `codex` row is
the platform branch working: at that date the same block on Windows read `unsandboxed` there,
because council passed `-s danger-full-access` on that OS (ADR-008's twelfth amendment; since
§9.2's 2026-08-29 amendment the Windows row reads `ro:enforced` too, with its own dated detail)
— and it reads it from `postureClaim`, not from a second platform test in the preflight.
### 9.43 the agy seat stops pretending a lost thread resumed (2026-08-09)
`STATE.md` carried this as an unowned gap for as long as it took to write the entry. **The agy
seat could not tell a lost thread from a resumed one, and the stream would not say.** Measured
2026-08-09 against agy 1.1.11 during the wire-fixture capture (PR #174; the record is
`internal/council/vendors/testdata/wire/README.md`, under *what could NOT be captured, and why*):
handed a `--conversation` id it does not hold, that CLI **does not error**. It opens a NEW
conversation, answers the brief normally, and reports `status: "SUCCESS"`, exit 0.
That is the whole difficulty in one sentence. Every other seat resolves the question for the
room: Claude Code returns a `result` frame whose `errors` array says *"No conversation found with
session ID: …"* (PR #178); codex writes `no rollout found` to stderr and exits 1; grok
does the same with a 404; the Cursor seat answers `session/load` with -32602 and opens a fresh
one in the same process (§9.36). agy claims success either way — so a room reading status and
exit code, which is every honest thing the seat did before this, rendered **a continued
conversation over a reply that had no history behind it.** Not a crash and not an empty column:
a plausible answer under a `restored` mark that was no longer true.
**The tell, and it is the only one the capture surfaced: the `conversation_id` that comes back is
not the one that was asked for.** Nothing read it. Reading it is this change.
**The comparison lives in council, not in the adapter, because neither half is where the other
is.** `ParseEvent` sees one line at a time and never learns which id the turn requested; dispatch
knows the request and never sees the stream. So `specFor` now *returns* the id it asked to
resume — the one fact a caller cannot re-derive, since it is buried in a vendor-specific argv
position — dispatch records it, and `adoptSession` compares it against the id the vendor names
on its own session event.
**The vendor gate is structural, and that is the honesty argument rather than a taste in
plumbing.** The arithmetic (`asked != returned`) is vendor-neutral; the *conclusion* is not. A
CLI that re-keys a resumed thread while keeping its history would look identical on the wire, so
a room that compared ids for everybody would announce lost threads it had never measured — §4a.1's
inference, wearing a comparison's clothes. The claim is therefore made by the seat, through
`vendors.SilentResumeFork`, whose one method returns **the build the fork was measured against**
rather than being a bare marker: a seat cannot make the claim without naming its evidence, and a
vendor bump that fixes the behaviour leaves a version string that no longer matches the fixture
beside it. Only agy implements it. `TestOnlyAMeasuredVendorArmsTheForkComparison` pins that it
stays alone until somebody captures a second case, and the gate is applied at *dispatch* — a seat
that never enters `forkWatch` cannot raise the card at all.
**Three rulings on what the room then does, each of which had a plausible alternative.**
| | what happens | the alternative, and why not |
|---|---|---|
| the reply | **renders, untouched** | failing the column would throw away an answer the user paid for, to punish a bookkeeping mismatch. The turn succeeded; what was false was only the claim that it was informed by everything before it |
| the new id | **adopted as this seat's thread** | discarding it orphans a real turn — the reply happened *inside* that conversation — and leaves the room rebuilding the same forking invocation on every later turn |
| the card | **the calm lost-thread card already in use** | a second card would say the same fact in different words. The outcome is identical to a refused reattach — this seat is starting fresh — and only the body differs, because there the turn failed and the id was let go, here it succeeded in a thread nobody asked for |
So the column reads *"thread not restored — starting fresh"*, quietly, with the mechanics
demoted underneath it: the seat asked to resume its saved thread, the vendor answered in a new
conversation instead and reported success, the reply below is real and the history behind it is
not, and the next brief continues from this turn. **No warning mark**, for the reason
`settleRestoredThread` states: this is the same fact `reattachCard` says calmly at idle when no
thread came back, discovered a turn later, and spending the ⚠ on it blunts the mark that carries
real failures. The seat's probation ends here too — the restored id is gone by evidence rather
than by a turn's outcome, so `settleRestoredThread` has nothing left to decide and a later
failure on the *new* thread cannot be blamed on a reattach that was already reported.
**The fixture is derived, and it is labelled derived.** `testdata/agy-forked-conversation.jsonl`
is the real 1.1.11 capture with one textual substitution — every `conversation_id` value moved
from `2222…` to `3333…`, nothing else, not a key and not a token count. It sits **outside**
`testdata/wire/`, whose contract is real captures only: the forked turn itself was measured, but
that probe's stream was not kept, and a hand-edited file among the captures would silently
restate a measurement nobody re-ran. What it proves is the narrow thing it can: a turn that looks
entirely successful still delivers the mismatched id to the room.
**What was not done, and is not claimed: no live agy turn was driven for this change.** The
behaviour rests on the 2026-08-09 capture, and the code rests on that capture's fixture. A
re-measurement against a later build is what would retire `SilentResumeForkMeasuredAt`, and
until somebody runs one, this seat's claim names 1.1.11 and no other build.
#### 2026-08-16: the agy tail, measured — and why this seat still names no end of turn
§9.33's amendments settled the codex seat on `turn.completed` and left this seat alone. The second
of them said why in one sentence: **"this vendor's linger is not measured at all"**. That sentence
is now false. This block is the measurement that replaces it, and the verdict is a measured **no**:
the marker exists, the tail does not, and `vendors/agy.go` keeps its behaviour.
**Instrument and version.** `agy 1.1.13`, read from `agy --version` at run time rather than from
this file. agy self-updates, so a version quoted from a document is a version nobody checked;
§3.8's re-verification records the same discipline. The probe ran this seat's own argv
(`--output-format stream-json --disable-slash-commands --print-timeout 30m -p `) on a
brief-shaped prompt, in a throwaway directory outside any repository. It recorded the arrival time
of every stdout and stderr line, then the process exit. It polled
`%USERPROFILE%\.gemini\antigravity-cli` every 250ms for size and mtime changes. It read no file
content, and every prompt told the model to use no tools.
**Three trials, not two.** Trials 1 and 2 are one-sentence replies. Trial 3 asks for about 600
words, because a tail that scales with the reply or with the conversation database would not show
itself in two short turns.
| | trial 1 | trial 2 | trial 3 (long reply) |
|---|---|---|---|
| `init` | 12.756s | 4.380s | 3.403s |
| first `agent_response` delta | 14.181s | 6.005s | 9.487s |
| `agent_response` DONE | 14.181s | 6.005s | 10.091s |
| `checkpoint` | 14.603s | 6.606s | 10.694s |
| **`result` (the answer)** | **14.603s** | **6.606s** | **10.694s** |
| process exit | 14.917s | 6.655s | 10.829s |
| **tail** | **0.314s** | **0.049s** | **0.135s** |
**The marker half of the question is yes.** On all three trials `result` is the LAST line on
stdout, stderr stays empty throughout, and the line carries both the full `response` text and the
`status`. The adapter already parses it. A reliable answer-complete marker exists on this seat.
**The tail half is what fails.** 0.049s, 0.135s and 0.314s, against the 4.06s and 4.25s §9.33
measured on codex the same day. `EndsTurn` would move two things and neither survives those
numbers. The turn clock would stop at `result` rather than at the exit, which corrects at most
0.314s on a clock that renders whole seconds. The column would settle at `result` and read
`exiting` until the exit, which puts a status word on screen for a third of a second at most.
§9.33's settle exists to remove a false `streaming` that ran for four seconds. It does not exist to
add a true `exiting` that nobody can read. So this seat keeps the process exit as its end-of-turn
signal, and the decision is now written down rather than left as an unexamined default.
**Nothing rides this tail either.** The vendor's own state writes were polled rather than diffed,
so this is an observation of when writes stopped. On all three trials the final write to the
conversation database, to the four transcript files under `brain//.system_generated/logs/`,
and to the CLI log all land within one 250ms poll of the exit. There is no gap between the last
receipt and the death, because there is no gap to hold one.
**What is NOT claimed.**
- **The failing turn is unmeasured.** All three trials ended `status: "SUCCESS"`. §9.33's second
amendment settles a *failed* agy turn on its `status: "ERROR"` line, and that path's tail is not
covered here. Producing a failure on purpose needs either a posture flag this adapter no longer
passes (ADR-008's seventeenth amendment) or a lost thread this CLI answers with success (§9.43),
so no probe reached one.
- **The tool-using turn is unmeasured**, for §9.33's own reason: a probe gets no write access.
- **The size is not a constant.** Three trials, one box, one build, one day.
- **The boot is not the tail.** `init` lands 3.4s to 12.8s after the spawn. That is the operator's
wait and the seat bills it honestly. It is named here only so a later reader does not mistake a
slow start for a linger.
**The same capture refutes a claim in the adapter, and the claim is corrected rather than left.**
`ParseEvent`'s doc comment said a whole `agent_response` arrives as ONE delta when the step turns
ACTIVE, plus a trailing newline on DONE — therefore a `PhaseWaiting` case and never a streaming
one. At 1.1.13 both halves are wrong. A short reply sends no ACTIVE line at all, and one DONE step
carries the whole text. A long reply sends true incremental deltas about 200ms apart (trial 3:
9.487s, 9.687s, 9.889s, with the final chunk on DONE), and each delta continues where the last one
stopped, so nothing duplicates. **No code changes for this.** The seat declares `GranFinalOnly` and
`applyEvents` promotes the phase to `streaming` on the first chunk, which is the modest-claim rule
working exactly as it was written. Only the comment was stale.
**Spend:** three billed turns, all trivial prompts, no tools.
#### 2026-08-16: what agy reports when a turn needs the operator, and why that is `CapNone`
§9.40's needs-you strip reads council's own gate queue and nothing else. The 2026-08-15 research
asked whether agy's vendor-REPORTED `agent_state` (§2.1) could become a second source for it. This
block is the measurement. The verdict is a `CapNone` refusal with **two independent reasons**, and
each reason is sufficient on its own.
**Instrument and version.** `agy 1.1.13`, read from `agy --version` at run time. agy self-updates,
so a version quoted from a document is a version nobody checked; §3.8's re-verification records the
same discipline. Windows 11. Six billed turns ran through this seat's own argv (`--output-format
stream-json --disable-slash-commands --print-timeout -p `), from a throwaway directory
outside any repository. The probe timestamped every stdout and stderr line, then the process exit.
It then read the resulting transcripts for structural fields only: `type`, `status`, `step_index`
and `created_at`. It copied no credential, and no prompt content from this machine enters this
document.
**Three shapes, two trials each.** Shape A is a normal completing turn and sets the baseline. Shape
B drives agy to its `ask_question` tool. That is the pure needs-input case: the agent asks the
operator a question and writes nothing. Shape C asks for a file write, which is the permission case.
**The first reason: the waiting state never occurs in print mode.** agy answers on the operator's
behalf, and it says so in its own words on both paths.
- **A question is skipped.** In both shape B trials the agent asked, received no answer, and
continued inside the same turn. Its own next message reads: "I've presented the prompt to select
a file to rename, but it looks like you skipped it." Both turns ended `status: "SUCCESS"`, at
6.9s and 5.2s. Neither one waited.
- **A tool permission is auto-approved.** Shape C trial 2 recorded a `SYSTEM_MESSAGE` step carrying
this text verbatim: `stop hook blocked termination due to reason: The user has automatically
approved the artifact through their review policy. Proceed to execution.` The write landed.
`init.permission_mode` reads `request-review` on all six turns, so that value does not mean the
operator is asked.
- **`ask_permission` has never been reached.** It is one of the 56 tools in the `init` tools array.
Across 90 conversations and 4,344 transcript records there are **zero** `ASK_PERMISSION` records.
`ASK_QUESTION` has five.
A gauge fed from this seam would therefore have nothing to report. `--print-timeout` is not the
ceiling `agy.go`'s comment describes either. No probe turn reached it, because no probe turn waited.
**The second reason: the print-mode stream cannot name the state it does emit.** The disk keeps the
record type. The stream discards it. The two surfaces cross-walk by `step_index` exactly:
| shape | idx | stream `step_type` / `state` | `tool_name` | `duration_seconds` | disk `type` / `status` |
|---|---|---|---|---|---|
| A (baseline) | 1 | `unknown` / DONE | absent | 0.0015, 0.0011 | `CONVERSATION_HISTORY` / DONE |
| B (needs input) | 1 | `unknown` / DONE | absent | 0.0010, 0.0010 | `CONVERSATION_HISTORY` / DONE |
| B (needs input) | 3 | `unknown` / DONE | absent | 0.5977, 0.5945 | **`ASK_QUESTION`** / DONE |
| C (write) | 3 | `tool` / ACTIVE then DONE | `write_to_file` | 13.5304, 0.5917 | `CODE_ACTION` / DONE |
**The needs-input step and the every-turn preamble step are byte-identical in every field this
adapter reads.** Both arrive as `step_type: "unknown"`, state DONE, with no `tool_name` and no
`tool_info`. Only `duration_seconds` separates them, and that is a continuous measurement rather
than a marker. A threshold over it would be the invented vocabulary §4a.1 forbids. This is the
`grok` problem the research named, in its worst form: waiting and working do not merely share
bytes, because the vendor resolved the wait before it wrote the line.
**No liveness signal pairs with it.** The status vocabulary across the whole corpus is exactly two
values, `DONE` (4,180) and `RUNNING` (164). §3.8's re-verification already refused `RUNNING` as
liveness, because its oldest rows sit in conversations nothing has touched for days. All five
`ASK_QUESTION` records carry `DONE`. Three of those five predate this probe and come from real
interactive sessions on 2026-08-09 and 2026-08-15. No record has ever been observed in a state that
means an ask is outstanding. The `RUNNING` count is also unchanged from the 2026-08-15 re-read,
which is the independent check that six completed probe turns leave no `RUNNING` residue.
**`agent_state` is not on this surface at all.** It is a statusline-payload field (§2.1), and that
payload exists only inside an interactive agy session. Council drives this vendor in print mode and
never sees it. The candidate seam and the strip that would consume it sit on different surfaces,
which is the structural half of the refusal.
**One vendor flag behaves differently than assumed, and it changes nothing today.** Shape C trial 1
passed `--mode plan` beside the seat's `--disable-slash-commands`, and agy answered on stderr:
`warning: --mode plan has no effect while slash command expansion is disabled.` The seat does not
pass `--mode` (ADR-008's seventeenth amendment), so no behaviour moves. It does sharpen that
amendment: the read-posture flag it dropped was inert twice over on this argv. Trial 2 re-ran with
plan mode live and the write still landed, which corroborates `PARITY.md`'s Antigravity row at
1.1.13.
**A claim in `agyPlumbing` is refuted by this capture, and the code was deliberately not touched.**
That comment argues `unknown` is a fixed preamble slot agy declines to name, on the evidence that
every observed one sits at `step_index` 1, carries no tool name, and lasts under 5ms. Shape B
produced a second `unknown`, at index 3, carrying a real act. The suppression still reaches the
right outcome today, because agy skipped the question before the line arrived and there is nothing
actionable to draw. The stated reason is now wrong. This lane measures rather than changes code, so
the correction is recorded here and is owed in `internal/council/vendors/agy.go`.
**What is NOT claimed.**
- **The interactive seam is unmeasured.** §3.8 observed `agent_state` transitioning, and
`tool_confirmation_pending: true`, on agy 1.1.9 statusline payloads. A re-capture at 1.1.13 needs
a TUI session and a statusline capture command, and this probe changed no agy configuration.
Whether an interactive `ASK_QUESTION` sits `RUNNING` while it waits is the open question this
block leaves behind. It is the one measurement that could reopen the seam.
- **Six turns, one box, one build, one day.**
- **The model picked its own write path in shape C trial 1.** It ignored the process cwd and wrote
under `~/.gemini/antigravity-cli/`. Both files were removed after the run. That is model
behaviour rather than a seam property, and it is named so a later probe expects it.
**Spend:** six billed turns, all short prompts.
### 9.44 the composer was a gap under a rule, and the room's state floated below it (2026-08-09)
**Inspiration is named because it should be: Grok's CLI.** Its input is a rounded, clearly
bordered box, and the bottom border carries a right-anchored legend — `Grok 4.5 (high) ·
always-approve` — laid *on* the line rather than under it, with the remaining key hint
(`Shift+Tab:mode`) on a muted line below. The thing that reads well there is not the corners. It
is that the box says **where you act**, and its own frame says **what you are acting under**,
which leaves the line below free to be nothing but keys.
Council had neither. §9.26 closed the frame with two full-bleed heavy rules, and the composer sat
in the gap between the lower one and the mode line — a prompt glyph on an unbounded row, with no
mark anywhere saying that this strip of the screen is the one place typing does anything. Every
other region in the room is a reading area. The one region that is an *input* was the only region
with no shape of its own.
**So the composer is a box, and it is the only bordered element on screen.** Rounded corners
(`╭ ─ ╮ │ ╰ ╯`, and `+ - |` in the reduced set), a side and a cell of air on each row, and the
bottom border carrying the legend. Bordering exactly one thing is the whole design: a second box
anywhere would make this one a decoration instead of a signal, the same scarcity argument §9.26
made for a second rule weight and §9.28 made for the seat hue.
**The lower heavy rule is gone, and that is a correction to §9.26 rather than a cost of this
change.** §9.26's claim was that the two full-bleed rules were the only *closed shape* on screen
and that a closed shape earns the second weight. They were never closed — two horizontal lines
with nothing joining their ends is a pair of lines, and the weight was doing the work a shape
should have done. The box actually closes, by corners and sides, so it draws **light**: closure
moved from ink to geometry, which is a carrier that survives `NO_COLOR` outright. What the heavy
rule now says is narrower and true — *the chrome stops here and the seats begin* — and there is
exactly one line in a grid where that holds.
**The legend is split by lifetime, not by topic.** On the border go the facts that stay true until
a key changes them: the mode word (`VIEW` / `COMPOSE` / `GATE` / the page label) and, when the
guard is off, `a not asking`. Under the box go the keys, which change with the mode, the draft and
the turn. Nothing appears in both places — `statusLine`'s left-hand slot is now empty and is
deliberately not backfilled, because a slot that survives its content is how a footer becomes the
wall §9.11 spent a whole pass taking apart.
The cadence cell keeps its **key** on the way up, not just its words. `a not asking` was added
(§9.24-era footer work) to close the §9.17 defect where a permanently ungated room documented the
way back nowhere on screen; moving the words without the key would have reopened it one release
later. It is on the border, whole, and unsheddable.
**What it costs, stated plainly.** One row of body — `promptChrome` goes from 2 to 3, since the
box replaces the rule with a top border and adds a bottom one — and six cells of composer width,
which is `boxChrome`: a side glyph plus `gutter` cells of air, on each side. The air is `gutter`
rather than one cell because **the room spells its separator one way** — two cells each side of
every `│` it draws, which `TestTheRoomSpellsItsSeparatorOneWay` already holds for the header, the
key line and the column rails. A box welding its prose to its own sides would be a second grammar
for the room's only vertical mark, on the element the eye is meant to read as the frame.
**The help panel paid the row, the way it always pays.** Its budget was 17 lines to the pinned `?`
and is now 16, which pushed `ctrl+c / q` off page one and the WORKSPACE sentence — the
load-bearing line — off page two at the reference machine's 24-row room. Neither was allowed to
fall: both pages spend the blank row directly under their title instead, on §9.11's own ranking
that **a rule outranks a blank**. The title *is* a labelled rule, so the blank beneath it was the
one row on each page restating a boundary the row above already drew.
`MinWidth` (60) and `MinHeight` (10) are unchanged and still
resolve: at the floor the room is 2 header rows, 3 of footer chrome, a one-row composer and four
rows of reading area. When the frame is too narrow to lay the legend on the border with air each
side and a rule cell outboard of it, the border **closes bare** rather than truncating — a legend
cut in half is a claim cut in half.
**One golden moved further than the box did, and it is a gain rather than a surprise.** Emptying
`statusLine`'s left slot gives the key line back four cells and its gap, and at the tabbed tier
that is enough for `1-N seat` to stop shedding — `empty-tabs.txt` now names a key that always
worked and had been dropped for width since §9.29. Nothing was un-shed by hand; the ladder in
`hint.shed` is unchanged and simply has more room to not use.
**No new hues.** The border and the sides are `Rule()`, i.e. muted chrome; the legend keeps the
exact styles the mode word already had on the mode line, gate included. This change spends
*shape*, which the palette does not pay for.
### 9.45 the turn clock counted the operator's reading time as the vendor's work (2026-08-15)
**A gated column said `⋮ streaming 5m` while nothing was streaming.** The room was stopped on an
approval card, the vendor was blocked waiting to be told yes or no, and the five minutes on the
header were five minutes of a person reading a diff. The number was real wall clock and it was
still a false reading, because of the word it sat under: `streaming` is a claim that output is
arriving, and this column had a stopped process behind it.
That is the same failure `TestWaitingIsNotStreaming` was written for, arriving by a different
route. That test guards the WORD — a seat with nothing to show must not render like a seat that is
showing something. Nothing guarded the FIGURE under the word. A stopped seat wearing a moving
seat's clock is the honest-gauge rule broken in the one place §4a.1 cares about most: the
displayed value no longer describes the thing it is labelled with.
**So the turn clock splits in two, and both halves are measured.** The column header states the
VENDOR's own time — wall clock minus whatever of it the operator held — and the turn's separator
states the operator's share beside the number it came out of:
```
▸ 1 CC Claude Code ⠋ streaming 12s
ro:tools tokens
⚠ waiting on you: Write:
internal/council/clock.go
y approve n deny a stop asking
…
turn 1 ───────────────── you 4m48s
```
Twelve seconds of vendor, four minutes and forty-eight seconds of operator, five minutes of wall
clock — and the two figures add up to it, which is the property that makes the split worth
drawing rather than merely correct.
**The stopwatch runs over the SEAT, never over the card.** One assistant message can raise a
parallel batch of requests and each one blocks separately, so a seat can have three cards up at
once — and it is stopped ONCE. Summing the cards would bill one person's one wait three times. So
`PendingGate.StoppedAt` is when the seat stopped, not when the card was raised: the first card of
a stretch carries its own moment, every later card in the same stretch inherits that stamp, and
the stretch closes when the LAST card goes. Stamping each card with its own moment was the first
cut and it was wrong twice — the figure on screen jumped backwards when the first of two cards was
answered, and the charge lost every second before the newest card, because the stretch's start
left the queue with the card that owned it.
**Stretches accumulate across one turn, and reset with it.** `Column.GateWait` is the operator's
share of the turn in flight, `TurnRecord.GateWait` is the same fact for a turn in the transcript,
and `startTurn` files one and clears the other. A turn that asked three times reports all three
waits as one figure, because what it claims is the operator's share of THAT turn.
**Auto-approved calls contribute nothing.** `queueGate` answers three ways without ever drawing a
card — `autoApproveRoutine` (this shell command is routine), `isReadOnlyTool` (this tool changes
nothing), and `!Asking` (the user said stop asking) — and all three return before the stamp. Nobody
was asked, so nobody waited, and charging the operator for a decision they never saw would be the
room inventing a measurement.
**Zero and absent stay different, as they must.** The figure is a `runner.Span` rather than a
duration, for the argument already written on that type: a turn that raised no card is UNMEASURED
and renders nothing at all, while a card answered inside a second is a measured zero and renders
`you 0s`. `gate-clock.txt` pins the three states side by side — an open card counting, an absent
one, a measured zero — because each is only legible against the others.
**Two spellings, one fact, and the surface picks.** `waiting on you 4m48s` is the room's own
phrase, already on the approval card and on §9.40's `NEEDS YOU` strip, and it is twenty cells. A
three-up room at 120 columns gives each column thirty-six, where the long form is more than half
the width and pushes the separator's own clock and cost off the line. So the grid sheds the LABEL
and keeps the fact — `you 4m48s` — which is §9.18's order applied to a phrase instead of a name,
and the by-turn page, which is the full frame wide, says it whole. It is a width TIER rather than a
per-line measurement, the same way `stripHeader` and `stripBadges` choose a form: one rule a reader
can learn, instead of a cell that rewords itself when a neighbouring number grows a digit.
**Why the separator and not the chrome.** The live turn's separator carried its number and nothing
else, on the rule that a turn's clock and cost are in the header and the badge line and repeating
them a row later would say one thing twice. The operator's share is the exception, and it is there
because the chrome has no room for it: `▸ 1 CC Claude Code` and `⠋ streaming 12s` already spend
thirty-three of thirty-six cells, and the badge line's right edge belongs to the cost. The
separator is also where the figure lives for every turn already in the transcript (`historyMeta`),
so the live turn and the filed one state it in one place and one spelling. What it costs is that
the figure scrolls with the transcript — and what stays pinned in the chrome is the card itself,
which says the room is stopped on you for as long as that is true.
**Render stays pure.** The room stamps and the renderer subtracts, exactly as `Reattach.SavedAt`
does: `queueGate` stamps when the card goes up, `decideGate` and `dropGates` charge when it comes
down, and Render turns a stamp into an age against `State.Now`. An open card with no stamp — every
State a test types out by hand — adds nothing and does not make the span measured, because a
duration arrived at by arithmetic over an absence is the invented figure §4a.1 puts at the top of
the rejected list. `TestGateClockIsPureOverState` pins it, and no existing golden moved.
**The `--trace` line is deliberately unchanged.** `runner.TurnClock` already says a seat blocked on
an approval card is inside its `Stream` span, "because from the process's side that is exactly what
it is". That stays true: the runner cannot see the decision. A gate decision reaches a live seat
through `Session.SendAside`, which is documented as carrying "an interrupt, a gate decision, a
protocol reply" — undifferentiated bytes — so a gate span in the trace would need the runner to be
told what it was writing, which is a change to that package's shape rather than to this figure.
The room knows, and the room is where the split is drawn. Adding the span to the trace is a
separate measurement with a separate seam to build.
**Wall clock is still reported where wall clock is the claim.** `turnElapsed` — the by-turn page's
figure for how long the whole turn took — is unchanged and still selects the longest seat's raw
elapsed. That number is labelled as the turn's duration rather than as any vendor's, and the
operator's own reading time really is part of how long the turn took. The per-seat rule under it
carries the split, so the page states both without either one contradicting the other.
**Amended 2026-08-29: the word was left behind, and the split alone could not fix the reading.**
This section opens by naming the defect as a WORD problem — "`streaming` is a claim that output is
arriving, and this column had a stopped process behind it" — and then corrects only the number.
The result was `⠋ streaming 12s` on a seat with a blocked process behind it. Twelve seconds is the
honest figure and `streaming` is still a false claim, so a reader who scanned the header still read
a working seat. A corrected number under a wrong word is a wrong reading.
**While a card is up, the header says `needs you` and states no clock.** The word is
`needsYouWord`, which is `needsYouLead` in lower case rather than a second spelling of it. One state
now has one vocabulary in three registers: `NEEDS YOU` on §9.40's strip, `waiting on you` on the
card and on the long form of the operator's own figure, and `needs you` in the column header, where
every state word is lower case. The mark is `Warn`, the card's own glyph. The spinner is this room's
only moving cell and it means a turn in flight (§7.1 rule 4), so a spinner over a stopped process
makes the same false claim the word did. Colour is `SevWarn` and carries nothing extra: the phrase
survives `--ascii` and `NO_COLOR` on its own.
**The clock goes because neither figure is time spent in this state.** The vendor's twelve seconds
are frozen for as long as the card is up, so the number is not moving and it does not describe what
the seat is doing. The operator's four minutes are moving, and they already have one home — the
turn's own separator, where `historyMeta` states them for every filed turn as well. Printing them in
the chrome too would put one fact in two places on one screen. What the header loses is a figure
that had stopped; it returns the moment the card is answered, and
`TestWaitingOnYouIsNotStreaming` asserts that arm rather than trusting it.
**Nine cells, so no layout moves.** `needs you` costs exactly what `streaming` and `cancelled` cost,
which is the width `stripColumn`'s floor is derived from (`layout.go`). The strip header takes the
same substitution, because a folded seat can hold an unanswered card and a strip is the one width
where the reader has no card beside the word to read it against.
**Three conditions, and the queue is the only source.** `stoppedOnYou` requires an installed seat, a
turn in flight, and a card in `State.Gates` for that vendor. `State.gateStopped` answers the last
one, and it deliberately differs from `gateStoppedAt`: that function skips an unstamped card because
a duration derived from an absence is the invented figure §4a.1 rejects, while this one reports the
seat stopped, because the card's existence establishes that on its own. Every fixture in the package
is unstamped and every one of them draws the card, so a stamp-sensitive predicate would have put the
header back in contradiction with the card two rows under it.
**The by-turn page takes the same word.** `seatMeta` states one seat's turn on one rule, and a page
saying `streaming` while the grid says `needs you` would be two surfaces disagreeing over one queue.
It applies to a LIVE entry only. A filed record's phase is how that turn ended, and the queue only
ever describes now.
**`gated-vs-streaming.txt` is a new golden rather than an extension of an existing one.** It puts a
blocked seat and a working seat on one frame at the same wall clock, because the reader's question
is never "what does a blocked seat look like" — it is "which of these two is running".
`waiting-vs-streaming.txt` pins two claims about a VENDOR; this pins the same shape of claim about
the operator, and neither case is an edge of the other.
### 9.46 cursor hooks report a blocked action, but never report a request to a human (2026-08-16)
**Environment and evidence class.** The survey drove `cursor-agent` **2026.08.11-e8db854** on
Windows 11, from a PowerShell parent. This build is NEWER than the build every other cursor record
in this document pins (`2026.08.04-aaa8809`), so read the tables below as the current build and the
older records as the older build. The event catalogue and the configuration paths come from a
source read of the installed bundle. Every claim about what fires comes from a live run. The rig
put a recorder on each event. The recorder appended the raw stdin payload, a UTC timestamp and the
leading bytes to a per-event log.
**Why the survey ran now.** §1 recorded the needs-input seam as a watch item and named Hooks as the
supported surface for it. §8's roadmap items 3 and 5 carry the same item. Nobody had asked the
narrower question: does any hook event mark NEEDS INPUT, as opposed to marking the absence of
completion? The answer decides whether the future needs-you strip can source this vendor.
**The catalogue, from the installed build.** `index.js` holds the event registry as a single map.
It carries **21** events:
```
beforeShellExecution beforeMCPExecution afterShellExecution afterMCPExecution
beforeReadFile afterFileEdit beforeTabFileRead afterTabFileEdit
stop beforeSubmitPrompt afterAgentResponse afterAgentThought
sessionStart sessionEnd preCompact subagentStart
subagentStop preToolUse postToolUse postToolUseFailure
workspaceOpen
```
Before this survey the repository knew three of these names, plus `preCompact` named but never
used.
**The Claude compatibility map has exactly two holes, and they are the two that matter.** The same
file maps Claude Code's hook events onto cursor's own, for the imported-configuration path. Eight
map across. Two map to `null`, and the bundle lists them together as the unsupported pair:
| Claude Code event | cursor event |
|---|---|
| `PreToolUse` | `preToolUse` |
| `PostToolUse` | `postToolUse` |
| `UserPromptSubmit` | `beforeSubmitPrompt` |
| `Stop` | `stop` |
| `SubagentStop` | `subagentStop` |
| `SessionStart` | `sessionStart` |
| `SessionEnd` | `sessionEnd` |
| `PreCompact` | `preCompact` |
| **`PermissionRequest`** | **`null`** |
| **`Notification`** | **`null`** |
`Notification` is the Claude event that fires when the agent waits for a person. `PermissionRequest`
is the other one. Cursor's own catalogue has no equivalent of either, which is why the import drops
them rather than renaming them. This is the verdict in one line, read off the vendor's own table.
**Where cursor-agent reads hook configuration.** The loader reads seven paths. Four are cursor's
own and three are Claude's:
| scope | path on Windows |
|---|---|
| enterprise | `C:\ProgramData\Cursor\hooks.json` |
| team | `\.cursor\managed\active-team-hooks\hooks.json` |
| user | `~\.cursor\hooks.json` |
| project | `\.cursor\hooks.json` |
| claude user | `~\.claude\settings.json` |
| claude project | `\.claude\settings.json` |
| claude project local | `\.claude\settings.local.json` |
The three project-scoped entries sit behind a boolean in the loader. The loader also refuses any
config path that contains a symlink. The project scope is what let this survey run without changing
the operator's own configuration for most of its arms.
**What fires on which path.** Two trials per arm unless the table says otherwise. A dash means the
event did not fire on any trial.
| event | print mode (`-p --trust`) | ACP, project scope | ACP, user scope |
|---|---|---|---|
| `workspaceOpen` | fires | — | — |
| `sessionStart` | fires | — | — |
| `preToolUse` | fires | — | fires |
| `beforeShellExecution` | fires | — | fires |
| `afterShellExecution` | fires | — | fires |
| `postToolUse` | fires | — | fires |
| `postToolUseFailure` | fires, on a hook denial only | — | not observed |
| `sessionEnd` | fires | — | — |
| `afterAgentThought` | fires, and kills the turn | — | not tested |
| `beforeSubmitPrompt` | — | — | — |
| `afterAgentResponse` | — | — | — |
| `stop` | — | — | — |
| the other nine | — | — | — |
**Two results in that table are new, and one confirms an older record.** First, **ACP honours the
user scope and ignores the project scope.** A project-scoped config fired nothing at all on ACP,
over two trials, not even `sessionStart`. The same file fired eight events in print mode. This
agrees with the loader's gate and with `PARITY.md`'s row that the ACP protocol has no
workspace-trust step. Second, **ACP fires no lifecycle event.** `sessionStart`, `workspaceOpen` and
`sessionEnd` are print-mode only, so a needs-input consumer on ACP gets tool events and nothing
that brackets the session. Third, `afterAgentResponse` still does not fire on ACP, which confirms
§7.16's 2026-08-15 amendment at this newer build.
**What a blocked moment looks like in bytes.** A hook that returns `permission: "deny"` replaces
`afterShellExecution` and `postToolUse` with one `postToolUseFailure`. That payload is the only
affirmative block marker on the whole seam:
```json
{"tool_name":"Shell","error_message":"Command execution was blocked by a hook: telltale-seam-deny
…","failure_type":"permission_denied","duration":0,"tool_use_id":"…","is_interrupt":false,
"hook_event_name":"postToolUseFailure","cursor_version":"2026.08.11-e8db854"}
```
`failure_type` is a free string, not an enum, and the hook path writes only two values into it:
`permission_denied` when a hook denies, and `error` when a fail-closed hook errors. Both describe a
refusal that already happened. Neither describes a wait.
**The awaiting-human moment is where the seam collapses.** Two arms produced a real human decision
point. In print mode a hook returned `permission: "ask"` with no person present. On ACP the client
answered `session/request_permission` with `reject-once`. **Both produced the success-shaped event
sequence** — `preToolUse`, `beforeShellExecution`, `afterShellExecution`, `postToolUse` — with no
`postToolUseFailure` anywhere. The only difference from an allowed command is one field:
| arm | `afterShellExecution.output` |
|---|---|
| allowed | `"[ERROR] - (starship::print): Under a 'dumb' terminal (TERM=dumb).\n\r\nseam-probe\r\n"` |
| `ask`, nobody to ask | `""` |
| human rejected over ACP | `""` |
An empty `output` is also what a silent successful command writes. So the seam encodes "a person was
asked and said no" and "the command printed nothing" with the same bytes. That is §4a.1's
zero-versus-absent collapse, arriving from the vendor rather than from a render path, and it is the
reason a needs-you strip cannot be built on these events.
**The affirmative signal exists, on the other seam.** The ACP wire carries it plainly. It is a
blocking JSON-RPC request from the agent to the client, and the turn stops until the client answers:
```json
{"jsonrpc":"2.0","id":0,"method":"session/request_permission","params":{"sessionId":"…",
"toolCall":{"toolCallId":"…","title":"`echo seam-probe`","kind":"execute","status":"pending",
"content":[{"type":"content","content":{"type":"text","text":"Not in allowlist: echo"}}]},
"options":[{"optionId":"allow-once","name":"Allow once","kind":"allow_once"},
{"optionId":"allow-always","name":"Allow always","kind":"allow_always"},
{"optionId":"reject-once","name":"Reject","kind":"reject_once"}]}}
```
This names the pending tool call, gives the reason, and lists the choices. It is everything a
needs-you strip wants. It is **not a hook**, and only the process that drives the ACP session
receives it. The council seat already reads it (`acpPermission` in `cursoracp.go`). A passive gauge
cannot, because there is no file and no second reader.
After the client rejects, the wire reports `tool_call_update` with `"status": "completed"` and no
`rawOutput`. That is the fourth shape §7.16's amendment predicted and could not attribute; this
survey attributes it, because the rejection here was the client's own and was known in advance.
**The verdict.** **No cursor hook event marks NEEDS INPUT.** The seam reports a completed refusal
(`postToolUseFailure`, `failure_type: permission_denied`) and it reports completion. It never
reports a wait. The vendor's own Claude-compatibility table says the same thing by mapping
`Notification` and `PermissionRequest` to nothing. For this vendor the needs-input signal lives on
the ACP wire, in `session/request_permission`, and it is available only to a seat that drives the
session. §8's roadmap item 3 should be read against that: cursor's entry belongs under the council
seat, not under the hook relay.
**Two traps for whoever builds on this seam.**
- **The BOM is on every event.** Every payload captured here began `EF BB BF 7B` — a UTF-8 BOM,
then `{`. `cursorhook.Parse` does not strip it, and does not need to today, because
`afterAgentResponse` is the one event that does not fire on the paths telltale drives. Any new
event routed into that parser fails on the first byte. `internal/cursorstatus/stdin.go` holds the
working strip.
- **A command hook on `afterAgentThought` kills the turn.** Three turns registered it and all three
died with `RetriableError: WritableIterable is closed` after the tool call. Six turns without it
completed. A compiled recorder in place of a PowerShell one changed nothing, so this is not hook
latency; the event fires inside the response stream and registering it breaks that stream.
**What is NOT claimed.**
- **The interactive TUI is unmeasured here.** This survey drove print mode and ACP only. §7.16's
per-surface table already records that the interactive console behaves differently, so the dashes
above are measured absences on two paths, not a claim about a third.
- **`stop` never fired on either path, and its interactive behaviour is unknown.** It is the natural
turn-complete event and it stayed silent, which is worth knowing before anyone designs against it.
- **The ACP rejection arm has one clean trial, not two.** The second trial lost its connection
mid-turn, a vendor-side failure this survey saw on other arms too. The wire shape matched on both;
the full event set was captured once.
- **No subagent, MCP, compaction or tab event was exercised.** They are listed above because the
build declares them, not because anything drove them.
**The operator's own configuration was restored.** Most arms used a project-scoped
`\.cursor\hooks.json` in a throwaway directory. The two ACP user-scope arms needed
`~\.cursor\hooks.json`, because ACP ignores the project scope. That file was copied first, the copy
was verified at SHA256 `3A9F05582D99DFEEB95E705559789F3B41D01DF1292F811D1A94834A54DFCB3C`
(146 bytes), the test configuration preserved telltale's own `afterAgentResponse` entry, and the
original was restored and re-verified at the same hash and length. No credential store was read or
copied at any point.
### 9.47 the room raced fourteen times and could not say who won (2026-08-29)
`/arena` has been building a record since it shipped, and nothing could read it. Every race
leaves an `arena/t/` branch per seat and every adoption leaves an
`adopt/t-`; both outlive the room by design (§9.37's kept-until-deleted ruling),
and `arenaRaceNumber` already reads the first namespace to number the next race — "the refs
are the one record that shares the leftovers' lifetime". What the ROOM kept of a race was
`Column.Arena`, a per-turn fact the next dispatch clears, and `TurnRecord` never carried it.
So a repository holding fourteen races and nine adoptions could not answer *which seat do I
actually take*, and the operator answered it by reading `git branch` in a second terminal —
which is the one thing §9.17 says a command surface exists to remove.
**`/arena record` reads the refs and states the standing, one line per seat.** It is a verb
inside `/arena` rather than a new room word, and that is a budget decision with a name:
`refuseUnknownCommand` prints the whole vocabulary on one line against a hard width, and that
line's own comment records `/adopt` as "the last cheap one — the next verb has to find its
characters somewhere else". `/arena drop` had already established the shape (a sub-verb, the
exact form only, anything longer races as prose), so the record costs the refusal nothing and
the help panel nothing. It is taught the way `x`, `/adopt` and `/arena drop` are: by this
section, and by the notice the command itself prints.
#### Derived from the refs, never stored
**Nothing is written and no new file exists.** The obvious build was a counts file under
`~/.telltale`. It was rejected before it was written: `CLAUDE.md` enumerates the writes the
gauges are ratified to make — three relays and the event sink, each with a test pinning its
serialized form — and a fourth exception is an owner-level edit to that contract, not a
feature's side effect. A tally over refs the repository already holds needs no such grant,
and it cannot go stale against them, because it IS them. Two `git for-each-ref` scans, over
the two namespaces `freeAdoptBranch` and `arenaRaceNumber` already scan, through the same
`gitOut` argv.
**The read happens in the command handler and the page renders from State**, exactly as
`ArenaResult` is computed in `finishColumn`. A body whose content came from a subprocess
inside `Render` would make every golden depend on the repository the tests happen to run in,
which is the purity rule `TestRenderIsPure` exists to hold.
**The verb is not refused mid-turn**, and that is what separates it from the other two arena
verbs. `/adopt` and `/arena drop` mutate worktrees a race is writing; this one reads refs. A
record read during a race is a measurement of a moment already past, which is what every
other reading in this room is.
#### What the refs can say, and the three things they cannot
This is the honesty boundary of the feature, and the page states both halves rather than
implying the first.
- **They CAN say who entered a race and whom the operator adopted from it.** An adopt branch
exists only on an adoption that LANDED — `undoAdoptBranch` deletes the branch a failed one
cut — so a surviving `adopt/t-` is a merge that happened, on the operator's own
`y`.
- **They CANNOT say a rank, a phase word, or that a seat was cut with `x`.** Those are
turn-scoped and die with the room. **So this surface never claims a LOSS.** A seat that
entered a race the operator never decided is counted as UNDECIDED and reported beside the
rate, never inside it. A race with a give-up is exactly that case, and the give-up is the
most probable ending a five-seat race has (§9.37's 2026-08-17 amendment) — folding it into
a denominator would be the room scoring a seat for a race nobody judged, which is §9.22's
refused "cross-seat agreement mark" arriving through a side door.
**The narrower case is stated rather than solved: a seat CUT with `x` in a race the
operator then decided for somebody else counts as one that was not adopted.** The refs
cannot tell a stall from a worse answer, and no honest reading of them can. What answers it
is the WORD on the page. It is `adopted`, never `won`, and that is the whole mitigation:
`0 of 4 adopted` is literally true of a seat that stalled four times, and it makes no claim
about why. A column headed `won` would make one, on evidence that does not exist.
- **A dropped branch leaves the record.** `/arena drop` deletes an arena ref by design, and
`adoptSeat`'s own notice offers that drop as the next command — so the winner's arena
branch is the one most likely to be gone. Two consequences, both handled: an adopt ref is
treated as evidence the seat ENTERED that race, which is what keeps a rate off the far side
of 100%; and the page says it counts over the branches the repository STILL HOLDS, in a
line of its own.
**The window sentence is a LINE, not the rule's meta.** `labelRuleIn` drops a rule's meta
whole when the width will not take it — correct for a count, wrong for the sentence that
bounds the claim, which would then vanish exactly where the room has least room to make it.
That is the act ledger's own ruling on its retention line (§9.22, amended 2026-08-17), and
this page's claim is bounded the same way. The clipboard document carries it too, and there
it matters more: a table pasted into a review a week later has nothing else saying these
counts were ever bounded.
**Only the refs this room minted are counted.** `arena/t/` and
`adopt/t-` with `freeAdoptBranch`'s numeric collision suffixes, parsed back
against the same two functions that write them. A hand-cut receipt is refused — the real one
is `adopt/t9-claude-helpers`, which the first live adoption left behind before the 2026-08-11
ruling gave the verb a spelling of its own. That is an UNDERCOUNT, said out loud here rather
than papered over, and it is the honest direction to be wrong in: a looser parse would credit
a seat for a branch somebody merely named after it. It is `dropRacer`'s judgement about paths
— "no state this room's arena created can have that name" — applied to refs.
#### The three renders, and the one figure this page is entitled to compute
- **A seat with no ref at all is ABSENT: `never raced`.** Not 0%. §4a.1's founding rule, on
the surface where a zero would be read as a verdict about a vendor rather than as a count.
- **A seat whose races were all undecided has NO RATE**, and its races are reported as
undecided. Inventing a denominator out of them is the same error one step down.
- **A seat the operator decided against is a MEASURED zero and prints one: `0 of 4 adopted
0%`.** The distinction between that line and the one above it is the whole feature.
**The rate never appears without its count, and the count comes first.** The two counts are
what was measured; the percentage is arithmetic over them, which §7.12 names as the one kind
of computed figure this product may show — "telltale's measurement of telltale's own
observations" — and that carve-out is conditional on the reader being able to check the
division. A bare `67%` would be the claim with its evidence removed, which is the same defect
as a total printed without its window (§7.15).
The seat name takes that seat's own hue and nothing else on the line does (§9.28's ratified
exception), for the turn page's reason: this is a stack of seats in one column, so position
answers nothing about who is being described. Every distinction the page makes is carried by
a WORD first — `never raced`, `no decided race`, `undecided`, `adopted` — so `--ascii` and
`NO_COLOR` lose nothing.
#### The keys, and what the page deliberately does not get
`t` closes it, because `t` is already what this room means by "give me the grid back" from a
full-frame body. A key of its own would be a second thing to remember for one act, and no key
at all would be a body reached by a typed command with no keyed way out — the help panel's
missing `?` with a whole surface behind it. The mode word is `RECORD`, against `TURN` and
`ACTS`: §7.8's always-on statement of what is on screen has to tell the room's three
full-frame bodies apart, and this is the one whose subject is neither a turn nor the keymap.
It carries no coordinate because it has none.
`y` takes the page, on §9.22's own argument — a copy key that took the focused column's reply
from behind a body the reader is looking at would break the one claim that earns it a footer
cell. `Y` is the same document for the same reason the page's is: there is no per-seat focus
for a narrower key to address.
**No scroll cell, and no scroll keys.** The record is one short line per seat, so the only
geometry that can clip it is the height floor, and there `recordCell` draws the overflow
marker instead. A footer naming arrows that move nothing is the false promise §7.8 forbids —
the same reason `f` and `tab` are absent and are swallowed rather than left to change the
grid invisibly.
**Deliberately not built, and each one is a ruling rather than a backlog item:**
- **Elo, or any rating.** The sweep that proposed this feature said it plainly and it is
right: Elo is overkill for one operator's vote volume, and a rating is a number with no
measurement under it. A plain adopted-of-decided tally is the honest shape.
- **A blind-review mode**, the other half of the sweep's candidate. §9.34 states that the
PERSON is never blinded — "columns stay labelled by vendor; the blind applies to what the
models read" — so an opt-in user-blind arena inverts a stated position and is an owner's
ruling, not a builder's. It is not built here and nothing here assumes it.
- **Anything cross-repository.** The record is one repository's refs, because that is what a
room is pointed at and what `/cd` moves. A tally across every repo the operator has ever
raced in would need a store, which is the write this whole section refuses.
- **A rank or a phase in the tally.** They are not in the refs. Carrying them would mean
filing `ArenaResult` into `TurnRecord` and persisting it, which buys a richer number by
taking on the store — and the number it buys is one the operator's own adopt decision
already summarises.
Verified offline. `record_test.go` pins the tally's arithmetic against hand-written ref lists
(one race counting once however many refs it left, an adopt ref proving entry, the numeric
collision suffix not double-counting), the refs this room did not mint being ignored, zero and
absent staying apart, no rate ever printing without its count, an undecided race never
reaching a denominator, the window sentence surviving the narrow width, an unreadable record
rendering as unavailable rather than as an empty one, `t` giving the grid back without opening
the turn page, the clipboard document agreeing with the screen and carrying the window, and
the verb taking only its own word — measured in a read-only room, where a longer draft is
refused by name and therefore provably reached the race path. `arena-record.txt` and its
`--ascii` twin are the frame. The one test that touches git builds its refs with `arenaBranch`
and `adoptBranch` in a temp repository, so the parser is asserted against the functions that
write the names rather than against the strings the test types. No test here spawns a vendor.
**Nothing in this section is a claim about vendor behaviour, so no live vendor run is owed.**
What IS owed is one live open against the reference box's own leftovers — 27 `arena/t`
branches and the `adopt/*` refs beside them are recorded in §9.37 — to confirm the counts a
real pile of refs produces and that the page reads at the room's own geometry. Stated here
rather than implied paid.
### 9.48 the race said what changed and never whether it worked (2026-08-29)
`/arena` measures everything about an attempt except the one thing an operator adopts on.
It reports what each racer CHANGED — the live stat, the settled `git diff --stat`, the full
patch, the commit receipt (§9.37) — and `/arena record` reports which seat the operator TOOK
(§9.47). Neither says whether the attempt WORKS. The room's own founding note admits it: rank
is arrival order, and the only clock that ranks a race is the room's. So the operator answered
the question by hand, once per seat, in a second terminal — `/cd` into each kept worktree and
run the same command — which is the act §9.17 says a command surface exists to remove.
**`/arena check ` names one command; every racer runs it in its own worktree, and
each attempt's block says PASS or FAIL from that run's real exit code.**
#### The grammar, and the two shapes it is not
It is a sub-verb inside `/arena`, for §9.47's budget reason: `refuseUnknownCommand` prints the
whole room vocabulary against a hard width, and that line's own comment records `/adopt` as
"the last cheap one — the next verb has to find its characters somewhere else". `/arena drop`
and `/arena record` had already established the shape, so a third costs the refusal nothing
and the help panel nothing. It is taught the way they are: by this section, and by the notice
the command itself prints.
**It is NOT a file in the repository, and that is a ruling rather than a preference.** The
obvious build was `.worktreeinclude`'s sibling — a `.arenacheck` the racer trees inherit — and
arena.go's own seeding doc already refuses exactly that shape: agent-deck pairs seeding with
repo-carried setup scripts, and council took "copy only, never execute", because a repository
that can run a command on the machine by merely CONTAINING a file is a different product with
a different threat model. A command a person typed into their own room is that person's act. A
command a clone brought with it is not. The parked byte-level trust question is untouched by
this feature, which is the point of not touching it.
**It is NOT a new room word.** `/check` would have cost the refusal line a re-wording and,
worse, a second meaning for a word this codebase already spends: the write gate, the gate
cards, `gatehook.go`. Two facts cannot wear one word (§9.13) — which is also why the verdict
is spelled `FAIL` and never `failed`. The room already spends `failed` on a phase, and a seat
that finished cleanly while its check exited 2 is a different fact from a seat whose process
died.
**The cost of taking free text, stated the way `parseArenaDrop` states its own.** This is the
one `/arena` sub-verb that cannot close its grammar with a length cap, so a brief opening with
the word `check` is at risk. What protects it is a PATH lookup on the first word: a draft whose
first word after `check` is not a program this machine can run is refused by name, handed back
to the composer, and neither raced nor set — nothing spawns and nothing is billed. The narrow
case that survives is a brief opening `check `; the notice names
exactly what was set, and `/arena check off` takes it back in one line. A path-bearing first
word (`./scripts/check.sh`) skips the lookup on purpose, because it is resolved against the
RACER's tree rather than against the room's own directory.
**The command is room state and survives `/cd`**, which is `/write`'s rule rather than an
oversight: posture is room state and moving the workspace does not quieten it either. The
mitigation is that every result names the command it ran, so a command left over from another
repository is visible in the verdict rather than assumed behind it.
**It does NOT survive the room.** `room.json` holds session ids and a workspace — keys and
numbers, never content (`CLAUDE.md`'s read/write boundary) — and a command is neither. The
consequence is deliberate rather than reluctant: a saved command would run in a session whose
operator never typed it, which is the one property the "a person typed it into their own room"
argument above rests on. Re-naming it is one line.
#### Four rulings, and each is a line the code may not cross
- **The exit code is the only source.** PASS is exit 0; FAIL is any other code; both are read
back from the process. Nothing here infers a verdict from output, from the diff, from a
duration, or from a model's opinion. **An LLM judge is refused by ruling twice over** —
§9.2's refusal of "a ranking stage, a chairman, or any synthesis hop" and §9.44's declined
cross-seat quality mark — and it stays refused even wearing an estimate's `~`, because `~`
marks a figure telltale COMPUTED and an opinion is not a computation. A measured exit code is
the opposite case, and it complies with ADR-001 exactly as the diff does.
- **A command that could not run is its own state.** A missing binary, a tree that could not be
entered, a run the deadline stopped: none of them is a FAIL. They render as
`check unavailable: `, because "this attempt failed the check" and "nothing measured
this attempt" are the degraded-vs-zero distinction §4a.1 exists to keep apart. `Exited` is a
field of its own and is the ONLY gate on a verdict, so a check record's zero value can never
read as a pass.
- **No command named is ABSENT, and absence draws nothing at all.** Not a dash, not a 0, not a
pending word. It is nil `Seed`'s rule on a second field: a room that was never asked for a
check has no check to report. Absence is a whole-race property rather than a per-seat one —
the command is the room's, so either every racer with a tree carries a check or none does.
- **The room captures no output.** The exit code is the whole claim. A failing command's stdout
would put unredacted subprocess text on a screen whose vendor streams all pass through a
`Redactor` first, and reading it would need the scroll surface the gate cards were already
refused. The block names the command and the worktree; the operator re-runs it there.
#### Where it runs in the finish line, and what that costs
**The check runs LAST, after the diff is read and after the attempt is committed, and the
order is a ruling.** A check that ran first would park its own build output on the arena branch
wearing the racer's name — the false receipt, §4a.1 pointed at a write. So nothing a check
writes can reach the stat, the patch or the commit.
What it CAN reach is a later `/adopt`, which commits a dirty attempt before merging it. That is
said rather than swept up: the run is bracketed by two `git status --porcelain` reads, and a
check that found a clean tree and left a dirty one prints one sentence saying so. The room does
not reset a tree the operator did not ask it to reset, and `/adopt`'s own y/n card names every
command it will run before it runs one. A tree that was ALREADY dirty, and a state that could
not be read, both claim nothing — the field reports a measured change, never the absence of
one.
**It runs off the render loop**, for the reason `arenaSetup` moved off it (§9.37's 2026-08-17
amendment): a `go test` inside `Update` is a room that draws no frame and reads no key for
minutes. `finishColumn` queues the run, `Update` drains the queue on the event batch and again
on the spinner tick, and the result arrives as a message — `arenalive.go`'s pattern, with one
difference that matters.
**The check's lifetime is the ROOM's, not the turn's.** The last racer landing ends the turn,
and a check that started at that moment must still be allowed to finish and report onto a
column that is still on screen — so the run hangs off `roomCtx`, and teardown kills it for the
same reason quitting kills every other child this room started. A stale run is dropped by
comparison (the vendor and the turn number), never by hoping the timing worked out.
**The run is a process TREE, and it is contained like one.** `runner.RunContained` is a new
one-shot sibling of `runner.Start` — no streaming, no parsing, no clock record, because a check
is not a turn — and what it takes from `Start` is the Windows job object and the unix process
group. That is not belt and braces: `proc_windows.go` exists because `codex` resolves to an npm
`.cmd` shim, and `npm test` has exactly that shape, so a deadline that killed only the direct
child would leave the real work running two processes down with nothing on screen to say so.
`contained_test.go` asserts it the way the vendor path is asserted — a helper that spawns a
grandchild writing to a file, and a cancel that has to stop the FILE growing. Both of `Start`'s
own limits carry over unchanged and are recorded there: the microsecond window before the group
is assigned, and unix's need to be killed on the way out.
**A cancelled turn runs no check.** ctrl+c is the operator saying stop, and a room that
answered it by starting a subprocess per seat would be ignoring the one act it exists to obey.
A seat cut on its own with `x` is NOT that case: §9.37's give-up ruling says a given-up seat
lands like any other finisher, and its partial work is as worth checking as its diff is worth
reading.
**The bound is ten minutes, and the anchor is this repository's own suite.** `CLAUDE.md`
records `go test ./...` at ~455s locally and ~4m22s in CI, which is about the slowest command
an operator would name here; ten minutes is past twice that. The margin is the decision rather
than the figure — what this bound must never do is kill a run that would have finished, since a
killed run yields no exit code and therefore no verdict at all. A run the clock stops is
reported as unavailable with the bound named.
**Checks may overlap, and that is stated rather than serialised.** Seats land at different
moments, so in practice they stagger; two that land together run together. The worktree adds
are serial because they contend for the repository's own refs (§9.37), and nothing analogous is
true here — each check runs in its own tree. Serialising them would make one slow check hold
every other seat's verdict, and the room measures nothing about the machine's capacity, so the
bound would be chosen from nothing.
#### Deliberately not built
- **Any LLM judge**, per the ruling above. It is the half of the sweep candidate that fails
§9.2 and §9.44, and a mark does not soften it.
- **Repeat sampling** (N attempts of one seat), the candidate's other half. §9.37 rules every
attempt a fresh session across all seated vendors and rules a one-seat race an ordinary turn
in a worktree; N-of-one contradicts that as written and collides with the
`arena/t/` identity scheme `lifecycle.go` re-mints from refs. It needs an owner's
ruling rather than a builder's, and nothing here assumes one.
- **Per-attempt cost and diff-size columns.** They already ship — the stat is §9.37's, and the
vendor-reported cost is `turnview.go`'s, absent where the vendor reports none.
- **A check that gates anything.** It reports; it does not refuse an `/adopt` and it does not
re-order a rank. The founding posture is offer-never-take, and a room that blocked an
adoption on its own reading of a test would be taking.
- **A check on an ordinary turn.** The question is whether an ATTEMPT works, and an ordinary
turn writes into the operator's own tree rather than into an attempt.
Verified offline. `arenacheck_test.go` pins the grammar (set, report, clear, and the refusal
that hands a brief back), one run per racer in that racer's own tree with the named argv,
absence drawing nothing, a cancelled turn queueing nothing while a given-up seat still runs,
the stale-message drops, the ordering — a check that writes into the tree reaching neither the
stat nor the commit — and the four renders, with `arena-check.txt` and its `--ascii` twin as
the frame. `contained_test.go` pins the run itself: an exit code coming back as itself, a
missing binary carrying no code at all, and a cancelled run stopping a GRANDCHILD. Colour is
asserted separately against the room's existing severity pair, per the goldens' own split. No
test here spawns a vendor: `countSpawns` stubs the check as its fourth spawn var, and
`TestMain` panics on any model-path run whose command this machine could actually resolve.
**One test runs a real process, and the process is this test binary.**
`TestPassAndFailComeFromARealExitCode` calls `runCheck` directly rather than through the
guarded var, with `os.Args[0]` and a helper `-test.run` as the command, and asserts that three
exit codes come back as themselves, that a missing binary comes back with no verdict at all,
and that a run which writes a file is measured as having written one. Without it the feature
would rest on a stub returning the answer it was handed, which pins the render and nothing
about the claim.
**The live half is owed.** No live race has run under a check. What is unmeasured is a real
vendor's attempt meeting a real suite in a real worktree, and specifically whether a check that
runs for minutes reads well on a column whose turn has already ended. One `/arena check` and
one `/arena` against a brief that changes files pays it. Recorded in `STATE.md` rather than
implied paid.
### 9.49 the patch was on screen and there was no way to say anything about it (2026-08-29)
`d` has flipped a racer's arena block between the diffstat and the whole patch since §9.37,
and the block prints the branch and the worktree path under it. What the operator could do
with any of that was read it. To say *this hunk is the problem* they retyped the hunk into
the composer by hand, and to open the tree they selected an abbreviated path off a
fixed-width column and repaired it in another terminal. The help panel, meanwhile, named
none of it: `d` shipped and no page ever said the key existed.
**Three keys close that, and the surface they make is the patch view itself.** A cursor
points at one hunk of the drawn patch, `D` quotes that hunk into the composer draft, and `o`
hands the operator the worktree — started in their own editor, or copied to the clipboard.
#### The comment is a DRAFT, and that is the whole design
The lane this was taken from (Vibe Kanban, Parallel Code, Sculptor) routes an inline comment
straight back to the agent as its next instruction. **This room cannot do that honestly.** A
race attempt is one-shot in a worktree: the seat that wrote the patch has finished, and
§9.37 rules every attempt a fresh session, so there is no conversation for a comment to
resume. A "reply to this attempt" control would therefore have to invent one — a new race, a
new turn dressed as a continuation, or a queue.
So the quote lands in the **live composer draft**, and nothing else happens. It is visible on
the next frame, it is editable with every key the composer already has, and `enter` — a
person, pressing a key — stays the only thing in this room that spends a quota. That is the
standing rule about queued drafts (one visible, editable draft, never auto-send) reaching a
feature that was invented to break it, and it is asserted as a spawn count and a turn counter
rather than as prose: `TestDQuotesIntoTheDraftAndSpawnsNothing`.
Two consequences are stated rather than hidden. The next brief runs in the **room's**
workspace, not in the racer's worktree — so the fence names the worktree path, which is the
difference between a comment a seat can act on and one it can only agree with. And an empty
draft is seeded with the racer's own `@mention`, because silence routes to claude (§9.9) and
a comment about codex's attempt reaching claude would be the room misdelivering the
operator's words. The seed is one token, the footer's route cell resolves it on the very next
frame, and a draft that already says something is left alone — it may carry a route the
operator chose, and a second mention would be the room editing a line being written.
#### The quote is fenced as DATA, because that is what §9.15 settled
The fence is `quote.go`'s, for `quote.go`'s reason: this text goes from one program's output
into another language model's input, which is a prompt-injection path whatever wrote it. The
wording differs in exactly one way, and it matters — it names the material as **measured `git
diff` output** rather than as another participant's answer, because that is what it is. A
fence that mislabelled its contents would be the room narrating over its own evidence.
The **whole hunk** crosses even when the drawn frame cut it off partway. The cursor points at
a hunk and the hunk is the unit git itself framed; sending half of one because a render cap
fell inside it would hand a seat an incomplete measurement under a fence claiming to carry
git's output. What the frame decides is where the cursor may GO, not what a quote CONTAINS.
A quote that would push the draft past the composer's own cap is **refused whole**, never
truncated — `paste.go`'s atomicity rule reached by a different key, and for `paste.go`'s
reason: the Antigravity seat takes its prompt on argv, so a draft past that size is one no
seat could be handed.
#### The cursor never scrolls, and that is a refusal rather than a limit
The drawn patch is capped at `arenaDiffScreenLines` and does not scroll. A cursor free to
walk past that cutoff would need a second scroll surface inside a column — the device this
product already refused for the gate card's preview — so **the cursor points inside the frame
the reader can see and refuses to leave it by name.** Both ends say which end they are, and
the far end says *the cursor does not scroll* and names the two routes to the rest of the
patch the cutoff line already carries (`y` copies the whole diff, the worktree holds it).
The anchor is a **hunk**, not a rendered line, and the reason is that a rendered line has no
name anything can resolve: "line 412" is a position in a capped render of a one-megabyte
diff. A hunk header is git's own anchor, printed in the patch, and it survives the cap, the
yank and the quote.
The parser reads git's framing and no heuristic. Inside a hunk every line is prefixed by git
(space, `+`, `-`, `\`), so a line that is not prefixed is not in the hunk — which is what
makes it safe against its own input. A patch is whatever five language models wrote into five
worktrees, and this repository's own diffs routinely ADD lines beginning `+++`, `---` and
`diff --git`; a parser that scanned for those anywhere would cut a hunk in half at the first
of them.
#### The keys, and why none of them is new vocabulary
- **`D`** is SHIFT on the key whose surface it acts on, which is the spelling `T` already
earns beside `t` and `Y` beside `y`. It costs the help panel no row of its own for `T`'s
reason: it is taught on the row that teaches `d`.
- **`[` and `]`** step the cursor. These keys have always meant *step one unit of whatever
the body is showing* — a turn in the grid, a page in the by-turn projection (§9.22) — and a
patch is a body with a unit of its own. A third reading of one motion, not a fourth
binding. The footer's cell says `[ ] hunk` while a patch is open, because §7.8's contract
is that the mode line announces what every key means on every frame.
- **`o`** raises the worktree card. A key rather than a room command for `c`'s reason: no
vocabulary leaves the composer.
The cursor arms with the patch and costs no key at all: `d` opens the patch and puts the
cursor on the first hunk, so the patch view IS the review surface rather than a mode a reader
has to discover before the feature exists for them. Its mark is the room's own `▸` — reuse,
not a collision, because that glyph already means *this is what the keys address*, which is
exactly what it means here. It is drawn only on the FOCUSED column, since that is the column
`D` reaches.
**Both keys refuse over a full-frame body** — the turn page, the act ledger, the arena record.
`y` can follow a page because a page IS a document (§9.22, §9.47); these two cannot, because a
page has no hunk and no worktree, and a projection whose whole point is that the turn is the
unit deliberately has no per-seat focus for a narrower key to aim at. `D`'s claim is *the hunk
under `▸`*, and a claim a reader cannot check against the screen is the one thing this room
does not print.
#### `o`: the one process council starts that no vendor asked for
Spawning an editor is council's loud exception and it is taken deliberately. The gauges spawn
nothing; council spawns vendor CLIs already, and this is a **stricter** trigger than any
vendor spawn in the package has: two keys, in front of a card that names the program.
- **$VISUAL, then $EDITOR, then nothing.** An unset pair renders as unset and the card offers
the copy instead. Falling back to `notepad` or `vi` would be the room inventing a value the
operator never gave it — §4a.1 on a setting rather than on a gauge.
- **The whole variable is tried as a program first.** `C:\Program Files\…\code.cmd` is one
program with a space in it, and splitting on whitespace before asking the operating system
would turn the operator's own setting into a program named `C:\Program`. Only when the
whole string does not resolve is it read as `program arg arg`.
- **The card names the program and `y` starts THAT program**, resolved when the card armed
rather than when the key was pressed — `adoptOnto`'s contract, applied to a spawn.
- **No stdio is wired**, because the room owns the terminal. The honest consequence is that a
terminal editor opens where the operator cannot see it, and the notice says so rather than
letting them discover it. What is REPORTED is exactly what `Start` measured: the process
began, or the operating system said why it could not. Whether the editor then drew a window
is not observable from here, and the notice does not claim it — `yank`'s rule, one keystroke
over.
- **`c` copies the FULL path**, through the same native-helper-then-OSC-52 path `y` uses, and
a cancel writes nothing at all: an empty OSC 52 write is how a clipboard is CLEARED, so
"nothing happened" must not be spelled the same way as "your clipboard is now empty".
`startEditor` is a package var for the reason the other three spawn vars are: `TestMain`
makes it fail closed for the whole suite, and `countSpawns` stubs it with the vendors. A test
that pressed `o` then `y` on a box where `$EDITOR` resolves would otherwise open a window on
the desktop of whoever ran `go test`.
#### What the help panel had to give up
The panel's budget is hard (16 rows) and the arena keys had nowhere to come from, so the
line-wise and screenful scroll rows were merged into one. That merge is the category rather
than a saving — the bar this panel's own budget comment sets: they are one act at two scales,
which is the argument `f / t / T`, `y / Y` and `g / G` are each already merged on, and the two
rows had come to restate each other's "in compose too" clause besides. The row it freed pays
for `d / D / o`, and `d` is documented for the first time.
Verified offline. `review_test.go` pins the parser against a patch whose hunk body carries
three lines that are file headers in every other context, a deletion's `/dev/null` never
rendering as a filename, absent staying apart from hunk zero, the cursor stopping at the drawn
frame with the reason named, `[` and `]` yielding the turn hop when no patch is open, the mark
appearing only on the focused column and surviving `--ascii`, the footer cell following the
body, `D` spawning nothing and starting no turn, the quote following the cursor rather than
the column, the route seeded only into an empty draft, an over-large quote refused whole, four
refusals each saying their own reason, both keys refusing over each of the three full-frame
bodies, the card copying the path without spawning, `y` starting the program the card named in
the racer's own tree, an absent `$EDITOR` said rather than guessed, a stray key touching
neither clipboard nor process table, and the three keys named above the help panel's fold. `arena-review.txt` and its `--ascii` twin are the frame. No
test here spawns a vendor or an editor.
**Nothing in this section is a claim about vendor behaviour, so no live vendor run is owed.**
What IS owed is one live open: `o` then `y` against a real `$EDITOR` on the reference box, and
`D` on a real race's patch through to a dispatched turn. Stated here rather than implied paid.
### 9.50 the codex seat gets a second protocol, and the sandbox re-measurement that kept it unseated (2026-08-29)
§9.33 drove `codex app-server` on 2026-08-15, recorded a warm turn at **1.44 s** against
`codex exec --json`'s 5.33 s, and stopped: *"This is measurement only. It authorises no seat
change."* `STATE.md` carried the standing caution beside it — the Windows sandbox finding does
not port to that path, and **any seat move re-measures rather than inherits**. This is that
re-measurement, and the build it authorised.
**The verdict is split, and both halves are measured.** The protocol ships: a second
`runner.Protocol` beside the ACP one, parsed, tested, and pinned to a real capture. The SEAT
does not move: the room still dispatches codex through `codex exec --json`. What decided that
is not caution — it is a liveness defect measured on the new path that #311 had just cleared
off the old one.
#### Version pinned first, and the surface was driven before it was believed
Everything below is **codex-cli 0.149.1** on Windows 11 (`codex --version`), 2026-08-29 — the
same build #311 re-measured the `exec` seat on the same day, so nothing here is confounded by a
version difference between the two surfaces. `codex app-server generate-json-schema --out `
wrote the whole protocol at this build; every shape below was then either captured live or is
labelled a schema read. **Eight billed turns**, one to three per arm, in throwaway directories,
with files checked on disk rather than read out of the model's reply. Prompts are
**brief-shaped**: they open with `brief.go`'s own `--- operating context …---` fence, because
§9.39 records a seat that shipped broken for a day behind a green live test whose prompt began
with a letter.
#### The handshake, and what a warm thread actually costs
| stamped line | arm A | arm F |
|---|---|---|
| spawn returned | 0.035 s | 0.151 s |
| `initialize` response | 1.091 s | 0.673 s |
| `thread/start` response | 1.986 s | 1.032 s |
Two round trips, and the spread is this box rather than the vendor: five MCP servers start on
the same path and a `sessionStart` hook runs, both of them the operator's own configuration.
§9.33's warning stands unchanged — **that figure must never be quoted as a property of the
vendor.**
Then the two turns of arm A, one thread, one process, read-only posture:
| turn | sent | `turn/completed` | total |
|---|---|---|---|
| 1 (three shell attempts, hooks, MCP startup) | 1.986 s | 34.771 s | 32.785 s |
| 2 (a question only turn 1's history answers) | 34.771 s | 36.257 s | **1.486 s** |
Turn 2 was asked what file turn 1 had been told to create and answered `wrote-ro.txt`, from the
same pid. **1.486 s reproduces §9.33's 1.44 s at a newer build**, so the prize is real and it is
now measured twice, fourteen days apart, on two builds.
**The linger question re-opens and closes in the vendor's favour.** `codex exec` prints its last
line and holds the process open ~4 s (§9.33's 2026-08-16 amendment). Here there is nothing to
settle around: turn 2's `turn/start` went out on the same millisecond `turn/completed` landed.
The shutdown fact runs the other way and is the one worth carrying: **closing stdin does not
reliably stop this server.** Four runs exited in 1.5–3.3 s; one was still alive 15 s later and
had to be killed. The caller owns the kill.
#### The sandbox, re-measured on its own surface
`thread/start` takes a `sandbox` parameter whose enum is spelled identically to `exec`'s `-s`
values. That spelling is a coincidence of naming, not a shared mechanism, and the results are
not the same.
| probe | result |
|---|---|
| `read-only`, write via **direct `cmd.exe`** | **DENIED** — `Access is denied.`, exit 1, `status:"failed"`, no file on disk |
| `read-only`, write via the router's default `pwsh.exe` | **NO PROCESS** — `CreateProcessAsUserW failed: 5`, no `commandExecution` item at all |
| `workspace-write`, write inside the workspace | landed, exit 0, file on disk |
| `workspace-write`, write to `.git` (two independent turns) | **DENIED** both times, exit 1, no file |
| `windowsSandbox/readiness` (a free request, no turn) | `{"status":"ready"}` |
**The read posture enforces when a process starts.** The denied write is the strong arm, and its
own failure mode confirms it: the second call in that turn came back with cmd.exe's *own*
`'…' is not recognized as an internal or external command`, which a process that never started
could not have produced. So cmd.exe ran INSIDE the sandbox and obeyed it.
**And the seat still cannot be trusted to read, which is why it is unseated.** This protocol's
tool router wraps a shell command in `"…\WindowsApps\pwsh.exe" -NoProfile -Command "…"` unless
the model names a program itself, and pwsh cannot start under the Windows sandbox on this box —
the identical `CreateProcessAsUserW failed: 5` that 0.146.0 failed *everything* with. #311
measured `codex exec` retrying through cmd.exe and obeying the sandbox. On this path the retry is
the model's own move and **it did not always make it**: in two of three read-posture arms the
model abandoned the turn and reported it could not inspect. A read seat that cannot list a
directory is the 0.146.0-class defect arriving on a new path, and it is the exact failure the
`ro:enforced` badge would be sold against.
**Two things are NOT measured, and they are stated rather than inferred from the `exec` path's
answers — the whole reason this section exists is that those answers do not port:**
- **Writing outside the workspace.** The one arm that tried it wrote into a directory under
`%TEMP%`, which `workspaceWrite` permits by default (`excludeTmpdirEnvVar` defaults false).
The write landed and the arm proves **nothing**. Recorded as void, never as a permit. The
probe could not be re-aimed: this lane's scratch boundary is inside the temp root.
- **The `writableRoots` override on `.git`.** A per-turn `sandboxPolicy` naming `.git` was sent
and the model's own shell quoting broke the call before the sandbox saw it. So whether this
path can buy `.git` back — which `codex exec` measurably cannot on Windows (#311) — is open,
and it is the measurement the write posture's shape depends on.
#### What this protocol carries that `codex exec --json` hides
§9.33 named these and built nothing on them. They are now parsed, and still rendered nowhere.
- **`thread/tokenUsage/updated`** — `total` and `last` counts, plus `modelContextWindow`
(258400 on this account's model). That denominator is the one §3.2 records as absent from
every other codex surface; it is what would make a context percentage a read rather than an
invention.
- **`account/rateLimits/updated`** — `usedPercent`, `windowDurationMins`, `resetsAt` and
`planType`, per window, live on the socket. These are **quota** in §7.15's vocabulary — a
share of a window that resets — and never spend, and nothing here converts one into the other.
- **`hook/started` / `hook/completed`** — the hook's id, event, source path, source and
`durationMs`. Dropped by the adapter, and deliberately: the one hook captured was the
operator's own `sessionStart` script, and rendering somebody's local configuration as council
activity is the machine-specific claim §9.33 already warned this figure must never become.
- **Typed shell items** — `commandExecution` carries `command`, `cwd`, `processId`, `status`,
`exitCode`, `durationMs` and `aggregatedOutput`. Richer than the `exec` stream's item, and
`status:"failed"` is load-bearing here in a way it is not there: it was captured on a command
that never started, where there is no exit code to read.
**A free read surface, which is new and worth naming.** `account/rateLimits/read`, `hooks/list`,
`permissionProfile/list` and `windowsSandbox/readiness` are ordinary requests that answer
**without starting a turn** — driven live, no model spend. Nothing consumes them yet.
#### The trap this wire carries, measured rather than assumed away
**The deltas and the completed item are the same text.** Across four `agentMessage` items over
two turns, the concatenated `item/agentMessage/delta` payloads equal `item/completed`'s `text`
byte for byte, every time. This is §9.6c's whole-message repeat, present on a surface nobody had
parsed. A reader that took both would print every answer twice. The adapter consumes the deltas —
streaming is what the column is for — and spends the completed item as a **separator only**, per
item rather than per turn, because one turn carries several complete messages and a per-turn
guard would pass the first and fail the second. An item that never streamed still prints its
text, which is the safety net the ACP seat explicitly does not have.
#### The fork, taken conservatively
§9.36 ruled the Cursor seat's equivalent fork **WHOLESALE** under one-attempt probation, and the
numbers there justified it: the old path was ~13 s per turn against 1.1–1.8 s, with nothing left
to fall back to that anyone would want. **This fork is not that one**, and the difference is on
the safety side rather than the speed side.
A cutover here would move the codex seat onto a path where, at this build, a read-posture turn
was measured **failing to inspect at all** in two of three arms. The old path does not have that
defect — #311 measured it away the same day and paid for the badge. Trading a 5.33 s turn for a
1.49 s turn that sometimes cannot run a command is not the trade §9.36 made; it is the inverse
of it. And the two measurements the write posture would need — the outside-workspace boundary
and the `.git` override — are the two this lane could not close.
So: **dual, with the default seat unmoved.** The protocol, its parser, its refusals and a
version-pinned real capture all ship. `vendors.CodexAppServer` implements `Conversational` and is
deliberately absent from `Registry()`, with `TestTheRoomStillSeatsTheExecCodexAdapter` pinning
that absence as a decision rather than an oversight — it is the test that fails, and is read, on
the day somebody registers the seat.
**What the flip needs, named so the follow-up does not rediscover it.** One line maps
`model.VendorCodex` to this type; the seat then drives through `StartRPCSession` exactly as the
Cursor seat does. What must be re-measured FIRST is the read posture's **liveness** — a turn that
lists a directory and reads a file, under the sandbox the badge claims — because that is the
property the badge sells and the one this path was measured failing. The write posture needs the
two open sandbox arms above. Until then the badge rule is unchanged, and `unsandboxed` stays the
floor wherever restriction is unverifiable.
#### Verification
Fixture replay through the real protocol driver, with no process anywhere near a test.
`codexappserver_test.go` pins the handshake sequencing and the held brief it releases, the
posture riding `thread/start` rather than argv, the write posture NOT inheriting the `exec`
seat's Windows flag, the whole-message repeat read once and two messages not running together,
a message that never streamed still landing, command outcomes on both branches with the
vendor's own failure line, a `completed` status with no exit code staying Unknown, the brief and
the model's thinking both dropped, the turn ending exactly once with no cost, an unseen stop
status not rendered as an answer, a refused handshake being terminal, a refused resume costing
two round trips rather than the turn, the interrupt naming both required ids, every server
request being answered and an unrequested approval refused into the trace, `Decide` never
claiming to have answered anything, usage and limits captured while producing no events, and
garbage on the stream not losing the turn.
`wire_test.go` replays the real capture,
`vendors/testdata/wire/codex-app-server-0.149.1-turn.jsonl` — the sandbox arm itself, sanitized,
so a denial arriving as a success or the vendor's `Access is denied.` going missing fails there.
`go vet ./...` clean, full `go test ./... -count=1` green.
**Spend:** eight billed turns, all cheap prompts in throwaway directories. One of them bought
nothing: the `writableRoots` arm broke on the model's own quoting and is recorded above as owed
rather than as an answer.
**Not verified here: macOS.** Every arm ran on Windows 11. Whether `codex app-server`'s sandbox
behaves differently there is unmeasured, and `PARITY.md` is where that belongs.
### 9.51 the columns were panes, and the operator could not size one (2026-08-31)
The room drew a row of columns and gave the operator no control over how wide any of them
was. Width came from three places, and none of them was a key. `resolveLayoutIn` divided the
usable cells evenly. `FrameOwners` narrowed the frame from the ROUTE, which the operator sets
by addressing a turn and not by pressing anything. `Expanded` — the `f` key — gave the focused
column the whole frame and was a boolean with no middle position. A reader who wanted one seat
wide and the other seats still on screen had no way to ask for it.
This section names what the grid already is, then adds the one control it never had.
#### The vocabulary
**A pane is one drawn seat column, and the room is a single row of panes.** The word is new;
the thing is not. `VisibleColumns` decides which seats hold a pane, `Layout.widthAt` gives each
pane its width, and `columnsBody` paints the row. Nothing about the roster changes here.
The operator now owns two facts about that row:
- **which pane owns the reading width** — the SPLIT.
- **where a boundary between two panes sits** — the SIZE.
Everything else was already built, and this section restates it rather than rebuilds it. Focus
moves between panes with `tab`, `shift+tab`, `h`, `l` and the seat numbers (§9.12, §9.29). `f`
zooms the focused pane to the whole frame (§9.11). A second key for either act would give the
room two spellings of one thing, which is the defect §9.31 records under its own name.
#### The keys, and why they sit behind a prefix
`^w` arms the pane prefix. The next key is a pane key:
| Key | Act |
|---|---|
| `^w` `s` | split: the focused pane owns the reading width, and the other panes hold at `stripColumn` |
| `^w` `>` | grow: move the focused pane's boundary right by one step |
| `^w` `<` | shrink: move the focused pane's boundary left by one step |
| `^w` `e` | even: every pane gets the same width again, and the split clears |
| `^w` `c` | compare (added 2026-09-16): the focused pane joins the split's owner, and the two share the reading width. With no split in force it is a split. On a third seat it replaces the peer. `docs/room-identity.md` carries the rule |
| any other key | cancels the prefix, and is swallowed |
**An unrecognised key is swallowed and NOT re-dispatched.** A prefix that let the second key
fall through would read as tolerant, and it is dangerous in this keymap: `^w` then `q` would
quit the room, and `^w` then `ctrl+c` would cancel a turn. Those are two irreversible acts
reached by a chord the operator has already shown they did not finish. The footer says
`any cancel` for the same reason. A footer that named `esc` alone would imply that the other
keys still mean what they mean, and one of them is `q`.
A prefix, and not four more chords, because the keymap has almost no surface left. §9.49 states
the pressure in its own words: none of the three keys it added was new vocabulary, and each one
had to be taught on a help row that already existed. `[` and `]` there take a third meaning
rather than a fourth binding, and `h`, `j`, `k` and `l` are already focus and scroll. Four
top-level bindings would spend the last of that surface on one feature. One prefix spends one
binding and leaves the rest open.
**The prefix is view mode only.** In compose mode every printable character is draft text, which
is the contract `q`, `f`, `c` and `t` already keep. A prefix armed there would change what the
next letter does while the operator types a brief, and the operator would learn it by losing a
character.
**An armed prefix says so.** The composer box's bottom border reads `PANES` while the room waits
for the second key, at the rank `GATE` and `COMPOSE` already take (§9.44), and the footer names
the pane keys (five since `c` joined on 2026-09-16). §7.8 forbids a mode that changes what an unmodified key means without saying so.
This is such a mode, for exactly one keystroke.
#### The arithmetic, and the two invariants it may not break
The split reuses `weightedWidths`. `framePrimary` already marks which seats own the frame, and
`State.PaneOwner` joins `State.FrameOwners` as a second source for that mark. **The operator
outranks the route.** A split the operator asked for is a request. `FrameOwners` is an inference
from where a turn went. When the two disagree the request wins, and the next dispatch does not
clear it.
The split is pinned to a SEAT, by vendor, and not to whichever pane holds focus. A split that
followed focus would reflow the whole grid on every `tab` press. `tab` moves a marker today, and
§7.1 rule 4 does not budget for a keystroke that re-wraps two columns of prose.
The size is a per-seat bias in cells, `State.PaneGrow`, applied over whatever the base
apportionment is. One press moves one boundary: the focused pane gains a step and the pane to
its right loses the same step. The bias therefore sums to zero, and the row still fills the
terminal exactly. The rightmost pane takes its step from the pane to its left, because it has no
right neighbour to take it from.
Two invariants hold over every frame. `TestColumnsExactlyFillTheWidth` and
`TestPaneWidthsHoldTheirFloors` assert them:
1. **The panes plus their separators fill the terminal exactly.** A short row leaves a ragged
edge. A long row wraps, and the grid shears.
2. **No pane goes below `stripColumn`.** 18 cells is the width at which a column stops being a
seat and renders as a strip, and §9.18 measured that a strip below it cannot say the two
things a strip exists to say.
`normalizeBias` repairs a bias that no longer sums to zero. That occurs when a seat folds out of
the grid between the keystroke and the frame. The repair is deterministic, and it runs inside
`resolveLayoutIn`. A `State` that a test types out by hand therefore cannot produce a torn
frame.
#### The step is the separator's own width
`paneStep` is `1 + 2*gutter`, which is five cells: one rail and its two gutters. The value is
derived and not tuned. One press moves a boundary by exactly the gap the reader sees between two
panes, so the move is visible on the first press. A one-cell step re-wraps nothing on most lines,
and it reads as a key that did not work.
#### The floors, and the ladder below them
`minColumn` keeps its old job and gets no new one. It is the width below which the whole tier
drops to tabs, and `tierFor` still tests it against the EVEN width, before any operator bias.
**The operator therefore cannot size the room out of the columns tier.** A terminal too narrow
for a grid is too narrow before a pane key is pressed, and it stays that way after.
Below the columns tier the pane controls do nothing, and the room refuses to offer them: `^w`
does not arm at the tabs tier, over a turn page, over an arena record, or in a zoomed frame.
That refusal is the reason the footer needs no permanent pane cell at all — the pane keys are
named only while they are live, so there is never a frame that promises a key which does
nothing. It is §9.11's rule reached by a different route: that section drops `f` and `tab`
outright in a one-seat room, and this one never offers the keys in the first place.
The composer border is also silent below the columns tier, even when the split and the bias are
still stored. A legend describing a boundary the reader is not looking at would be the room
describing someone else's frame.
The stored split and the stored bias survive the narrow frame and return when the operator
widens the terminal, exactly as `Expanded` does. A control that forgot on a resize would punish
the operator for dragging a window.
The full ladder, widest first: even panes, then a split pane with strips beside it, then one
zoomed pane with a tab bar, then the tab tier, then `floorMessage`. Every rung was already built
except the second one, and the second one is this section.
#### Colour carries none of it
WORDS on the composer box's bottom border carry the pane state: `^w e panes split`, `^w e panes
sized`, or `^w e panes split, sized`. The key comes first and the state follows, which is the
shape `a not asking` already has (§9.24). The border names the state and the key that reverses
it, so the operator does not have to remember which press undoes a split.
`^w` is ASCII, the words are ASCII, and the feature spends no glyph and no hue. The legend
appears only while the operator has split or sized something, so the ordinary room's frame is
byte for byte what it was.
#### What this does NOT build
- **A second axis.** Panes split left to right, and never top to bottom. A seat's column is one
transcript, and a horizontal boundary through it would cut one document into two viewports.
Nothing measured says a reader wants either half on its own.
- **A pane that holds something other than a seat.** A pane is a seat's column. A pane that held
a file, a diff or a shell is a different product, and the arena review surface (§9.49) already
answers the one case that came up.
- **A key that adds or removes a pane.** `/seat` and `/unseat` do that (§9.31), and they do it in
the roster, where the effect on DISPATCH is visible. A `^w` key that hid a seat while the seat
kept answering would put a live vendor off screen with nothing anywhere to say so, which is
the failure `CollapsedColumns` and the notice row exist to prevent.
- **Mouse drag on a boundary.** §9.10 refused the mouse wheel, and its reasoning transfers
whole. The enum half of that section is MEASURED: there is no wheel-only mouse mode, so a
program cannot ask for a drag without claiming button reporting. The cost half is stated there
as INFERRED rather than measured — that button reporting suppresses the terminal's own text
selection — and it is repeated here at that same strength. Nothing new was measured for this
section, and a boundary drag would buy an input convenience with the room's output, which is
the trade §9.10 already refused on this surface.
#### Verification
`layout_test.go` sweeps every width from `columnsBreak` to 220 and every pane count from 2 to 4,
over the biases the keys can produce, and it asserts the two invariants above. It sweeps a bias
far larger than any keystroke writes, on purpose: `Render` is pure over `State`, so `State` is an
input this package does not control, and an invariant that held only because `paneResize` was
careful is one a hand-typed test could break by accident.
`panes_test.go` drives the keys through `Model.key`. It asserts that the prefix arms, that every
branch clears it, that an unknown key is swallowed rather than quitting the room, that no pane
key reaches the draft in compose mode, that the split does not follow focus, that one press moves
one boundary, that the boundary stops at the floor rather than appearing to move, and that the
arrangement is legible with `PlainStyles` and the ASCII glyph set — which is the test that would
catch a pane feature readable only in colour. Four goldens are new: `panes-split.txt`,
`panes-split-ascii.txt`, `panes-sized.txt` and `panes-keys.txt`. One golden changed, `help.txt`,
by one row, because the panel now names the prefix.
### 9.52 the room ended every agent on the way out and never said so, and the room that came back implied they had lived (2026-09-01)
Two defects, one sentence apart, and they are the same defect seen from each end of a quit.
**On the way out, the room says nothing.** `q` and `ctrl+c` reach `teardown`, which kills every
seat process, and that is correct: an agent that outlives the window that shows it is the
invisible state this product refuses. Nothing on screen says it happened. The operator learns
the contract by noticing, later, that a conversation is cold.
**On the way back in, the room says the wrong thing.** The reattach notice reports
`2/3 seats restored`, and the seat card reports `this seat's thread came back`. Both sentences
are true about the THREAD. Neither says one word about the PROCESS, and the process is the half
that died. A room that opens on four columns of restored threads reads as a room that was left
running. It was not.
This section rules both halves and it introduces one word.
#### The word is `rebuild`, and `reattach` and `rejoin` are not touched
`reattach` is taken. It means *resume the vendor session ids from `room.json`*, in
`Reattachment` (`resume.go`), in the goldens (`testdata/golden/reattached.txt`) and in the demo
script (`STATE.md`). It keeps that meaning exactly.
`rejoin` is reserved and deliberately unspent. A later host lane needs a verb for *a client
reconnects to a live process*, and spending it here on something that is not that would leave
that lane renaming a shipped word.
So the new verb is **rebuild**, and the three words name three different facts:
| word | what it is a fact about | when it is true |
|---|---|---|
| **reattach** | the FILE | the room read `room.json` and holds the saved ids |
| **rebuild** | the PROCESS | the room launched a NEW vendor process on a saved id |
| **rejoin** | reserved | *(a client reaches a process that was already running — nothing does this today)* |
The room performs the first two. It has never performed the third, and until something does, no
surface may use the word.
#### Rung 0 — the quit path states the contract it already keeps
Nothing changes about what quitting does. What changes is that quitting says it.
The room prints a closing line on stdout after the alternate screen is released. It is stdout
rather than a card because there is no longer a frame to draw a card in — the same reasoning
that already puts a failed save on stderr at that point.
The line reports three measured facts and infers none of them:
1. **How many vendor processes were ended.** Counted in `teardown`, from the seat registry it is
already ranging over. A room that spawned nothing reports a measured zero in its own words —
`no vendor process was running` — rather than `0 vendor processes ended`, because the two
sentences answer different questions and only the first one is true here.
2. **What survived on disk, and what did not.** The session ids and the turn number survived, at
the named path. The conversation did not. Saying only the first would let `room.json` be read
as a transcript, which it has never been (`resume.go`'s doc comment).
3. **What reopening costs.** Named as a rebuild, in rung 2's vocabulary, so the two ends of one
quit use one word.
The line is not a warning and carries no mark. Ending the seats is the room working.
#### Rung 2 — the rebuild happens at room open, and it says which of the two things it is
Today the saved session ids are spent on the FIRST DISPATCH: `seatProcess` launches the seat's
process with the saved id when the operator's first brief arrives. Everything that startup costs
is therefore charged to the first brief, and the operator waits for it while looking at a room
that appeared instantly.
Rung 2 moves the launch to room open. The room rebuilds every restorable seat as soon as it
opens, in parallel, and reports each one's progress in its own column.
**What that buys is latency, and it is worth stating exactly which latency.**
`runner/session.go` records the measurement this rests on: a one-word turn cost about 25 seconds
and about $0.23, *nearly all of it startup*. Moving the launch earlier moves the seconds and
does not move the dollars — a process that has started has run no model turn, so nothing is
billed until the first brief. So the claim is split, and the room states both halves:
- **The ~25 seconds are spent at room open instead of on the first brief.** That is what the
operator gets.
- **The ~$0.23 a seat is still billed by the first brief.** Rung 2 does not spend it early and
must not appear to.
Both figures carry a leading `~` and both name what was measured: one one-word turn, once. Four
seats is an extrapolation from one measurement, and the room says so rather than printing a
four-seat total as though somebody had counted it.
**Per-seat progress is measured at every step, and there are four outcomes.** They are carried
on the column's existing note, so this rung adds no field and no render path:
| state | what was measured | what the seat says |
|---|---|---|
| **rebuilding** | the spawn returned with no error | a new process is loading the saved thread |
| **rebuilt** | the vendor announced a session id, and it is the saved one | the thread came back, on a NEW process — and the one you left was ended |
| **forked** | the vendor announced a DIFFERENT session id | §9.43's existing sentence, unchanged |
| **failed** | the spawn failed, or the process exited before it announced anything | the vendor's own line, and the next brief opens a new session |
#### Where each half of the news lives, and why it is not the notice
The first build put the whole statement in the room notice, and the notice **is one line and
it is truncated, not wrapped**. At 120 columns the reattach sentence already fills most of it,
so the settled rebuild rendered as `… 2/4 seats rebuilt in 0s — NE…`. A cost clause that
disappears at a hundred columns is not a stated cost, and the clause being cut was the exact
one the rung exists to say.
The columns are the opposite shape: every note wraps, every seat has one, and `noteCard`
already draws a muted detail block under its title. So the split follows the shape of each
surface, and it lands where `reattachCard`'s own rule already points — the room fact in the
notice once, the seat fact in the seat.
- **The COLUMN carries the sentence that must never be lost**, and the measured cost under it
as detail. The cost is a *per-seat* fact (`$0.23 a seat`), so per-seat is the correct home
and not merely the roomy one.
- **The NOTICE carries the room fact.** `rebuilding 2 seats`, joined to the reattach sentence
while the rebuild runs; then `2/4 seats rebuilt in 24s — NEW processes, not the ones you
left` once it settles.
**The settled sentence replaces the reattach sentence rather than joining it**, and that is
the one thing this rung spends. Joined, it is cut. The reattach sentence is not lost by the
swap: it was the entire notice from room open until the moment the rebuild settled, so its
once-only clauses have been on screen for the whole window — and `telltale council ls` (§7.27)
can print them again at any time.
**`rebuilding` and `rebuilt` are two states and they must not collapse.** A launched process is
not a proven thread. `persistent.go` already refuses to claim otherwise on this exact path —
"Deliberately NOT reattached to the saved thread. Nothing has come back yet" — and the rebuild
keeps that rule. What promotes a seat from `rebuilding` to `rebuilt` is the vendor's own init
line arriving with a session id, which is a statement from the vendor and not a timer.
**The fork case is not new machinery.** A vendor that is asked to resume and answers in a fresh
conversation is §9.43's finding, and `adoptSession` already detects it through `forkWatch` and
already prints the honest correction. The rebuild arms `forkWatch` with the saved id exactly as
a dispatch does, so a fork at room open and a fork at turn 5 report identically. One mechanism,
one sentence.
**The seat card and the note say different halves, on purpose.** The existing reattach card says
`this seat's thread came back`, which is a claim about the thread and stays true. The rebuild
note under it says the process is new. Together they state what came back and what did not.
Neither alone would.
#### What rung 2 deliberately does not do
**It persists nothing new.** No transcript, no scrollback, no vendor output. `resume.go` ruled
this for the same data, and the ruling does not change because a process moved: duplicating any
of it would be a second copy of a private conversation in a place the user did not choose.
`room.json` stays session ids, a workspace, and numbers.
**It starts no process the room would not have started anyway.** Every seat it launches is a
seat the first brief was going to launch. The rebuild changes WHEN, never WHETHER. A room that
opened and spawned a seat the operator had not seated would be spending on a roster nobody
typed.
**It launches nothing that is not restorable.** A seat with no saved id, a vendor this machine
cannot run, and a seat that is not driven as a live process are all skipped. Each is skipped for
a measured reason, and none of the three is reported as a failure.
**It does not survive anything.** This is the sentence the whole rung is built around: the
agents did not live through the quit. They were ended, and new ones were started on the ids they
left behind. A room that rendered a rebuild as a continuation would be the most expensive lie
this surface could tell, because the operator would trust a history that no process holds.
#### Verification
The rebuild is exercised with the spawn vars stubbed (`countSpawns`), which is the package's
standing rule: a council test never starts a vendor. The kickoff is fired from `Init`, which no
test calls, so a model a test builds directly launches nothing at all.
Goldens pin the two rendered states apart — `rebuilding.txt` and `rebuilt.txt` — because
`rebuilt` and `survived` rendering alike is the regression this section exists to prevent. Both
take their strings from the model itself rather than from text typed into the test, so a
wording change moves the golden instead of quietly passing a stale assertion. **No footer hint
was added**, so no existing golden moved: every golden's last line is the footer's key hints,
and one new hint there rewrites about eighty-nine files at once.
### 9.53 a live seat shows Claude Code's own screen, and measures nothing from it (2026-09-01)
Every seat in this room is the same kind of thing: a column whose body is text council parsed
out of a structured stream. That is what makes the gauges honest, and it is also the whole
limit. The operator cannot see what the agent's own terminal shows. The permission prompt it
draws, the spinner it spins, the box it puts around a diff — none of that is in
`--output-format stream-json`, so none of it is in the room.
This section adds a second KIND of seat. A LIVE seat runs the vendor in its own interactive
mode, under a pseudoconsole, and draws that program's real screen inside the pane the seat
already owns. Everything else about the seat does not move.
#### The display-only contract
**The live screen contributes NO measured field.** Not a cost, not a token count, not a context
percentage, not a quota window, not a posture, not a phase, not a turn clock. Every gauge, every
badge and every number on a live seat still comes from the structured adapter path, or renders
absent under [§4a.1](#s4a-1)'s rule.
This is not caution. It is the honest-gauge rule ([ADR-001](#adr-001)) applied to the one input
that would break it. A screen of ANSI is a picture of a program, and reading a number off a
picture is inference. `claude.exe`'s own status row prints a dollar cost and a weekly-limit
sentence — both were captured in the spike, both are exactly the figures this room renders
elsewhere, and both are exactly the figures the room must NOT take from here. `CapNone` exists
to refuse inference. A field scraped off a repaint is a field this product spent its whole life
declining to invent.
So the seam is drawn at the type. The emulator lives on `Model`. The only thing that crosses
onto `State` is `State.Live.Grid`, a slice of already-decoded plain rows, and `Render` draws
those rows and nothing else. `applyPTY` writes `State.Live` and no other field of `State`.
`TestLiveSeatMeasuresNothing` asserts that as a property rather than as a review note: it feeds
the emulator a script carrying a cost, a quota sentence, a posture word and an elapsed time,
then compares the whole `State` before and after and demands the only difference is `Live`.
**The live pane is a SECOND process, and it doubles that seat's spend.** The structured session
keeps running; the live child is another `claude` on the same account. The room says so where
the pane is, in a word, because a surface that quietly billed twice would be the most expensive
thing this section could ship.
#### Only Claude Code can take the seat
`vendors.Persistent` has exactly one implementer (`vendors/claude.go`). Codex, Antigravity and
Grok are batch programs that exit every turn, so a pseudoconsole on one of them is a pane that
is empty between turns. Cursor speaks ACP JSON-RPC and draws no terminal UI at all; seating it
live would mean a different invocation that loses every measured field the ACP path supplies,
which is the display path bought with the structured one — the exact trade this section refuses.
The live invocation is not council's invocation. Council runs `claude --input-format stream-json
--output-format stream-json ...`. The live child runs `claude` with no arguments, which is a
different program mode with a different contract, and that is why it is a separate child rather
than a second reader of the first.
#### `os/exec` cannot spawn a ConPTY child
Go's `syscall.SysProcAttr` on Windows has no field for a `PROC_THREAD_ATTRIBUTE_LIST`, so
`EXTENDED_STARTUPINFO_PRESENT` cannot be honoured through `exec.Command`, and
`PROC_THREAD_ATTRIBUTE_PSEUDOCONSOLE` is the whole mechanism by which a child attaches to a
pseudoconsole. The spawn half is therefore a direct `windows.CreateProcess`, and it is the only
place in this repository that starts a process without `os/exec`.
The containment half is NOT rewritten. `windowsGroup`'s job object was measured working
unchanged on a ConPTY child: `AssignProcessToJobObject` succeeded, and closing the job handle
killed the child immediately where `ClosePseudoConsole` alone had left a `claude.exe` REPL alive
more than three seconds. `attach` is refactored to take a pid, `attach(*exec.Cmd)` calls it, and
the PTY calls the same function. One job-object implementation, two spawn shapes.
`golang.org/x/sys/windows` v0.47.0 already exposes `CreatePseudoConsole`, `ResizePseudoConsole`,
`ClosePseudoConsole`, `PROC_THREAD_ATTRIBUTE_PSEUDOCONSOLE` and `NewProcThreadAttributeList`, so
the pseudoconsole itself costs no new module.
One non-obvious requirement, measured: the child does NOT attach to the pseudoconsole unless
`STARTF_USESTDHANDLES` is set with all three std handles left at zero. Without it the child
keeps the PARENT's std handles, even with `bInheritHandles` false and even when the parent owns
no console at all — a `cmd /c echo` printed on the parent's own stdout and the pty stream stayed
empty. That one flag is the difference.
#### `CREATE_NO_WINDOW` is the trap, and it fails silently
`proc_windows.go` sets `CREATE_NEW_PROCESS_GROUP | CREATE_NO_WINDOW` and `HideWindow` on every
child today, and the comment there is correct for every child it was written for. It is fatal
here. Measured on 2026-08-31, a ConPTY child created with `CREATE_NO_WINDOW`:
- emits ZERO bytes — not even conhost's own preamble,
- accepts no input, so a write to its stdin does nothing,
- exits 0, and `CreateProcess` returns a NIL error.
There is no log line, no error and no output. The pane is simply blank. `DETACHED_PROCESS`
fails the same way and additionally exits 1. `CREATE_NEW_PROCESS_GROUP` and `HideWindow` are
both compatible and both measured working.
An implementer reaches this bug by copying the existing spawn helper, which is the most likely
thing to do, so the refusal is written down at the call site and the flag is named in
`pty_windows.go`'s doc comment rather than left to this file.
**There is no console flash to suppress.** That was measured with a differential `EnumWindows`
scan carrying a positive control, across every flag combination and on runs up to twenty
seconds: zero new visible windows. The reason is structural — a child holding
`PROC_THREAD_ATTRIBUTE_PSEUDOCONSOLE` attaches to the pseudoconsole's own headless conhost and
never asks Windows for a console, so `CREATE_NO_WINDOW` has nothing to suppress and its only
effect here is to break the attach. One caveat for whoever writes the regression test: ConPTY's
conhost DOES create a window of class `PseudoConsoleWindow`, `IsWindowVisible` reports it as
visible, and it never paints. A flash test must match `ConsoleWindowClass` and must not treat
`PseudoConsoleWindow` as a flash.
#### `fit` is not sufficient, and no width helper can be
`fit` (§9.5) is ANSI-aware for SGR, which is the `padRight` trap it was written for. It does not
solve this problem and cannot. `lipgloss.Width` measures display cells, and a cursor-move, an
erase-in-line, an absolute cursor position and a scroll-region set all measure as ZERO cells —
so `MaxWidth` passes them through untouched into telltale's frame, where they execute against
the host screen.
These were all found in the real captures, and each is a different kind of damage:
| Sequence | Damage if forwarded |
|---|---|
| `ESC[8;20;60t` | resizes the operator's REAL terminal window |
| `ESC[>0q` | the real terminal answers into telltale's own stdin — a reply to a question the host never asked |
| `ESC[?9001h` | switches the host into win32-input-mode, and Bubble Tea then decodes keys wrong |
| `ESC[?1004h` | enables focus reporting on the host, injecting focus events into its input |
| `ESC[?2004h`, `ESC[?2031h` | change how the HOST receives pastes and theme changes |
| `ESC[2J`, `ESC[13;1H` | absolute addressing against the pane's origin, painting over telltale's frame |
| `ESC]0;...BEL` | sets the host window title to the guest's path |
The escapes must be CONSUMED, not measured. That is an emulator, and it is mandatory.
#### Why `x/vt`, and not a hand-rolled parser
`github.com/charmbracelet/x/vt` at `v0.0.0-20260830003929-9f48cc723c1c` (MIT) is taken as a
direct dependency. Three reasons, in the order that decided it.
**A hand-rolled parser is not a smaller problem, it is the same problem with a smaller test
suite.** The table above is not the list of sequences to handle; it is the list found in twenty
seconds of two guests. Correctness here is not "parse the ones we saw" — a sequence this room
does not model and forwards by accident is a host-terminal corruption the goldens cannot see,
because `PlainStyles` neutralises telltale's own escapes and is blind to a vendor's. The failure
mode of an incomplete parser is invisible to every test this repo has.
**The dependency cost is exactly two modules.** Measured by diffing `go list -m all`: `x/vt` and
`github.com/charmbracelet/x/exp/ordered` (v0.1.0, MIT, a small generics helper). Everything else
it needs — `ultraviolet`, `x/ansi`, `x/term`, `colorprofile`, `displaywidth`, `uax29`,
`go-colorful`, `go-runewidth`, `cancelreader`, `uniseg`, `terminfo`, `x/sys`, `x/sync` — is
already in the graph because Bubble Tea v2 pulls it. `golang.org/x/exp` is NOT pulled in.
**It links against this repo's exact pins, proven rather than assumed.** A compatibility program
built against `bubbletea/v2 v2.0.9`, `lipgloss/v2 v2.0.6` and
`ultraviolet v0.0.0-20260811164956-006e29f97886` compiled and ran, and one type satisfied both
`tea.Model` and `uv.Drawable`. `x/vt`'s own `go.mod` asks for an OLDER `ultraviolet`
pseudo-version; minimal version selection promotes it to this repo's newer one and it still
works. That promotion is the standing version risk and it is measured green today.
The alternatives were surveyed and rejected on evidence: `hinshun/vt10x` has no tagged release
and no commit since 2022-03-01, and `liamg/darktile` is a terminal application whose emulator is
not importable.
The counted cost of taking it: `x/vt` has no tagged release either. It is a pseudo-version off
`main` from a vendor moving fast, so its API can change with no semver signal. That is a real
cost and it is accepted, because the alternative is owning a VT parser in a repository whose
test suite cannot see the bug class it would introduce.
#### The update path, which is dictated by `Render` purity
`Render` is pure over `State` — `TestRenderIsPure` renders twice and demands byte equality, and
`TestElapsedIsPureOverState` sleeps between the two renders and demands it again. A pane that
asked an emulator for "the current screen" from inside `Render` fails both. So the screen is
materialised in `Update` and `Render` draws finished rows:
```
ConPTY output pipe
-> reader goroutine, chunks into a BOUNDED channel (a full channel stalls the child)
-> waitPTY() tea.Cmd, shaped like waitEvents (block on one, drain up to a cap)
-> Update, case ptyBatchMsg
emulator.Write(chunk) -- the emulator is on Model, never on State
st.Live.Grid = snapshot() -- decoded rows cross onto State
re-arm waitPTY()
-> Render draws st.Live.Grid
```
The emulator sits on `Model` for the same stated reason `gateInputs` does: it holds the whole
raw stream, and `State` is what the renderer can reach. Only finished rows cross.
Coalescing falls out of the shape. One `Update` applies every chunk it drained but takes ONE
snapshot, so a flood costs one grid copy per frame rather than one per chunk. The emulator holds
the state, so nothing is lost by coalescing — which is the same argument `drainMax` already
makes for text.
**Resize is issued from `Update`.** `layoutFor` is pure and callable there, so the room computes
the pane's cell rectangle, and when it differs from the pseudoconsole's current size it calls
`ResizePseudoConsole` and `Emulator.Resize` together. Doing one without the other leaves the
grid and the pty disagreeing about the width. Never from `Render`.
#### What the pane draws, and what it drops
The pane draws the emulator's TEXT grid. It does NOT carry the guest's colours, and that is a
decision rather than an omission.
Three reasons. A golden may not embed ANSI (§9.5), and `PlainStyles` neutralises telltale's own
escapes and not a vendor's, so a coloured grid could not be pinned by a golden at all. Colour in
this room is always a SECOND signal for a distinction telltale is making, and the guest's
colours are not telltale's claims. And carrying SGR through a clip re-opens the bleed hazard
`fit` exists to close: a row clipped mid-run leaves a colour open into the next pane.
The cost is real and worth naming: a guest that distinguishes something by colour alone loses
that distinction here. Claude Code does not — its own output carries words and glyphs — but that
is a measurement of one guest, not a rule about guests.
The guest's own GLYPHS are kept verbatim, including under `--ascii`. `--ascii` governs
telltale's glyph set, and the live pane's chrome respects it. The grid is another program's
screen, which is quoted material in the sense CLAUDE.md already carves out for error text: a
room that rewrote what a vendor printed would be showing something the vendor did not print.
The pane does not scroll. A live terminal viewport is the emulator's current screen, and the
emulator is sized to the pane, so there is no hidden region for a scroll key to reach. The
guest's own scrollback stays inside the guest.
#### Honest degradation, and the build floor
**Only Windows build 10.0.26200 was measured.** ConPTY's documented floor is Windows 10 1809
(build 17763), and that floor is DOCUMENTATION, not a measurement made here. Older builds are
also reported to repaint more heavily, so the 97.3%-payload result may not hold below Windows
11. **This is the single largest unverified claim behind this section.**
The seat is therefore gated rather than trusted. `StartPTY` reads the running build with
`RtlGetVersion` and refuses below 17763 with a sentence naming the build it found and the build
it needs. A refusal renders as an `unavailable` card in the pane, which is the shape §4a.1
already rules for a field that could not be read: the room says what it could not do, and does
not draw a blank pane that looks like an agent with nothing to say.
**Non-Windows compiles and refuses.** `pty_other.go` carries the `!windows` build tag and
returns `unavailable on this OS`. A build tag rather than a runtime check, so a Unix build does
not carry a pseudoconsole path it can never take, and so the refusal is a fact about the binary
rather than a guess made at the moment somebody presses a key.
The other measured gaps, recorded so a later reader does not mistake silence for coverage: no
streaming agent turn was driven through a pane, no alternate-screen guest was exercised (`x/vt`
has two screen buffers and should handle the alternate-screen mode set, and that path is
untested here), the longest spike run was twenty seconds so nothing is known about handle or
memory growth over a long session, and the documented ConPTY close deadlock did not reproduce on
this build — a negative result on one machine, not a guarantee.
#### The spawn guard grows a sixth var, in the same change
`startPTYSession` is declared beside the other spawn vars in `persistent.go` and takes a
`runner.Spec`, so `refuseRealVendor` needs no fork — a second refuser would be a second thing to
keep honest. `TestMain` wraps it, and `countSpawns` stubs it with the restore in the existing
`t.Cleanup`.
This is not optional and it is not a follow-up. A council test never spawns a vendor, and a PTY
child is a vendor process like any other: the fact that its output is DISPLAY ONLY changes
nothing about whose account it runs on. A spawn that escaped the count would let
"nothing was spawned" pass over a vendor launched with a pseudoconsole attached, which is the
exact assertion those tests exist to make.
#### The seat is seated by a flag, and by nothing else
`--live claude` names the seat. Parsing lives in `ParseLive` in `internal/council` and `Run`
calls it before the alternate screen, so `cmd/telltale` hands over one field and learns nothing
about vendors — and an impossible seat is a line on stderr rather than a card behind a TUI the
user has to quit to read, which is the discipline `--brief` and `--trace` already follow. There is no key. A key
would have to be taught on the help panel and would be a second way to spend a second account
that the operator has not asked for; the flag is a decision made once, at the point where the
room is opened, and that is the right place for a control that doubles a bill.
**Nothing types into the pane either.** `PTYSession.Write` exists because the pseudoconsole's
input pipe exists whether or not anybody writes to it, and no council code calls it. So the
first cut is a WINDOW, not a terminal: the operator watches the agent's own screen and answers
it, when it asks something, in the seat's structured column beside it. Sending keystrokes needs
a key encoding, a focus rule and an answer to what `q` means while a pane has the keyboard, and
none of those is a display question. Owed, and named here rather than half-built.
#### Verification
`go vet ./...` clean, `go build`, and `go test ./internal/council -timeout 20m`. The live seat's
tests drive the emulator and the render path directly, never `startPTYSession`, which is the
same shape `arenacheck_test.go` uses for the check var.
Two goldens are new — `live-seat.txt` and `live-seat-ascii.txt` — and no existing golden moved.
The feature is entirely opt-in: `State.Live`'s zero value is off, no footer hint was added, and
no chrome row was spent, so every existing frame renders byte-identically. The goldens pin the
seat's chrome and a FIXED emulator grid produced by feeding a fixed byte script through `x/vt`
in the test. They pin what a golden can pin: that escapes were consumed, that the grid is
exactly the pane's rectangle, and that the display-only marker is a word rather than a colour.
**Not verified here: a live drive.** No `claude` interactive session was run through a pane by
the session that wrote this. The spawn guard makes that impossible from inside the suite by
design, and it is the same class of debt as the host's first live turn ([STATE.md](../STATE.md)):
an operator-driven check, owed and named rather than implied.
### 9.54 the room was a committee, and a crew's seats are busy one at a time (2026-09-02)
`dispatch()` opened with one line — `if m.turn != nil { "a turn is already in flight — ctrl+c
cancels it" }` — and that line was the whole difference between the room that existed and the
room the owner asked for. Seats were concurrent inside a turn and the room was serial across
turns: `@all` fanned one brief out to five processes at once, and the moment any of them was
still answering, a second brief to a different seat was refused. You could not hand codex a
refactor and, while it ran, hand grok the docs. The owner's ruling is that council is a **crew**,
not a committee answering one question at a time, and a crew's seats are busy or idle **one at
a time**.
#### What a turn is now
A turn is a fact about a SEAT. `Model.turn` is gone; `Model.turns` maps each seat in flight to the
dispatch it is answering, and `turnState` is the record of ONE DISPATCH — the seats one press of
enter sent a brief to, and what those seats share while they answer it: the turn number, the
route the header names, the arena's all-or-nothing bookkeeping. What moved down to the seat is
everything the operator can now do to one seat while its neighbours work: its process handle,
its own child context (so cancelling it kills its child and nobody else's), the cancellation
word, the give-up.
**The turn number stays one room-wide sequence**, and that is a ruling rather than a leftover. A
per-seat count was considered — "codex's turn 4, grok's turn 4" — and refused, because every
surface that already prints a turn number reads one coordinate: the separators, the by-turn page
(`PageTurns`), `/retry`, `room.json`, the reattach card. Two seats both on "their turn 4" would
leave the page able to open only one of them. So a turn number is a **dispatch** number: turn 5
is the fifth brief the room sent, whoever it went to, and a seat's own history is the subset of
those numbers it took part in. `Column.TurnN` already carried exactly this.
#### What the operator sees
- **A brief to a busy seat is refused for THAT seat, by name and turn.** `@codex …` while codex
is mid-answer: `a turn is in flight on codex (turn 4) — ctrl+c on its column cancels that
turn, or address another seat`, and the draft stays put. `@all` with codex busy goes to the idle
seats and says so: `sent to grok, agy — skipped: codex (turn 4), still on a turn; ctrl+c on
its column cancels it`. Both halves are measured — the seats in the new dispatch's live set, the seats
whose own turn refused them. The refusal is what stands between a persistent seat and a second
prompt written into a process mid-turn, which is the failure the room-wide wall was in front of.
- **The header names the newest dispatch and counts the rest.** `turn 5 → codex · 3 in flight`.
The count is `State.SeatsInFlight()`, measured over the columns, and it is printed only when
some seat in flight is on a turn OTHER than the one the cell names (`inFlightBeyond`, read off
`Column.TurnN`) — an `@all` turn with its three seats streaming reads `turn 3 → everyone` as it
always did, because `· 3 in flight` beside it would be the route restated as a number. A live
column with no turn number at all is not counted as another dispatch: zero means never
dispatched, and a count that read it as elsewhere would be inferring. The route retires when
ITS dispatch lands (`ts.n == st.Turn`), not when the room goes quiet.
- **ctrl+c has three meanings and the footer says which is live.** The focused seat's turn if it
has one (`ctrl+c cancel codex`); everything in flight when the focused seat is idle (`ctrl+c
cancel all`); quit when nothing is. With one seat in flight the label is the plain `cancel`
every earlier frame carried. The gate line's `cancel the turn` became the same label, because
the key reaches `viewKey` through `gateKey`'s fall-through and means there what it means here.
`x` is unchanged: the per-seat give-up with its card.
- **The room-wide refusals name the seats.** `q`, `/cd`, `/seat`, `/unseat`, `/read`, `/write`,
`/retry`, `/adopt` and `/arena drop` still need the whole room idle — each changes something a
busy seat was dispatched against — and each now says `a turn is in flight on codex (turn 4),
grok (turn 5) — …`. `c` and `u` became per seat: the thread or the tree they touch is the
focused seat's alone.
- **A race is the one turn that still owns the room.** `/arena` refuses while any seat is busy
(a race that skipped a seat is not a comparison), and every brief is refused while a race runs
(its racers are writing into worktrees a room brief would cut across). `race()` is the read.
- **A `/flow` hop waits on its own seat.** Hop N+1 dispatches the moment hop N's seat lands,
whatever the rest of the room is doing; the chain's death-on-teardown runs only when the hop's
own dispatch ends (`turnState.flow`). A hop whose seat is busy with an unrelated brief stops
the chain by name — `flow stopped at hop 2/3: @codex is still on turn 4 — …` — rather than
queueing behind the seat, because a chain that dispatched itself later, when a seat happened
to free up, is the room acting on its own at a moment nobody chose (§9.16's argument).
- **A rebuttal quotes what a busy neighbour last FINISHED saying.** The snapshot is taken per
seat at its dispatch; a seat mid-answer contributes its last filed turn (`settledReply`) and
nothing if it has never filed one. Quoting the half it has streamed would put half an argument
in front of another model as though it were whole.
#### The inbox
§9.40's strip named only the seats stopped on a gate. A crew has a second stall of the same shape:
an answer lands in a column the reader is not on, the room knows, and the reader has to go
looking. The strip now lists **seats whose turn ended since the reader last had the keys on
them** — `⚠ NEEDS YOU 2 Codex 3 Grok done 4 Cursor failed` — with the terminal phase word,
the same word the column header speaks, as the whole distinction between the two kinds of entry
(no glyph, no colour: it reads the same under `--ascii` and `NO_COLOR`).
*Amended 2026-09-03 (`LEDGER.md`):* the lead is `NEEDS YOU` only while a listed seat is
blocked on a gate. A strip of landed replies and nothing pending opens `UNREAD`, with no mark
(`unreadLead` in `internal/council/needsyou.go`); the entries and their phase words are unchanged.
Every entry is a measurement, on §9.40's own rule. A landing is two stamps the Model took itself:
`Column.Ended`, written in `finishColumn` while the seat still holds a turn (so the second
retirement a persistent seat or an ACP racer goes through cannot re-stamp it), and
`Column.LastFocus`, written by `setFocus` on BOTH the seat left and the seat entered, so it marks
the end of the reader's last look. The strip is `Ended.After(LastFocus)` over terminal columns.
Nothing stores what was acknowledged — the comparison cannot drift, which is the argument §9.40
made for deriving the gate half from `Focus`. The default-focus hole stays open for §9.40's reason
and closes itself: the focused seat is never listed, and the reader's first departure stamps it.
`.` is the strip's key: the next listed seat after the focus, wrapping. The footer names it only
while the strip has an entry (§7.8: never a key that does nothing); in compose it is a full stop.
#### Mechanics worth knowing
- **One event reader.** Every dispatch and every batch re-arms the pump, and two goroutines
reading one channel would deliver batches to `Update` out of order — an exit before the text it
followed. `Model.eventsArmed` makes `waitEvents` hand out one reader; `Update` clears it on the
batch. `sendTurn`'s Cmd can therefore be nil for a dispatch that DID start, so `applyArenaSetup`
reads `race()` rather than the Cmd.
- **The give-up and the cancel outlive the seat's turn.** `givenUp` and `cancelling` moved from
the turn to the Model, keyed by seat, and are cleared at the seat's next dispatch. A cut seat's
turn ends the instant its column lands, while its process is still draining — and on the
persistent seat still answering the interrupt with a failed `result` — and those echoes met a
seat with no turn, where `applyEvents` would have written the abort error over the give-up's own
note.
- **Teardown walks `dispatches()`** — every distinct record in the map — and reaps each one's
one-shot handles, racer handles, ephemeral sessions and context. `TestTeardownReapsEveryDispatch`
pins two dispatches; `teardown_test.go` still pins the racer.
- **Geometry.** `frameOwnersFor` counts a column still in flight as an owner, read off the
column's own phase, so a brief to grok while codex streams does not narrow codex's prose under
the reader. The no-mid-stream-reflow rule is about the room moving because a VENDOR did
something; a dispatch is the operator's act.
#### What this section does NOT change
`room.json` gains no field: sessions and the turn count are room-level and the save runs at each
dispatch's end. `--trace` clocks are per process and unaffected. Session resume per vendor is
unchanged — a seat is refused a second prompt while busy, which is the only new rule a resume
could meet. The spawn guard is untouched: every crew test dispatches through `countSpawns`.
#### Verification
`gofmt`, `go vet ./...`, `go test ./... -count=1`, `go build ./...`, and the windows/amd64 and
darwin/arm64 cross-builds, all clean; `go test -race ./internal/council` once. Five goldens are
new — `two-in-flight`, `busy-seat-refused`, `inbox-landed`, `inbox-landed-ascii` — and the
existing goldens that moved moved on ONE cell each: the footer's cancel label, on frames where
two or more seats are in flight, which is the key's new meaning drawn honestly. No header of an
existing frame moved.
**Not verified here: a live crew.** No two vendors were run concurrently by the session that
wrote this. What the suite pins is the room's bookkeeping over stubbed processes; what only a live
run can show is a persistent seat taking its NEXT brief cleanly after a per-seat cancel, two
one-shot seats' event streams interleaving through one reader without a stall, and a `/flow` hop
landing beside an unrelated seat mid-answer. Owed and named rather than implied, on §9.53's rule.
### 9.55 a crew's writers share one tree, and the room becomes the integrator (2026-09-02)
§9.54 made the seats concurrent and left them where they were: every ordinary turn ran in
`State.Workspace`, and a worktree existed only inside `/arena`. The moment two writing seats
could answer at once, two writers were in one checkout — which §9.37 had already ruled out
for a race in one sentence, "four writers in one shared tree are not four answers, they are
one trampled tree". And the containment could not be a card: only the stream-json seat can
be asked before a write (`canGate`), the other four act unasked. So the crew's containment
is the race's, made structural and made permanent: **a writing seat gets its own worktree,
cut once from the room's HEAD and reused for every writing turn after it, and the room is
the integrator** — nothing a seat writes reaches the room's tree except by `/adopt`.
#### What changed
- **One worktree per writing seat, by default.** In write posture the first brief to a seat
cuts `-seat-` beside the workspace on `seat/`, from HEAD, and the
seat's process runs there (`seatDir` is the one read every spawn path — `specFor`,
`seatProcess`, the flow receipt, `/hand` — goes through). Read posture and a `/flow` read
hop keep the shared tree: a tree cut for a hop that cannot write is containment for a
hazard that does not exist. `--shared-tree` is the opt-out, a FLAG rather than a room word
because it decides where five processes write, the same class of fact as the workspace.
The cut runs off the render loop under arenasetup.go's deadline, serially, with the step
named on the footer (`worktree: preparing worktree for codex…`) and ctrl+c stopping the
setup; a tree already cut is kept. A tree an earlier room left is found by name and
reused, and the notice says so; a directory at that name that is not the seat's worktree
is refused by name. Nothing is seeded into a seat's tree and no brief file is written into
it — those are affordances for a fresh throwaway tree, and a seat's tree is reused.
- **The badge says which holds** (`Column.Containment`, stamped at dispatch from the same
read the spawn uses): `wt: seat/codex`, `shared tree`, or `⚠ shared tree · ` for a
fallback the room could not avoid, with git's own sentence in the notice when it happened
and the whole reason on the `?` postures page. At a three-seat column's width the reason
sheds before the word and the mark stays, because a clipped reason is not a reason
(§9.11); the granularity word sheds first, on stripBadges' own order. A seat never
dispatched wears no badge: no process, no directory, no claim. `--ascii` keeps the mark
as `!` and the separator as `-`.
- **`/adopt ` merges a seat's branch, the race's way** (`adoptSource`). The card
names `git merge --no-ff seat/codex` onto `adopt/seat-codex`, refuses a dirty room and an
empty branch exactly as before, and after the merge **resets the seat's tree and branch
onto the new HEAD** — a reset that fails degrades the notice, never the adoption, and says
what moves the tree by hand. Resolution order is a ruling: the arena branch when the
seat's CURRENT turn is a race attempt (that is the block on the column), else the seat
branch, else the race receipt. Hybrids work across two seat branches on
`adopt/seat-+`; a hybrid across a race attempt and a seat branch is refused,
because one receipt names one kind of branch. The donor is not reset — only its named
paths were taken.
- **`/arena record` counts seat adopts apart** (`ArenaRecord.SeatAdopts`, from
`adopt/seat-[-k]` refs): their own sentence under the window, inside no rate. A
race adopt is a verdict among competing attempts at one brief; a seat adopt is the room
taking one seat's ordinary work with nobody else in the running, and folding it into a
rate would score a seat for races it never entered.
- **`/flow` fans with `&`** (`FlowStep.Stage`): hops joined by `& @seat` are one stage and
one dispatch, each seat handed its own task; the stage after `->` dispatches when the
stage's last seat lands (`StageDone`), carrying every predecessor's reply as its own
labelled fence, in landing order. The header names the stage (`hop 1/2 @codex & @grok`).
Two refusals a fan adds, at parse: a seat named twice in one stage, and a stage mixing
write and read hops — a stage runs at ONE posture, because a fan is one dispatch and
§9.16's table gives a dispatch one posture. One `y` releases a writing stage and the card
names every target. A busy seat stops the whole stage by name, on §9.54's rule. An
ampersand not followed by a mention stays prose.
- **`/hand `** (`handcmd.go`) puts ``'s stat and patch — against the point
its branch parted from the room's HEAD, read with `add -N` so created files count — into
the draft addressed `@`, fenced as measured git output with the tree and branch named.
A draft, on §9.49's whole argument; `enter` is the only thing that spends. The cap is the
composer's: a patch past it is cut at a hunk boundary and the closing fence states the
cut and the way to the whole (`y`).
#### What this section does NOT change
`room.json` gains no field: a seat's tree is rediscovered on disk by name. The arena's own
worktrees, seeding, brief file, ranks, `x`, `u` and `/arena drop` are untouched — a racer's
badge now reads `wt: arena/t/`, which is the tree it already stood in. The spawn
guard is untouched: every test here dispatches through `countSpawns`, and the git that
runs is against a temp repository.
### 9.56 the gauges route, and a real run can be shown without a vendor (2026-09-02)
Two changes, one premise: §9.54 made the seats a crew, and a crew changes what two existing
things are FOR. The quota readings (§9.21) used to say "this turn may not land"; on a crew they
answer "which seat has room for this brief". And the room's event stream, which every column
already applies, is the one thing that could show the room to someone who has not paid for five
seats — if it were kept, and if what was kept could be shown honestly.
#### The readings as routing
The routing cell qualifies the route. `→ codex · 5h 94% used` when the draft addresses a seat
whose relayed window is at or above `--headroom-warn` (default 90) and under a hundred;
`→ everyone · codex 5h 94% used` on a route that names a set, so the seat is named. `@auto` is
a route word (`Route.Auto`) that resolves against State: among seated, idle seats with a measured
reading, the most headroom in the SHORTEST window — the window that resets soonest, since that is
the one the next brief spends from — ties to seating order. The cell states the pick before
enter, `→ auto: grok (5h 12% used)`; enter rewrites the draft to `@grok …`, dispatches, and the
notice repeats the choice. The header and room.json record a turn to grok, which is what
happened, and no surface learns a fifth route shape.
Every rule is §4a.1's, restated for a rank:
- **Nothing is invented and nothing is aggregated.** The hint copies one window's label and
percentage. `@auto` ranks headroom (a hundred minus the vendor's own figure) and prints the
figure it ranked. No total, no average, no dollar.
- **An absent seat is never ranked.** Cursor and grok have no reading anywhere (§7.17), an
unrelayed Claude has none today, and a rank needs a number. With no measured seat in the room
the cell says `→ auto: no measured reading` and enter refuses — `@auto needs a measured
reading; none of the seated seats has one` — rather than falling back to the default route.
A brief handed to the readings must not go quietly to Claude because the readings were empty.
- **A stale reading is absent, not a number.** A window whose reset has passed is dropped by
quotacache on read and by `measuredWindows` between reads, when the room's own clock passes
it. A reading past `quotaAgeWarn` is the alarm's to name as stale (§9.21) and is absent here.
- **The hint stops one short of a hundred.** A full window is the alarm's cell, with the
warning mark, and one fact in two cells on one line is the drift vendorTag's comment warns of.
- **A busy seat is never picked.** §9.54's refusal would meet the pick a keystroke later.
The threshold is council's own pick, which is why it is a flag: the default is a boundary no
vendor published, so the operator can move it, and the cell prints the vendor's figure beside
whatever threshold applied. `quotaFullPercent` is unchanged and still the only threshold this
package did not choose.
#### The recording, and the boundary it is written under
`--record ` writes the room's event stream as it happens: the room line (seats, postures,
posture, workspace as the header drew it), each dispatch as its seats hold it (turn, route, each
seat's brief as `startTurn` echoed it, whether it ran on a persistent process), every
`runner.Event` in every batch `applyEvents` receives, and each gate decision the operator made.
JSON lines, millisecond offsets from the monotonic clock, one flat record type so a fixture can
be typed by hand. `runner.Gate.Input` is the one field dropped: never rendered, a Write's whole
file content, and a replay has no vendor to hand it back to.
**This is content, and it is not argued as one of the numbers-and-keys writes.** CLAUDE.md's
read/write boundary holds room.json, the quota relay and the token relay to numbers and keys.
The recording is the event sink's kind of exception — "different in kind", contained by scope
rather than redaction — and it is contained three ways: an explicit path the operator typed
(a path under `~/.telltale` is refused before a byte is written, and an existing file is
refused rather than overwritten or extended); an explicit flag typed at the door (no key in the
room starts one); and off by default (the recorder is nil, every hook on it is a no-op, no test
builds one). `recording.go` carries the argument at length; this paragraph is the decision.
**No redaction, on purpose.** The recorder writes what `applyEvents` saw. A recording that
differed from the run would be a second truth; the replay puts the same bytes through the same
redactor at the same choke point, so the frames match and the file underneath is the raw
stream. The review the README requires of every frame is given a tool instead:
`telltale council replay-check ` lists the workspace, the seats, every session id, every
tool line and gate card, and the size of the prose — and says it did not read the prose.
#### The replay
`--replay ` opens a room over the file. `Run` returns to `runReplay` before `LoadRoom`,
the brief, the trace, the relay or a host is consulted. The State is built from the room line
— seats, labels, postures, granularities, the abbreviated workspace with `Home` left empty — so
a replay on a laptop with nothing installed draws the recorded room. Each dispatch line puts the
recorded seats on the recorded turn through `startTurn` and `holdTurn` (a `turnState` with no
handle and a no-op cancel); each event line goes through `applyEvents`; each gate line takes the
card down the way `decideGate` does, with the wait charged from the recording's clock.
**The clock is the recording's, and `Render` stays pure.** `State.Now` on a tick is the
recording's start plus the wall time elapsed since the replay opened, times `--replay-speed`;
each record stamps its own offset on the same clock and never moves it backwards. The two wall
reads the live retirement path makes — `finishColumn`'s Elapsed and the charged gate wait —
are stamped from the recording before the event lands, using the guards those paths already
have. A fixture played through `Update` therefore renders the same bytes every time
(`replay`, `replay-gate`, `replay-ascii` goldens), and two plays of one file are one frame.
**Labelled on every frame, in words.** `⚠ REPLAY` in the header where WRITE/READ sits;
`REPLAY` first on every column's badge row (the whole row at strip width); `⚠ REPLAY nothing
here is live` where the compose footer's routing cell and `enter dispatch` would be;
`ctrl+c quit` on the view line. A replay is a real run played back — not §8's invented
recording, because somebody ran it and every event was measured — and the label is what the
honesty rule asks in return: a replayed frame must never pass for a live one.
**Three refusals, and nothing else.** Enter, the card's `y`/`n`/`a`, and the per-seat verbs
say `this room is a replay; nothing here is live`. `ctrl+c` and `q` quit in every mode.
`saveRoom` returns on the replay flag, so a finished turn and a teardown write nothing;
`TestAReplayNeverTouchesRoomJSON` plants a sentinel room under the suite's sandbox home and
reads it back unchanged. No spawn var is reached: `Init` starts the feed and nothing else,
which is the spawn guard's own "no test calls Init" argument met from the other side.
#### What a recording does not hold
The operator's cancels and give-ups (a cancelled column replays as the vendor's own exit),
focus and scrolling, the `--brief` file's text, and `--trace`'s clocks. A recording with a long
idle gap replays the gap at the recorded pace; `--replay-speed` is the remedy, not a cap.
#### Verification
`gofmt`, `go vet ./...`, `go test ./... -count=1`, `go build ./...`, the windows/amd64 and
darwin/arm64 cross-builds, and `go test -race ./internal/council`, all clean. Goldens new:
`containment-badges`, `containment-badges-ascii`, `containment-badges-expanded`,
`flow-fan`, `arena-record-seats`, `arena-record-seats-ascii`. Goldens moved: the two
`slash-refusal` frames, on one line — the refusal's vocabulary gained `/hand` and paid for
it with "room" and an article. No other existing golden moved: a column dispatched before
this section carries no containment claim, so no frame built before it changed.
**Not verified here: a live crew in its trees.** No vendor was run by the session that wrote
this. What only a live run can show: a persistent seat resuming its thread after the respawn
that moves it from the workspace into its tree (the same `--resume` composition `/cd`
measured, in a directory that did not exist a second earlier); a one-shot seat honouring the
tree as cwd on a resume (`codex resume` rejects `--cd`, and the tree is passed as `Dir`);
two seats writing at once into two trees and `/adopt` folding each in turn; the fanned
`/flow` stage's two artifacts landing beside each other and the join reading both; and the
`⚠ shared tree` badge at real widths on a real fallback — SEEN 2026-09-03 on every seat of a five-seat room opened in `~` (not a git repo) at about 180 columns, with the `· ` clause shed at that width; the clause itself at a wider room is still unchecked. The rest owed and named rather than implied,
on §9.53's rule.
darwin/arm64 cross-builds, and `go test -race ./internal/council` once. New goldens:
`route-headroom`, `route-auto`, `replay`, `replay-gate`, `replay-ascii`. No existing golden
moved: a room with no near-full reading and no `@auto` draws the routing cell it drew before,
and a live room's header, badge rows and footer are untouched by the replay flag.
**Not verified here: a live recording.** No vendor was run by the session that wrote this, so
a recording of a real room EXISTS as of 2026-09-03 — the owner ran `--record demo.jsonl` over a five-seat gated write room at turn 3 with one `@all` brief, and `replay-check` read it clean: 75 records over 24s, five seats, one dispatch, 59 text events, no tool calls, no gate cards (the file carries verbatim prose and session ids, so it stays local). `--replay` of that file has NOT yet been driven; the replay fixture in the suite is still synthesized (fake ids, fake paths,
three seats, one card, one finished seat). What only a live run can show is a `--record` of a
five-seat room with real timing, that its replay reads as the room did, and what `replay-check`
lists off a real capture. The hero decision stays the owner's.
**Measured 2026-09-03: `--replay` of a real recording, and it found a defect.** The owner
recorded a four-seat gated write room (`--vendor claude,codex,agy,grok`) over seven dispatches:
1863 records over 39m36s, `replay-check` clean. `--replay` of that file drew every persistent
seat as `failed` right after each dispatch, with an elapsed figure of an hour and the note "the
vendor process ended mid-turn". The live room had shown those turns done. The mechanism, read
off the file: each dispatch to a persistent seat was followed 0.1s to 11s later by a `done`
from that seat, before its first text of the turn. That is the seat's PREVIOUS process ending
after the respawn that replaced it (`stopProc`). The live room attributes that exit to the
old process by one liveness test on the current one, and discards it. The recording carried
only the vendor name, so the replay handed it to `applyEvents` as the new turn's terminal
event.
Two fixes, and the reason for two. The recorder now writes from inside `applyEvents`, after
the stale-exit guard (`Model.staleExit`, the one liveness test the `KindDone` and `KindError`
branches used to make inline), so a file carries what the room applied and not what the
channel delivered. That is the fix at the source, and it is exact. The reader
(`markStaleExits`) repairs a file recorded before the guard, which the first real recording
is: a process's exit is the last event it emits, so a persistent seat's exit that another event
from the same seat follows before its next dispatch was not this turn's end. A replay consumes
such a line and does not apply it, and `replay-check` prints how many it will skip. The same
rule reads the first-turn retreat to a batch adapter correctly, where the room also discards
the exit and the seat goes on. Three tests pin it over a fixture with one stale exit and one
real death (`stale-exit.jsonl`).
Two smaller findings from the same file. The room line listed five seats and the room-open
rebuild reported `5/4 seats rebuilt`: a seat `--vendor` left out was still marked `Restored`
on reattach, so the rebuild launched it and the recorder listed it. `reattach` now marks a
seat restored only if `State.seats` says it takes turns, and the room line lists the columns
`VisibleColumns` draws. And `replay-check` printed one bare `grok` line per tool RESULT (an
act carrying an outcome and no text, 112 of them for one seat); those fold into one count per
seat. The replay half of the owed measurement is done; the hero decision still stays the
owner's, and the rendering of that recording belongs to the density pass.
**2026-09-16: the replay draws its provenance on the room line.** Before this date, the
recording's stamp lived only in `replay-check`'s stdout, and the scrubbed claim lived on the
entry notice and the closing notice. The entry notice is gone at the first dispatch, so a
reader who looked at a frame in the middle of a replay saw `REPLAY` and nothing that said which
run it was or whether the prose was real. The room line (`roomline.go`, `replayFact`) now
prints the provenance on every replayed frame, first among the room facts: `recorded
2026-09-03 21:14 -0400` for a capture, and `recorded 2026-01-01 09:00 UTC · scrubbed: the
shape is real; the date and every word are synthesized` for a scrubbed file. Each clause
follows §4a.1. The date is the file's own stamp in the zone the recorder wrote. A file with no
readable stamp draws no date; it never draws the epoch or this machine's clock. A scrubbed
file's stamp is `scrub.go`'s constant, so the clause says the date is synthesized with the
words. The vendor CLI versions are NOT drawn, because the format carries no version field
(`recordLine`) and a version read off this machine would describe a room that ran somewhere
else. A version field on the room line, written when the room learns one, is owed and is not
started here. The two notices keep their words. The cost is one room-line row on every replay
frame, so the `replay`, `replay-gate`, `replay-ascii`, `demo-gate`, `demo-final` and
`demo-ack` goldens moved by that row, and `demo-compare` and `demo-compare-route` are new (the
compare is `docs/room-identity.md`'s 2026-09-16 section; `panes-keys` moved by the `c compare`
cell and `panes-compare` is new). `TestTheReplayFactIsHonestAboutWhatTheFileCarries`
pins each clause.
### 9.57 three seats stay up between briefs, on a reading rather than a run (2026-09-02)
Until this change the room kept exactly two processes alive across turns: the Claude seat
(`claude -p --input-format stream-json`, §9.8, measured) and the Cursor seat (`cursor-agent acp`,
§9.36, measured). Codex, Antigravity and Grok each paid a whole process per brief — `codex exec
--json`, `agy -p`, `grok --single=` — and two of those cold starts were measured on the reference
box at **5.6 s** (Cursor, before §9.36) and **6.4 s** (Antigravity). A crew tool that dispatches a
brief every few minutes spends most of a seat's wall clock on startup that way, and a seat that
starts fresh every turn has no channel on which to be asked anything. This section records the
move of all three to long-lived processes, what each was built from, and — the part that matters
more — the fact that **none of the three shapes has been driven from this repository.** Every
badge on those columns says `unmeasured` and names the version it was read at, and the checklist
at the end is the price of removing the word.
**This is a departure from §9.50's ruling and says so.** §9.50 left the codex app-server protocol
unseated because a read-posture turn on it was measured failing to inspect in two of three arms,
and named the read posture's liveness as the measurement owed before any flip. That measurement
has not been made. The seat moved anyway, on the crew ledger above, with three things holding the
honesty line: the measured batch invocation stays one step away as the fallback, the seat owns
its own kill, and the badge says what nobody has watched. The owed measurement is now the FIRST
item of the checklist rather than a precondition, and that ordering is a decision, recorded here
so nobody reads the registry and assumes the debt was paid.
#### What was built, and from which pages
| seat | live shape | fallback (measured) | built from |
|---|---|---|---|
| Codex | `codex app-server`, one process, `thread/start{cwd, sandbox, approvalPolicy}`, `turn/start` per turn, `item/*/requestApproval` answered through the room's card | `codex exec --json` (codex.go, 0.149.1) | the protocol capture of §9.50 at **0.149.1**; the app-server README on `openai/codex` main and `app-server-protocol/src/protocol/v2/{shared,item}.rs`, read 2026-09-02, for the `approvalPolicy` enum (`untrusted`, `on-request`, `never`), the v2 decision enum (`accept`, `acceptForSession`, `decline`, `cancel`), the approval params (`itemId`, `command`, `cwd`, `reason`, `grantRoot`), and `turn.status: "interrupted"`. Installed build **0.152.1**, undriven |
| Grok | `grok agent stdio`, the ACP client of §9.36 under a second dialect, `session/new{cwd}`, `session/prompt` per turn, `session/request_permission` answered by kind through the room's card | `grok --single=` (grok.go, 1.0.4) | docs.x.ai/build/cli/headless-scripting and zed.dev/acp/agent/grok-build, read 2026-09-02, for the subcommand and the `--cwd` / `--resume` flags this seat deliberately does not pass; agentclientprotocol.com/protocol/schema for `loadSession`, the option `kind` values and the `cancelled` outcome. Grok Build **1.0.13**, undriven at the time; driven 2026-09-04 (the checklist below) |
| Antigravity | `agy --input-format stream-json --output-format stream-json`, one process, `{"event":"user","message":{"content":…}}` per turn, `result` ends the turn | `agy -p` (agy.go, 1.1.13) | antigravity.google/docs/cli/headless and the changelog entry for **1.1.15** (2026-08-19), read 2026-09-02, for the flag, the envelope, "one turn per message in a single conversation", and "close stdin … the process exits after the input pipe is closed and the current turn completes". Installed build **1.1.24**, undriven |
Three shapes in the package carry the move:
- **`vendors.LiveFallback`** names the measured batch adapter behind each live seat.
`FallbackRegistry()` is the whole room after a retreat, and it exists so that state can be
constructed rather than scripted — the give-up and re-send suites build "an ordinary one-shot
seat" from it, because the default registry no longer drives one. `Registry` became a var for
that reason, on the spawn vars' precedent.
- **`vendors.GracefulStop`** is the seat owning its kill. §9.50 measured a closed stdin NOT
reliably ending `codex app-server` (four exits in 1.5–3.3 s, one alive at 15 s), so the
app-server protocol's `Closing()` cancels any held approval and interrupts an open turn,
`Grace()` bounds the wait at 4 s, and the runner's job-object kill is unchanged behind both.
The ACP client and the Antigravity seat implement the same pair.
- **`acpDialect`** is everything the shared ACP client may vary per vendor, and it is three
fields: the read-posture mode id (cursor: `plan`, measured; grok: none), whether permission
answers use the measured option spelling or pick by kind from the request (cursor: fixed; grok:
by kind, `cancelled` when the kind is not offered), and whether `session/load` waits for the
server to advertise it (grok only). `cursoracp.go` is `acp.go` now, with every measured
sentence intact.
**Codex's approval policy is a choice, and the alternative is named.** Read posture asks for
`never`; write postures ask for `on-request`, the vendor's own interactive default, under the
`workspace-write` sandbox — so the vendor asks when it wants more than the sandbox allows and the
room cards that. `untrusted` would ask about every command off the vendor's trusted list, and the
pages read for this do not say whether a command approved under it then runs outside the sandbox.
A policy that might trade the sandbox for a keystroke is not one to adopt unmeasured. The gated
posture never reaches the seat as itself: `spawnPosture` collapses it to write for every
Conversational seat, and whether a request becomes a card or an automatic yes is the room's
decision when one arrives.
**Two things the codex badge did, in opposite directions.** On Windows the read level HOLDS at
`ro:enforced`, because the same 0.149.1 session measured the sandbox on the app-server path — a
write through cmd.exe denied with no file on disk — and the detail now carries that path's
sharper liveness residual. Off Windows the level DROPS to `ro:requested`: the macOS enforcement
was `codex exec`'s (2026-08-05, 0.146.0), every app-server arm ran on Windows, and a seat move
re-measures rather than inherits. That is the one badge this change lowered.
**What is unchanged, stated so it is not inferred.** `canGate` still names one seat. Two more can
now be ASKED — a request the vendor raises reaches a person — and neither has a coverage
measurement, so both stay `WRITES` with `asks · unmeasured` where the argument lives. Grok's cost
figure comes only from the `--single` fallback: the ACP prompt response is `{stopReason, _meta?}`
and nothing read says what grok puts in `_meta`, so on the ACP seat cost renders **absent**,
never zero. Antigravity's `--conversation` on the stream argv is unmeasured for composition, and
the §9.43 fork tell is what makes sending it safe: the seat implements `SilentResumeFork`, so the
room compares the id it asked for against the id `init` reports.
**What the core does with the two interfaces (landed 2026-09-02, in the crew integration).**
This paragraph said "nothing yet" while the six crew lanes ran apart; the core was patched once
they merged, and it now reads both. `LiveFallback`: a seat whose protocol reports `Dead()` — on
the brief that finds it dead in `handTurnToSeat`, or on the failed-turn event the handshake
refusal produces — and a persistent seat whose process dies on its FIRST turn before it names a
session, both retreat to the batch adapter for the rest of the room (`fallback.go`). The retreat
is on the SAME dispatch: the brief the operator typed goes down the batch branch with the
operating brief applied, the column stays on its turn with a note naming the invocation, the
badge reads `seatShape(v, true)` through `postureClaimFor`, and the seats in flight beside it are
untouched. `GracefulStop`: teardown, a respawn, and a cancel that has to become a kill all run
`stopProc` — `Closing()` down the pipe, `runner.Session.CloseInput`, `Grace()`, then the kill
that §9.50 measured necessary — with teardown waiting on every seat's grace at once. Both are
pinned over stubbed sessions in `fallback_test.go`; the live run each one owes is on the
checklist below, and STATE.md carries the consolidated list.
#### The live measurements owed, as a checklist
Each item names the command to run on the reference box and the sentence it would let the badge
change. Capture the wire under `vendors/testdata/wire/--.jsonl` per the
README there; a claim that is not captured is not made.
- [x] **Codex, read liveness (the §9.50 debt).** PAID. MEASURED 2026-09-04 at codex-cli 0.151.0, Windows 11: a hosted read room (`telltale council --host --read --cd --vendor codex`, driven through the plain client with stdin piped), the brief "list this directory and read README.md" three times; every turn ran `cmd.exe /c 'dir /a'` and `cmd.exe /c 'type README.md'` and settled `done (exit 0)`. Three of three inspected. The transcript is filed privately (desk/research, 2026-09-04).
The Windows detail's "two of three read turns" is retired by this run; it was true at 0.149.1
and is not true at 0.151.0 on this path. What this run does NOT re-measure: the sandbox's
refusal of a write under `-s read-only` (no write was attempted), so PARITY's `ro:enforced`
still rests on the 2026-08-29 measurement.
- [ ] **Codex, the approval flow, both branches.** A write-posture thread (`workspace-write`,
`on-request`) asked to write OUTSIDE the workspace and to reach the network: does
`item/commandExecution/requestApproval` or `item/fileChange/requestApproval` arrive, does the
vendor BLOCK until answered, does `accept` run it and `decline` stop it, and does the file
land or not. This is what lets `asks · unmeasured` become `asks · measured`.
- [x] **Codex, `turn/interrupt` and teardown order.** MEASURED 2026-09-05 at codex-cli 0.151.0,
Windows 11, five runs of `codex app-server` through the same `codex.cmd` shim `doctor` names as
the seat's entry point, with the seat's own `initialize`, `thread/start{sandbox:"read-only",
approvalPolicy:"never"}` and `turn/start` frames on this repo as cwd (transcript filed
privately, desk/research). **The interrupt half passes clean:** `turn/interrupt` on a running
read turn produced `turn/completed` with `status: "interrupted"` on every run, 0.10–0.13 s
after the frame. **The teardown half fails the grace:** with no turn live `Closing()` sends
nothing, and after stdin close the process exited 0 after 6.79, 1.76, 6.69, 7.77 and 6.42 s —
four of five over the 4 s grace, so the room's reaper ends this seat before it ends itself
most of the time. No stderr on any run. **DIAGNOSED the same day, same build, forty runs, stdin
closed with a thread open:** the shim is innocent (`cmd.exe /c codex.cmd`, `node.exe codex.js`
and the native `codex.exe` each took 6-8 s, five runs each); no shutdown frame exists in the
schema, and `thread/unsubscribe` and `thread/archive` before the close changed nothing; a process
with no thread exits in 0.03 s; with `--disable hooks` the same process exited in 0.06-0.08 s
with its MCP servers still configured, and with `mcp_servers={}` and the hooks kept it took
4.5 s. The cost is the operator's own `SessionEnd` hooks in `~/.codex/hooks.json`, one of which
took 4.5-6.6 s run by hand, and the server writes nothing to stdout while they run. The room may
not disable them (the fleet's guard hooks ride the same flag), so the fix is the last resort the
chip named, with the diagnosis attached: `Grace()` is 20 s, a bound on the operator's hooks
rather than on the vendor. Five runs after, through the shim with an interrupted turn: 4.46,
4.33, 14.73, 7.55 and 2.45 s, all inside it. The 14.73 s run is the same hook under load.
- [ ] **Codex, macOS.** The read sandbox on the app-server path, on the Mac, before the
off-Windows badge may return to `ro:enforced`. Record in PARITY.md.
- [ ] **Codex, the exec fallback trigger.** Run the seat against a build without `app-server`
(or an unauthenticated one) and watch what the room shows; this is the arm the room's retreat
(`fallback.go`) is built for, and the note and the `exec · unasked · fallback` badge are what
it should show.
- [x] **Grok, the handshake and a turn.** PAID 2026-09-04 at grok 1.0.13 (5e9a58528b76), `grok agent stdio` driven
with the seat's own frames from a scratch cwd (the transcript is filed privately, desk/research).
`agentCapabilities`: `loadSession: true`; `sessionCapabilities: {list, resume, close}`;
`promptCapabilities.embeddedContext: true`, image and audio false; `mcpCapabilities: {http, sse}`;
and an `_meta` block advertising `x.ai/hooks` with blocking events `pre_tool_use`, `stop`,
`subagent_stop` and decisions `deny`, `block`. `authMethods`: `cached_token`, `grok.com`.
**`session/new` at 1.0.13 returns `sessionId`, `models` and `_meta` — no `modes` and no
`configOptions`.** The frame this adapter's header quotes (a `modes{currentModeId:"agent"}`
block) was 1.0.4's; at 1.0.13 the server advertises no mode at all, so what `session/set_mode`
does for the read posture is now an open question (the permissions item below). A fenced
brief streamed back as `agent_message_chunk` (853 chunks), beside `agent_thought_chunk`,
`tool_call`/`tool_call_update`, `available_commands_update` and `session_info_update`, and
resolved `stopReason: end_turn`. **A brief beginning `/` is NOT eaten on this path**: `/help`
reached the model as text and was answered with a description of the TUI's slash commands.
stderr carried `BatchLogProcessor.ExportError … network error` lines throughout — grok's own
OTLP exporter pushing at a collector that was not running (§7.16a), harmless here.
- [x] **Grok, a permission request.** PAID 2026-09-04 at grok 1.0.13 (5e9a58528b76), Windows 11,
with the seat's own frames (the transcripts are filed privately, desk/research, one per arm;
`vendors/acp.go`'s header carries what each showed). Two findings, and a third the item did
not ask for. **The server does not ask by default.** A write session given "run `mkdir zzz`
and create probe.txt" raised NO `session/request_permission` in two trials; the wire showed
`_x.ai/session_notification{pending_interaction, kind: "permission"}` followed at once by
`interaction_resolved`, `_x.ai/sessions/changed` reported `"yolo":true`, and both trials
put `zzz` and `probe.txt` on disk. After the prompt `/always-approve off` — handled by the
agent as a host turn, no model call — the `write` tool DID raise the request: `toolCall.kind:
"edit"`, options `allow-edits-session` (allow_always), `allow-once` (allow_once),
`reject-once` (reject_once), cursor's spelling to the letter and the tool's `_meta` naming
the `opencode` namespace. `reject-once` ended the turn `stopReason: cancelled` with no
`probe.txt`; `allow-once` put it on disk; `mkdir zzz` through `run_terminal_command` asked in
NEITHER trial and landed in both. The seat sends no toggle, so its badge now reads `acp ·
unasked · measured at 1.0.13` and the write detail says why. **The read posture's mode.**
`session/new` advertises no `modes` at this build, and `session/set_mode` answers `{}` to
every id — `plan`, `agent`, `read`, `bogus-mode-xyz` — so the acceptance is no evidence;
only `plan` and `agent` echoed a `current_mode_update` in both runs. Under `set_mode plan`
the same write brief ran no shell and wrote nothing in the workspace, two of two: the agent
wrote its plan into `~/.grok/sessions//…`, raised the server request
`_x.ai/exit_plan_mode{planContent}`, the client's empty result for an unknown request was
read as "the user wants to revise the plan", and the turn ended. A read brief under `plan`
still ran `list_dir` and `read_file` and answered from the file. So `grokDialect.readModeID`
is `plan` since this date and the ACP seat's read badge is `ro:requested`, on exactly the
evidence class cursor's is: a mode the model obeys, two trials, never `ro:enforced`. The
`acpDialect` bullet above that says "grok: none" was true until this date. **`x.ai/hooks` is
not a wire protocol**: it is the agent's on-disk hook system (`~/.grok/hooks`, listed by the
agent-handled `/hooks-list`; every tool call produced a `hook_execution` notification naming
the operator's own hook). Declaring the block under `clientCapabilities._meta` changed
nothing — the initialize response was byte-identical and no hook-shaped request reached the
client — so a client cannot hold a posture through it. Also measured: closing stdin ends the
process in 2.2–4.4 s (exit 0, twelve of twelve), above the 2 s grace, so the kill lands
first; and a `/`-brief is eaten when its name is an advertised command and passed to the
model when it is not (`/help`, item above). Not measured: macOS, and whether the seat should
send `/always-approve off` itself so the room's card can gate this seat's file writes — that
is a design choice, recorded as open in STATE.md rather than made here.
- [x] **Grok, cost — HALF PAID 2026-09-04 at grok 1.0.13 (5e9a58528b76).** No `total_cost_usd` anywhere. But the
prompt response's `_meta` now carries `usage` with `inputTokens`, `outputTokens`,
`cachedReadTokens`, `cacheCreationTokens`, `reasoningTokens`, `modelCalls`, `apiDurationMs`,
a per-model `modelUsage` map keyed `grok-4.6-build`, and **`costUsdTicks`** (129,988,800 on a
two-call turn). That is a cost figure in a unit nobody has measured: a tick could be 1e-9 or
1e-8 USD and the two readings differ tenfold, so under §4a.1 the seat may render the TOKEN
counts and must keep the cost cell absent until the tick is measured against grok.com's own
billing page for the same turn. The adapter header's "no cost, and no token usage anywhere"
was true at 1.0.4 and is not true at 1.0.13; the header says so now. Owed: the tick's unit.
- [x] **Grok, `session/load`.** PAID 2026-09-04 at grok 1.0.13 (5e9a58528b76). `loadSession:
true` is advertised on every handshake and honoured: `session/load{sessionId, cwd,
mcpServers}` in a NEW process, from the SAME cwd, answered with `models` and `_meta` (no
session id, as the adapter assumes) after streaming the whole prior conversation back as
`session/update` lines — `user_message_chunk`, `agent_message_chunk`, `tool_call`,
`hook_execution` — which is the replay `acp.go`'s guard drops; a prompt after it recalled the
earlier turn's files by name. The refusal shape is `-32603 "Path not found."` with `data.code:
FS_NOT_FOUND`, and it is the SAME for an unknown id and for a known id loaded from a
different cwd: the store is keyed by cwd, so a moved room cannot resume. The process survives
the refusal and `session/new` answers in it, which is the branch the cursor capture measured
at `-32602`. The `costUsdTicks` unit (the cost item above) is still owed.
- [x] **Antigravity, the stream handshake.** PAID 2026-09-05 at agy 1.1.26, Windows 11, the seat's
own argv and line shape from a scratch workspace (transcripts filed privately, desk/research).
Two `{"event":"user",…}` lines down one stdin: one pid, the same `conversation_id` on both
`result` events (`num_turns` 1 then 2), and the second turn answered a question only the first
could ("copper"). stdin closed → exit 0 after **0.08 s**, well inside the 3 s grace. No stderr.
- [x] **Antigravity, `--conversation` under stream input.** PAID 2026-09-05 at agy 1.1.26: the id
from the run above, resumed in a NEW process with `--conversation ` after `--input-format`
— the SAME id came back, `num_turns` 3, and the word was recalled. Resumed, not forked, on this
path. The unknown-id arm (§9.43's fork, measured on `agy -p` at 1.1.11) was NOT re-run under
stream input and stays as `agystream.go` states it.
- [x] **Antigravity, `--print-timeout` under stream input.** PAID 2026-09-05 at agy 1.1.26, with a
stand-in: `--print-timeout 20s`, turn one answered, **35 s idle**, turn two answered on the
same pid. The bound is per turn and idle does not count against it; a 30m bound therefore does
not end a seat half an hour into a room. The 30m value itself was not waited out.
- [x] **Antigravity, the fallback trigger.** MEASURED 2026-09-05 at agy 1.1.26, and it does NOT
produce the shape the retreat keys on. With the seat's batch argv, no `-p` and no
`--input-format`, a stream line on stdin was read as the PROMPT: `init` with a fresh
`conversation_id`, three `step_update`s, then `result` `status: SUCCESS` with the answer the
JSON asked for, exit 0, **zero stderr**, 14.3 s. No first-turn death, no missing session — a
silent misparse. So on 1.1.26 the retreat is never triggered by this arm; it remains built for
a build that lacks the flag, and this reading confirms only that 1.1.26 is not that build.
#### Verification
No vendor was started. `go test ./internal/council/vendors -count=1` and the same with `-race`
pass over synthesized fixtures: the app-server handshake, both approval methods answered each way
in both vocabularies, the interrupt cancelling held approvals first, `Closing()` in order and
silent when idle, every fallback trigger reported by `Dead()`, the approval policy per posture;
the grok dialect's handshake, its `loadSession` gate against the cursor dialect's measured
behaviour, a permission answered by kind and cancelled when the kind is absent, the read
posture's own refusal, interrupt and closing order, no cost on the prompt response; the
Antigravity stream's argv, envelope, `result` ending the turn, the fork fixture replayed, and its
refusals. `internal/council/seatshape_test.go` pins every badge word and the fallback registry.
`go vet ./...`, `GOOS=windows GOARCH=amd64 go build ./...` and `GOOS=darwin GOARCH=arm64 go build
./...` are clean. The `help-postures` golden moved by exactly the codex detail's new sentences.
### 9.58 the codex seat failed in both rooms and said only `exit status 1` (2026-09-01)
Two live drives on 2026-09-01 and 2026-09-02 showed the same seat in the same state. The
hosted read room drew `codex — failed (exit 1)`. The gated write room drew `Codex ✗ failed 12s`
with `⚠ exit status 1` under it. Neither room showed a sentence, and neither room showed an
answer. The failure was in both postures and in both rooms, so it was not a posture defect and
it was not a host defect.
#### The measurement
The seat was last measured at codex-cli 0.149.1 ([§9.2](#s9-2)'s 2026-08-29 amendment). The
machine ran codex-cli 0.151.0. A chip re-ran the seat's exact argv outside council, from a
scratch directory, with the prompt on stdin, three times: the read seat's first turn, the write
seat's first turn, and the resume shape. Every run ended the same way:
```
{"type":"thread.started","thread_id":"..."}
{"type":"turn.started"}
{"type":"error","message":"You've hit your usage limit. ... try again at 11:45 PM."}
{"type":"turn.failed","error":{"message":"You've hit your usage limit. ... try again at 11:45 PM."}}
exit 1, stderr empty
```
Two facts follow from the capture, and they are different facts.
**The turn failed because the account had no quota.** That is the vendor's condition and council
cannot change it. The three argv shapes parsed and each produced `thread.started`, so no flag
moved between 0.149.1 and 0.151.0. The sandbox probes could not run on a turn that never
reached a tool, so §9.2's claims stay pinned at 0.149.1 and [STATE.md](../STATE.md) carries the
re-measurement as owed.
**The room lost the sentence because the vendor moved it.** Through codex-cli 0.147.0 a failed
turn wrote its reason to stderr and put zero bytes on stdout, and `codex.go` recorded that as
the reason it modelled no error frame: the runner's exit event carried the stderr tail, and that
was the whole failure signal. At 0.151.0 the reason rides stdout as a `turn.failed` frame and
stderr is empty. The adapter dropped the frame as an unknown type, on the rule that an unlisted
type is dropped rather than guessed at, and the exit event arrived with `exit status 1` and an
empty tail. Both rooms rendered exactly what they were handed.
**Which path this is, after [§9.57](#s9-57).** The drives above ran on a build that seated
`codex exec --json`. On the same day this section landed, §9.57 seated `codex app-server`
first and kept `codex exec --json` as the measured fallback. The app-server path reports a
failed turn on its own `turn/completed` with `status: "failed"`, as a `KindError` that ends the
turn, so the sentence reaches the room there already. The adapter change below is the fallback
path's, and the exit guard below is what the fallback needed: a spawn-per-turn seat is the one
whose failure sentence a process exit can overwrite.
#### What changed
**`turn.failed` is parsed, and its sentence is the card's note.** The event takes agy's shape
(`vendors/agy.go`): a `KindError` with exit code 0, no error and no `EndsTurn`, because the
process has not exited. That puts it on the branch [§9.33](#s9-33) built and
`failedturn_test.go` pins. The column settles at the vendor's sentence and the exit retires it.
**The `error` line is not parsed.** In the capture it always paired with a `turn.failed` that
carried the same text. Whether a lone `error` line ends the turn is unmeasured. A column failed
on a line that may be a recoverable hiccup would be the room inventing the verdict, and
`turn.failed` IS the verdict.
**The exit no longer overwrites the sentence, in either room.** A spawn-per-turn failure now
produces two events: the vendor's sentence, then the process exit. The second arrived last and
replaced the note, in `dispatch.go` and in `councilhost/room.go` alike. Both now hold one rule.
A process exit that lands on a column already failed with a note keeps that note, because only
the runner's exit event carries `Err`, and the only way a column is failed with a note before
its exit lands is a failure the vendor reported in its own stream. In the single-process room
the exit's own sentence moves to the note's detail line, so the card reads the vendor's reason
as its title and `exit status 1` as its body, and nothing the exit said is dropped. In the
hosted room the exit code was already on the head line, so the seat keeps its note and gains
`(exit 1)` beside the phase. A bare exit on a column nobody had failed is unchanged: it still
names itself.
**The failure stays `Unclassified`.** `runner.FailureClass` is grounded in strings captured off
a run that positively never reached the conversation. A usage-limit refusal says nothing either
way about the thread: the resume shape produced `thread.started` with the requested id and then
died the same way. Whether the room may keep a restored thread across that is a ruling under
[ADR-008](#adr-008)'s sixteenth amendment, and this section does not make it.
#### Verification
`go vet ./...`, `go build`, `go test ./internal/council -timeout 20m`, `go test ./internal/councilhost`.
The capture is pinned as `vendors/testdata/wire/codex-0.151.0-turn-failed.jsonl`, sanitized to
its thread id only, and `TestCodexFailedTurnWireIsPinnedAt_0_151_0` replays it. The adapter's
three new tests pin the verdict, the empty-payload verdict, and the dropped `error` line. The
dispatch test feeds the real adapter's event and then the runner's exit shape, and asserts the
note, the detail, the settle and the retirement. The hosted-room test does the same over
`Room.Apply`.
**Not verified here: a live turn that answers.** Every turn the chip ran died at the usage
limit, so the fix was checked against the captured failure and never against a codex turn that
succeeded at 0.151.0. The 0.147.0 fixture still replays green, which is the evidence that the
success path did not move in the parser. The first live turn after the account has quota is the
check, and it is the operator's.
### 9.59 the agy racer wrote its attempt into the vendor's own scratch directory (2026-09-03)
A recorded race (`2026-09-03-rebuttal.jsonl`, turn 9) asked four seats to add `haiku.md`. Claude,
Codex and Grok wrote the file in their worktrees. The Antigravity column said `no changes
against 5664d51` and ranked 3rd of 4. Its trace showed why: `write_to_file` on
`~\.gemini\antigravity-cli\scratch\sailboat-telltales\haiku.md`, then on `~\.gemini\antigravity-cli\scratch\haiku.md`.
The seat also listed its own scratch directory and read a task log under its own `brain\`
folder. The rank was honest. The seat never touched its worktree.
#### The measurement
The adapter set `Spec.Dir` to the worktree, and the vendor's `init` line reported that cwd. The
transcript agy stored for the turn holds the cause, in its own system prompt: *"The user does
not have any active workspace. If the user's request involves creating a new project, you should
create a reasonable subdirectory inside the default project directory at
C:\Users\sanle\.gemini\antigravity-cli\scratch."* Every past headless conversation on this box
that had no `--add-dir` carried the same sentence, back to the first arena on 2026-08-08. The
conversations that did name an active workspace were the interactive ones, opened in `~`.
Eight probes at agy 1.1.25, each one `agy --output-format stream-json --disable-slash-commands
--print-timeout 3m … -p ""` from the named directory:
| cwd | extra flags | where `probe.md` landed | turn |
| --- | --- | --- | --- |
| arena worktree (`.git` is a file) | none | `~\.gemini\antigravity-cli\scratch` | SUCCESS |
| the plain scratch repository (`.git` is a directory) | none | scratch | SUCCESS |
| a directory with no `.git` | none | scratch | SUCCESS |
| `code\telltale`, an exact `trustedWorkspaces` entry | none | scratch | SUCCESS |
| arena worktree | `--mode accept-edits` | scratch | SUCCESS |
| seat worktree | `--add-dir ` | **nowhere**: `write_to_file` reported DONE, no file, stderr `a tool required the "write_file" permission that headless mode cannot prompt for, so it was auto-denied` | CANCELED, empty response |
| seat worktree | `--add-dir --mode accept-edits` | **the worktree** | SUCCESS |
| arena worktree | `--add-dir --mode accept-edits --input-format stream-json`, one user line on stdin | **the worktree** | SUCCESS |
A ninth run resumed the CANCELED conversation with `--conversation --add-dir --mode
accept-edits` and asked for a second line. It edited the file in the worktree, echoed the same
id, and reported `num_turns: 2`.
So two facts, and each one needs its own flag. The cwd is no workspace to this vendor. Neither
git shape nor the trust list changes that, and `--add-dir` is the flag that names one. A write
inside a named workspace needs a permission that print mode cannot ask for, and `--mode
accept-edits` is the vendor's edit-only grant. A write into the vendor's own scratch directory
never needed the grant, which is why the defect looked like a wrong directory and not like a
denied write.
#### The fix
`baseArgs` in `vendors/agy.go` takes the workspace and the posture. It passes `--add-dir
` in every posture, and `--mode accept-edits` in the write postures. Both precede
`-p`, by the adapter's standing rule. The stream-json session builds from the same function.
`PostureWriteGated` takes the write argv: this seat has no gate channel, and the posture
contract says a gated seat may do anything write mode allows.
The read posture keeps the workspace named and the edits unaccepted. That is the closest thing
to a read posture this vendor has offered: the one in-workspace write measured under it was
denied. **The badge does not move.** A single denied write is not a sandbox. `run_command`
still runs under the operator's own allow rules in `settings.json`, and a write outside the named
workspace still lands, so the column stays `unsandboxed` and its detail now says which two flags
council passes and why. `--dangerously-skip-permissions` stays refused: the edit grant is not
the approve-everything class ADR-008's fifth and seventh amendments refuse.
One side finding, recorded in [§9.37](#s9-37)'s AGENTS.md table: with a workspace named, the
seat ran `list_dir` and then `view_file AGENTS.md` before it wrote, on a brief that never named
the file, and then wrote the file the brief in AGENTS.md asked for. That is the Claude row's
fact. The codename probe was not run.
#### Verification
`go vet ./...`, `go build ./cmd/telltale`, `go test ./...`. `TestAgyNamesItsWorkspaceOnArgv` pins
`--add-dir ` ahead of `-p` on the first turn and the resume, and pins that an empty
workspace sends nothing. `TestAgyAcceptsEditsOnlyWhenTheRoomWrites` pins the grant on both
write postures, on both paths, and its absence on read. The two posture tests now pin that the
grant is the WHOLE difference between the postures. `TestAgyStreamSessionKeepsEveryFlagAndNoPrompt`
pins both flags on the stream session.
**Verified live by the operator, 2026-09-04, race t13.** The section shipped saying the council
TUI takes no scripted input, so the racer's exact argv had been run by hand instead (rows seven
and eight above) and the live race was left as the operator's check. That check has now run:
four seats, the same `haiku.md` brief, in `~\Desktop\telltale-rooms\scratch`.
The Antigravity column reported `4th of 4 · done · 8m19s`, `committed 3c9df8c`, and a stat of
`haiku.md | 3 +++`. Its trace showed `write_to_file` and `view_file` on
`scratch-arena-t13-agy\haiku.md`, then four `run_command` steps inspecting the tree with `git
status` and `git diff`. On disk afterwards, `3c9df8c` on `arena/t13/agy` carries `haiku.md` and
three insertions, and the file holds the haiku. Agy's own scratch directory received nothing:
the `haiku.md` sitting there is still the turn-9 one, dated 2026-09-03.
**The rank is the honest part.** This seat placed 3rd of 4 before the fix, on a worktree it had
never touched, and 4th of 4 after it, on work it actually did — the slowest of the four at
8m19s, where the other three finished in 29s, 36s and 1m41s. The fix did not make the column
win. It made the column true.