# Troubleshooting **Say to your agent first:** *"Run a brain health check and fix what you find"* — this routes to the maintain skill, which runs `gbrain doctor` and either auto-fixes or prints the exact repair command; your agent can run the whole loop (*"Get my brain health score to 90"* uses the remediation planner with a cost cap). The sections below are for when you want the manual path. **Agents:** every gbrain error and refusal carries `code` and a `fix` whose `next` says who acts and whether to stop and ask. Follow the [agent operator protocol](../protocol/AGENT_OPERATOR_v1.md) first; `gbrain errors ` explains any code offline. This page covers symptoms that are not a single error. ## Symptom table **Who acts**: *agent* runs it; *agent, after the user agrees* means relay the consent and stop first; *user* means the user runs it in their own terminal; *brain host* means the operator of the machine that runs the brain. **Consent** lists the effects (`paid`, `destructive`, `credentials`, `egress`, `persistent_install`) that need the user's agreement; "none" means the agent can just run it. **Verify** is always read-only. | Symptom | Next step | Who acts | Consent | Verify | |---|---|---|---|---| | Recall returns conversation facts like "Date said …", "Attendees said …" or "Summary said …" | `gbrain repair conversation-labels` previews them ([conversation labels](repair.md#conversation-labels)); apply the printed command with `--yes` | agent, after the user agrees | `destructive` | `gbrain doctor --only conversation_label_facts --json` is ok | | Conversation facts are dated 1970-01-01 | same preview: `gbrain repair conversation-labels` (an undated page's pre-fix facts carry a `segment 1970-01-01` context); add a `date:` to the page frontmatter before re-extracting it | agent, after the user agrees | `destructive` | `gbrain doctor --only conversation_label_facts --json` is ok | | `remember` or `put_page` refused `source_changed` "The canonical file contains an uncoordinated local edit" on many pages nobody edited, or doctor `lost_caller_writes` warns | upgrade to v0.60.157.0 or later (a database-only v0.13.1 `validate: false` stamp was read as a file edit, #6429); then `gbrain repair failed-writes --source ` previews the refused writes and prints the apply command ([failed writes](repair.md#failed-writes)) | agent; the apply after the user agrees | none to preview | `gbrain doctor --only lost_caller_writes --json` is ok | | A write (`add_timeline_entry`, `put_page`, `remember`, any `mutating` tool) refused `invalid_params` "unknown parameter … write operations always reject unknown parameters; nothing was written" | the call named a parameter the tool does not declare (for example `content` instead of `detail`); the `suggestion` names the submitted key and the declared one; resend with the declared name. Reads keep the `mcp.strict_params` warn grace period | agent | none | the resent call returns no `invalid_params` and no `unknown_param` warning | | Facts extraction or embedding reports `(cause=AIConfigError model_not_found)` or `gbrain models doctor` says `model_not_found` while the key is valid | the provider refused the model, not the key (a 404, or a 403 "project does not have access to model"): check the model id with `gbrain models`, enable the model for the project in the provider console or set another model (`gbrain config set models.facts `) | user (provider console); agent for the config change, after the user agrees | none | `gbrain models doctor` reports the model `ok` | | A gbrain call failed with a code you don't recognize | follow `fix.next` ([protocol](../protocol/AGENT_OPERATOR_v1.md)); `gbrain errors ` | agent | as the fix's `consent` | the fix's `verify` | | A graph read omits a relationship you know existed, or says someone still works somewhere they left | the default read returns relationships true today: repeat with `status: "all"` or `as_of`; to record an end, add a dated timeline line (`Ended works_at [[companies/x]]`) or `add_link ... valid_until` ([temporal edges](temporal-edges.md)) | agent | none | `gbrain doctor --only edge_validity --json` | | A command exited 3 (`confirmation_required`) | relay `user_message`; run `fix.command` only after the user agrees | agent, after the user agrees | the payload's `effects` | the fix's `verify` | | [Database unreachable or `GBRAIN_DB_ACCESS `](#database-unreachable) | `gbrain engine status --probe`, then `gbrain db-repair` (diagnose); `gbrain db-repair --yes` applies safe fixes | agent; `--yes` after the user agrees | none to diagnose; `--apply-rewrites` rewrites the config URL (undoable) | `gbrain engine status --probe --json` | | [PGLite `RuntimeError: Aborted()` at startup](#pglite-aborted) | automatic repair on the next command; else `gbrain pglite-repair --dry-run`, then `gbrain pglite-repair --yes` | agent; `--yes` after the user agrees | `destructive` (a WAL backup is kept and the restore command printed) | `gbrain doctor --only connection --json` | | [`expected N dimensions, not M` on import](#embedding-dimensions) | `gbrain doctor` prints the exact `gbrain config set …` or `gbrain migrate embeddings --to ` command | agent, after the user agrees | `migrate embeddings`: `paid`, needs the brain's writer lock | `gbrain doctor --only embeddings,embedding_width_consistency --json` | | [Keyword-only recall; the user wants semantic search](#embedding-dimensions) | the readiness fix: `gbrain init --force --embedding-model --path ` (keeps pages and facts) | agent, after the user agrees | `credentials`, `paid`; needs the brain's writer lock | `gbrain doctor --only embeddings --json` | | [Doctor residue (`timeline_history`, `derived_visibility`, `safe_index_pending`)](#doctor-residue) | preview `gbrain doctor --remediation-plan --json` or `gbrain repair`; then `gbrain repair --apply` | brain host, after the user agrees | `destructive` (rewrites derived rows) | `gbrain doctor --only timeline_history,derived_visibility,safe_index_pending --json` | | Low health score | preview `gbrain doctor --remediation-plan --json`; then `gbrain doctor --remediate --yes --target-score 90 --max-usd 5` | agent, after the user agrees | `paid`; `destructive` with `--include-repairs` (approve with `--expect ` from the preview) | `gbrain doctor --json` | | [Sync held a file (`Held : invalid_frontmatter …`, `git_held_files`)](#held-files) | `gbrain sources status `, then preview `gbrain repair frontmatter --source ` and apply the printed `--apply --expect --yes` | agent; the apply after the user agrees | `destructive` (rewrites the previewed lines of the files) | `gbrain doctor --only git_held_files --json` | | [Sync held a file for its facts or takes fence (`Held : invalid_fence …`)](#held-fence-files) | nothing for most holds: the next maintenance run repairs and commits the fence. To see or do it now, preview `gbrain repair fences --source ` and run its printed `--apply --expect `; to pause automatic repair, `gbrain config set fences.repair.enabled false` (ask the user first). A `manual` reason needs the edit the hold names | agent; the apply needs no extra consent | none (model repair spends within the `fences.repair.*` caps; raising them is `paid`, the user's call) | `gbrain doctor --only fence_integrity --json` | | A managed catch-up on a live checkout keeps dying (`revision_conflict`, `pinned_git_worktree_conflict`, `page_identity_changed`, `write ECONNABORTED`), or you want an agent to drive the recovery | `gbrain sync status --source --json` (cursor, `committed_last_10m`, each hold and the last error with `class` / `safe_actions` / `needs_human`); when nothing moved and `needs_human` is false, `gbrain sync unblock --source --apply`, then the sync it prints ([runbook](sync-unblock-runbook.md)) | agent; a person only when `needs_human` names a slug | none | `gbrain sync status --source --json` | | [Sync held a file with `worktree_dirty`](#worktree-dirty-hold) | commit (or restore) the file, then `gbrain sync unblock --source --apply` and the sync it prints | agent; whoever edits the file commits it | none | `gbrain sources status --json` | | A managed sync of a source far behind HEAD prints `staged catch-up: synced to , short of HEAD`, or refuses `request_too_large` with `One commit exceeds the bounded cursor size` | Staging is automatic: `gbrain sync --source --no-pull` takes stage after stage (each the furthest first-parent commit whose delta fits 100,000 entries and 16 MiB) until it reaches HEAD, and a stop resumes at the stored cursor; a scheduled sync takes one stage per run. When one commit alone is above the bound, exclude part of it (`--exclude`, or `gbrain config set sync.exclude`) and rerun; `--full` above the bound still refuses | agent | none | `gbrain sync status --source --json` (cursor target reaches HEAD) | | The drain stopped `connection_lost` | check the database and pooler (`gbrain doctor --json`), then rerun the same sync; it resumes at the stored cursor | agent | none | `gbrain sync status --source --json` | | [Managed catch-up stuck, or `Held : preparation_stalled`, or the drain stopped `preparation_abandoned` / `preparation_systemic` / `write_capacity`](#catch-up-stuck) | `gbrain sources writer status --source --json` (what the stuck write was doing: step, `waiting_on`, owner, version), fix what it names or upgrade the owner, then `gbrain sources retry-held ` and the same sync with its options | agent; a restart or upgrade of the owner after the user agrees | none to inspect; `persistent_install` when the fix restarts a service | `gbrain sources status --json`, then `gbrain doctor --only git_held_files,persistence_write_stall --json` | | [The catch-up is parked: `data_moving: false`, doctor `managed_sync_not_moving`, or `gbrain sources writer movement` exited 1](#managed-sync-not-moving) | `gbrain sources status --json` prints the exact writer-status command; run it: `gbrain sources writer status --source --json` ends each running claim in a `claim.next` (`claim_running`: rerun after `retry_after_ms`; `claim_overdue`: `tell_user_to_run`, restart the named owner pid) | agent; the restart after the user agrees | none to inspect; `persistent_install` when the fix restarts a service | `gbrain sources writer movement ` (exit 0 when the data moves again), then `gbrain doctor --only managed_sync_not_moving --json` | | [The drain stopped `drain_stalled` with `cause: owner_wedged_here`](#managed-sync-not-moving) | before the ceiling: wait `retry_after_ms`, then rerun `next.command` (`safe_to_loop` is true); past it: restart the owner process the stop names (`gbrain serve`, the jobs worker or autopilot), then `gbrain sources retry-held ` and the same sync | agent; the restart after the user agrees | none to inspect; `persistent_install` when the fix restarts a service | `gbrain sources writer status --source --json` (the claim is gone or publishing), then `gbrain sources writer movement ` | | [Doctor warns `two_consumers_on_host`](#two-consumers-on-host) | `gbrain sources writer status --json` lists the consumers alive on this host (`host.consumers`: pid, kind, mode, age); let the shorter-lived one finish (a `gbrain sync` run ends on its own) or restart it so it defers to the resident one | agent; a restart after the user agrees | none to inspect | `gbrain doctor --only two_consumers_on_host --json` | | [Doctor warns `consumers_without_heartbeat`](#consumers-without-heartbeat) | upgrade and restart the process whose pid and kind the check names (it predates the heartbeat table and never defers) | brain host, after the user agrees | `persistent_install` when it rewrites services | `gbrain doctor --only consumers_without_heartbeat --json` | | [Doctor warns `host_identity_mismatch`, or maintenance jobs report `owner_unavailable` / `host_mismatch` for weeks](#host-identity-mismatch) | set `GBRAIN_HOME` on that process's supervisor to the value the check prints (the owner's home; config appends `.gbrain` itself) and restart it; never edit or delete a `host.json` | brain host, after the user agrees | `persistent_install` when it rewrites a service definition | `gbrain doctor --only host_identity_mismatch --json` | | [`fence_repair` or chronicle refused `owner_unavailable`, reason `host_mismatch`](#owner-unavailable) | `gbrain sources writer status --source --json` on the brain host (the owner host id); run the step there, or give the worker and the shell one `GBRAIN_HOME`. Never copy or regenerate `host.json`. Do not retry from the same host | brain host | none | `gbrain sources writer status --source --json` shows this host's id as the owner | | [`owner_unavailable`, reason `transfer_in_progress` or `clone_in_progress`](#owner-unavailable) | wait 30 s and retry the same step; the worktree is draining for a writer transfer or recovering from a topology clone | agent | none | `gbrain sources writer status --source --json` shows state `active` | | [A page write refused `owner_unavailable`, detail `unbound_source`, on a classic Postgres brain](write-refusals.md#unbound-sources-on-postgres) | don't retry; write the page as `.md` in the source's checkout, commit it, and run `gbrain sync --source ` (a remote agent gives those steps to the user); database-only pages or managed mode are the user's choice | agent, after the user agrees (remote: user) | none | `gbrain get --source ` | | [`owner_unavailable`, reason `binding_missing`, `incarnation_changed`, `local_path_missing` or `coordination_path_missing`](#owner-unavailable) | `gbrain sources writer status --source --json`; the registration is missing or outdated and re-binding or re-registering the checkout is the host administrator's decision | brain host | none | the same status read shows an active binding for this host | | [`gbrain repair fences` skipped a file `sync_in_progress` while a sync runs](repair.md#fence-repair-during-a-sync) | nothing: a write in flight or the running sync's frozen manifest still names that file; every other candidate repairs now and the next run picks this one up (`gbrain sync --source --no-pull` finishes the sync) | agent | none | `gbrain doctor --only fence_integrity --json` | | [A Google or GitHub item is held (`connector_held_items`)](#held-connector-items) | `gbrain sources status `, fix the cause, then `gbrain sources retry-held ` and `gbrain sync --source ` | agent, after the user agrees | `egress` (fetches from the provider again) | `gbrain doctor --only connector_held_items --json` | | [`gbrain migrate --to` refused with `writer_coordinator_required`](../ENGINES.md#engine-migration-refused) | follow the refusal's `fix`: relay its `user_message` (stay on PGLite and share with `gbrain mcp expose`, or leave the brain as it is) | agent, after the user agrees | `persistent_install`, `egress` (for `gbrain mcp expose`) | `gbrain doctor --no-migrate --json` | | [A write was refused with a named reason](#write-refused) | the reason's recovery in [write refusal reasons](write-refusals.md) | as the refusal's `fix` | as the refusal's `fix` | the refusal's `verify` | | [Managed sync blocked with `checkpoint_validation_timeout`](#checkpoint-validation-timeout) | `gbrain repair request-indexes --apply` when an index is missing or INVALID, then the printed `gbrain sync --source --no-pull --retry-failed …` | brain host | none | `gbrain doctor --only persistence_request_indexes --json` | | [Doctor warns `persistence_session_timeouts` (`session_timeouts_not_applied`)](#session-timeouts-not-applied) | `ALTER ROLE SET statement_timeout = '5min'` on the database, or accept it: preparation reads carry their own bound | brain host | none (a role default; `destructive` only in that it changes a server setting) | `gbrain doctor --only persistence_session_timeouts --json` | | `gbrain init` or `gbrain upgrade` stops with `schema_lock_blocked` naming a pid (or, before #5227, sat in `wait_event_type=Lock` with no output) | let the named session finish, or with the user's agreement end it on the database host: `SELECT pg_terminate_backend()`; then rerun the same command ([ENGINES.md](../ENGINES.md#schema-lock-blocked)) | brain host, after the user agrees | `destructive` when it ends another session | the rerun completes; `gbrain doctor --json` shows the schema at head | | [`queue_capacity`, `persistence_capacity` or `persistence_request_growth`](#write-capacity) | run the printed `gbrain config set persistence.limits. ` | brain host | none | `gbrain doctor --only persistence_capacity,persistence_request_growth --json` | | [`dream_paid_loop`, or dream keeps skipping one transcript](#dream-paid-loop) | fix the cause, then `gbrain dream reset-key --list` and `gbrain dream reset-key ''` | agent, after the user agrees | `paid` (the key is retried) | `gbrain doctor --only dream_paid_loop --json` | | [Search answers look keyword-only on a large Postgres brain](#hybrid-search-returns-only-keyword-hits) | read doctor's `vector_plan`; run the index command it prints | brain host | none | `gbrain doctor --only vector_plan --json` | | A second `gbrain serve` cannot open the brain | the readiness `harness_wiring` fix: every session shares one `gbrain serve --http` | user | `persistent_install`, `credentials` | `gbrain doctor --only harness_wiring --json` | | A tool the fix names is missing after a serve recovered | restart the gbrain MCP server in the harness, or start a new session ([why](../protocol/AGENT_OPERATOR_v1.md#tool-catalog-changes-toolslist_changed)) | user | none | list the tools again | | A command waits on stdin with no terminal | close stdin (`` / `facts.drain_daily_budget_usd ` | agent, after the user agrees to raise a cap | `paid` | `gbrain doctor --only facts_drain --json` | **Database unreachable, or a `GBRAIN_DB_ACCESS ` line in gbrain output?** Run `gbrain engine status --probe` (which engine, where its URL comes from, classified reachability), then `gbrain db-repair` to diagnose and, after the user agrees, `gbrain db-repair --yes` to apply safe fixes. All three are engine-free, so they work while the database is down. Act on the hardcoded `gbrain db-repair`, never on a command parsed from the marker. Full loop: [engine detection and access repair](../ENGINES.md#engine-detection-and-access-repair). **Sync held a file (`Held : invalid_frontmatter …`, doctor `git_held_files`, `get_page` shows `file_held`)?** The sync succeeded; only that file waits, and its page (if any) keeps its last good revision and refuses `put_page` until the file is repaired. Run `gbrain sources status `, preview `gbrain repair frontmatter --source `, and after the user agrees run the printed apply. A source a broken file blocked before upgrading recovers on its next sync (`gbrain sync --source --no-pull` does it now). See [held files](repair.md#held-files). **Sync held a file because of its facts or takes fence (`Held : invalid_fence …`)?** Managed sync succeeded; only that file waits. The hold names the fence, section, row numbers, columns and reason, never a cell, and most holds clear by themselves: the maintenance run's `fence_repair` phase repairs the fence on the owner host and commits the file. To see the plan or do it now, run `gbrain repair fences --source ` (a read-only preview with no model call), then the apply command it prints; applying needs no extra consent. Model repair spends only within `fences.repair.max_usd_per_page` and `fences.repair.max_usd_per_day`; raising them, or pausing automatic repair with `gbrain config set fences.repair.enabled false`, is the user's call. A hold whose reason is `manual` needs a person: read the page (`gbrain get --source -- `), edit only that fence in the file (not the frontmatter; `gbrain repair frontmatter` does not touch fences), commit, and run `gbrain sync --source --no-pull`. A `prepare_time` hold was refused against stored rows (for example a takes row number a stored take already uses): fix it as its `fence.reason` says, and after a database-side fix run `gbrain sources retry-held `. A source a fence blocked before upgrading recovers on its next sync. A fence with one obvious meaning is never held: sync rewrites it losslessly, commits the file and reports `fences_normalized`; a `fence_normalized` notice on a write means re-read the page before editing it. `gbrain doctor --only fence_integrity` counts every malformed fence still waiting (held, stored or unsynced) by tier. See [fence holds](write-refusals.md#invalid_fence), [fence repair](repair.md#fences) and the [fence format](fence-format.md). **Sync held a file with `concurrent_write`?** A Git edit raced a database-only write; nothing was overwritten. Preview `gbrain sources reconcile --preview`, let the user choose, then `gbrain sources retry-held `. See [concurrent_write](write-refusals.md#concurrent_write). **Sync held a file with `worktree_dirty`?** The working tree holds uncommitted bytes for the file that match neither the pinned commit nor the page (an agent mid-edit on a live checkout), so sync held it instead of overwriting the edit and kept going; nothing is lost and the run finished. Commit the file (or restore it), then `gbrain sync unblock --source --apply` re-screens every held file that is now committed and prints the sync to run; a later commit that changes the file re-screens it on its own. A file held three times is reported `needs_human` by `gbrain sync status`: find out what keeps editing it. Bytes committed at HEAD past the pinned target are imported as HEAD has them, never held. See [worktree_dirty](write-refusals.md#worktree_dirty) and the [sync unblock runbook](sync-unblock-runbook.md). **Managed catch-up stuck (one write `running` for minutes while the rest waits), held N files as `preparation_stalled`, or the drain stopped with `preparation_abandoned`, `preparation_systemic` or `write_capacity`?** Every managed write's *preparation* has a deadline (`persistence.sync_preparation_ms`, default 120 s, for sync members; `persistence.maintenance_preparation_ms`, default 120 s, for maintenance writes; a hard ceiling `persistence.preparation_ceiling_ms`, default 600 s; `persistence.max_preparation_attempts`, default 2; all behind the `preparation_deadlines` write switch), so one stuck write is cut off at its budget, but a cut-off preparation that ignores cancellation keeps its root blocked until the ceiling (600 s from the claim's start), when the owner frees the root and the next claim holds the entry; the rest of the source keeps publishing meanwhile, and a sync beside such an owner prints `stalled s on ` with the owner's pid and kind from the allowance to the ceiling instead of stopping. One wedge therefore costs up to ten minutes and one held file, never the run. The runbook branches on what you see; every command below is read-only unless it says otherwise, and the files themselves are never the problem in this branch. 1. **While it is stuck.** The progress line prints `stalled s on ` instead of an ETA. Run `gbrain sources writer status --source --json`: the blocked request's `claim.stall` names the step (`origin_check`, `knowledge_publication`, …), `waiting_on` (`git`, `fs`, `db`, `pool`, or an honest `unknown`), `step_age_ms`, the budget, and the owner process (kind, pid, gbrain version). An owner whose version predates this release never gives up: upgrade it and restart it (`gbrain serve`, the jobs worker or autopilot; the drain's own process ends with the drain). Otherwise wait out the budget: the write is released, retried once, then held. A release does not always free the root: a preparation that ignores cancellation holds its root barrier until the ceiling, and `writer status --json` ends the claim in a `claim.next` that says so (`claim_running` with `retry_after_ms` to the ceiling, or `claim_overdue` as `tell_user_to_run` naming the owner pid once the ceiling has passed or the owner reads `restart_required`). 2. **Held files (`preparation_stalled`).** `gbrain sources status --json` lists them with the step each stalled on; no `gbrain repair` applies. After the cause is fixed (or gbrain upgraded), run `gbrain sources retry-held ` and then the same sync with the same options, which the hold's `fix` and `retry-held` print, for example `gbrain sync --source --no-pull --no-embed`. If that sync still has a backlog, it finishes the backlog and then re-screens the scheduled files in the same run. Verify with `gbrain sources status --json` (the held count drops) and `gbrain doctor --only git_held_files --json`. A held deletion keeps its page until that retry. A stalled hold also re-screens by itself when the gbrain version changes. 3. **The drain stopped `preparation_abandoned` (exit 0).** A preparation in the sync's own process was still running past its allowance (its budget plus 30 s) and cannot be cancelled in-process, so the sync exited to end it; `drain.stall.step` says where. Rerun `next.command` (`safe_to_loop` is true): the next pass holds the entry if it stalls again. 4. **The drain stopped `preparation_systemic` (exit 1).** Too many writes stalled in one run (more than `sync.hold_escalate_count` / `sync.hold_escalate_pct` of the run's screened imports, or five in a row with no page committed between), so the run stopped with one diagnostic instead of holding every file. `breaker` carries the rule, the counts and the last step. Inspect the writer as in step 1, fix the cause, then rerun the same sync: it re-freezes the stopped entry with a fresh request and re-screens the files that run held. 5. **The drain stopped `write_capacity` (exit 1).** Other writers' requests held the sync writer's outstanding-request cap (`persistence.limits.principal_outstanding`, default 100; on one host the CLI sync shares the `local_cli` principal with fact-fence adoption and `dream`) for the whole 30 s window, so the sync could not admit its next write. It waited and printed `waiting for write capacity (N outstanding of M)` first; nothing failed and the cursor is intact. `gbrain sources writer status --source --json` lists the outstanding requests. Let them finish, or raise the cap (`gbrain config set persistence.limits.principal_outstanding 200`), then rerun the same sync. Reference: [drain stops](write-refusals.md#drain-write-capacity). 6. **A `put_page`, `remember` or maintenance write finished `preparation_stalled`.** Nothing is held: the receipt names the step. Read the writer as in step 1, then resubmit a foreground write under a new `request_id`; the next maintenance run submits its write again. 7. **A manual fence defect among the holds** (`invalid_fence`, reason `manual`, or `repeated_marker`) is a different branch: see [held fence files](#held-fence-files). `gbrain sources status` keeps the kinds apart and routes each to its own command. Turning the `preparation_deadlines` switch off (`gbrain config set persistence.preparation_deadlines false`) restores the previous behaviour exactly: no deadlines on sync or maintenance writes, no attempt counter, no `preparation_stalled` holds. See [`preparation_stalled`](write-refusals.md#preparation_stalled), [managed sync drain stops](write-refusals.md#managed-sync-drain-stops) and the [timing knobs](live-sync.md#catching-up-a-large-backlog-on-managed-postgres). **A Google or GitHub item is held after repeated failures (doctor `connector_held_items`, `gbrain waiting` says coverage is partial)?** Run `gbrain sources status `, fix the cause, then after the user agrees run `gbrain sources retry-held ` and `gbrain sync --source `. See [held items](google-connect.md#held-items). **A managed sync is blocked with `checkpoint_validation_timeout`?** Run the printed commands on the brain host: `gbrain repair request-indexes --apply` when an index is missing or INVALID, then the printed `gbrain sync --source --no-pull --retry-failed …`. Doctor's `persistence_request_growth` warns before lifetime request IDs run out. **PGLite crashes at startup with `RuntimeError: Aborted()` (often right after a macOS upgrade)?** Not a macOS incompatibility — the OS-upgrade reboot killed gbrain mid-write and tore the data dir's WAL. gbrain repairs this automatically on the next command (data preserved, backup kept); if auto-repair is disabled or skipped, run `gbrain pglite-repair --dry-run` to diagnose and `gbrain pglite-repair --yes` to repair in place. Full recovery ladder (repair → rebuild → engine switch) in [`docs/ENGINES.md` — Troubleshooting: startup abort](../ENGINES.md#troubleshooting-startup-abort-runtimeerror-aborted) and [`docs/INSTALL.md`](../INSTALL.md#pglite-crashes-at-startup-runtimeerror-aborted). **`gbrain import` fails with `expected N dimensions, not M`?** Run `gbrain doctor`. It will print the exact `gbrain config set ...` or `gbrain migrate embeddings` command to repair the mismatch (`migrate embeddings` re-embeds through the provider: effect `paid`, ask the user first). You should not need to delete `~/.gbrain`. Fresh `gbrain init --pglite` auto-detects your embedding provider from API keys: set `VOYAGE_API_KEY` (or `OPENAI_API_KEY` / another provider key) in the environment — or in `~/.gbrain/config.json`, which init also reads — before running init, or pass `--embedding-model :` explicitly. With multiple keys set, init fires an interactive picker (non-TTY auto-picks the Voyage default when its key is present). With no keys at all, init continues keyless (keyword-only search) with a loud notice; add a key later and re-run `gbrain init --force --embedding-model voyage:voyage-4` to enable embeddings (effects: `credentials` and `paid`, since every page is embedded through the provider; it needs the brain's writer lock, so stop a running `gbrain serve` first; pages and facts are kept), or pass `--no-embedding` up front to make keyless explicit. `gbrain config get embedding_disabled` reports whether embedding is off and which plane says so (the config file and the database setting; `true` on either turns it off); `gbrain config set embedding_disabled true|false` and `gbrain config unset embedding_disabled` change both, and the `init --force --embedding-model` enable path clears both. See [`docs/integrations/embedding-providers.md`](../integrations/embedding-providers.md) for the full provider matrix and [`docs/operations/headless-install.md`](../operations/headless-install.md) for Docker/CI sequencing. **`gbrain doctor` warns `default_source_local_path`?** Your `default` source has no `local_path` AND that null pointer is provably breaking write-through (the repo fallback is another source's own working tree, or file-backed default pages have no resolvable root). A null `local_path` on its own is the designed fallback topology and reports ok. The repair is a pointer update, never a file move: `gbrain sources set-path default ` prints the prior value before changing it and refuses a path that nests inside or swallows another source's tree (exit 6; `--force` bypasses). **Say to your agent:** *"Run a brain health check and fix what you find"* — the maintain skill runs `gbrain doctor` and applies the printed repair. **A Gmail, Calendar or GitHub connector source still has an old `local_path`?** Connector sources sync from their provider, so autopilot ignores their `local_path`: once a connector has synced (or tried to) at least once, autopilot syncs it on every interval and runs its database phases with no checkout. A connector that has never synced stays idle until you run `gbrain sync --source ` once; autopilot prints that command the first time it skips one. To remove the stale pointer, run `gbrain sources set-path --clear` (connector sources only; it refuses a filesystem source and a connector bound to a canonical owner). `config.syncEnabled=false` still opts a source out. **Say to your agent:** *"My Gmail source points at an old folder. Clear it and keep it syncing."* **A save, sync or background effect was refused with a named reason?** Reasons such as `file_database_drift`, `ambiguous_source_path`, `physical_root_device_changed`, `cursor_processing_options_conflict`, `take_row_collision`, `invalid_source_uri`, `queue_capacity` and parked effects (`targets_parked`, doctor `parked_effects`) each come with a recovery step in the error's `suggestion`. [Write refusal reasons](write-refusals.md) explains each one and its recovery. **Say to your agent:** *"My save was refused. Explain the reason and show me the fix before you run it."* **`gbrain doctor` warns `timeline_history`, `derived_visibility`, or unsealed pages under `contextual_retrieval_coverage`?** Preview the fix with `gbrain repair`, then apply one kind at a time with `gbrain repair --apply` on the brain host. See [repair residual damage](repair.md). **`gbrain doctor` warns `timeline_orphans`?** Timeline rows from an earlier version of a page are still in the database after the dated bullet was edited or deleted. Preview with `gbrain extract timeline --prune-orphans --dry-run`, then remove them with `gbrain extract timeline --prune-orphans` (add `--source-id ` to limit it). Rows no page version ever produced, such as enrichment and meeting fan-out, are kept. **A doctor check says "Not verified"?** The check could not read its input, so it reports `warn` instead of `ok`; `gbrain doctor --json` carries `details.code: "not_verified"` and the underlying reason in `details.reason`. Fix the cause, then re-run `gbrain doctor`: | Check | What it could not read | Fix | | --- | --- | --- | | `multi_source_drift` | a source's `local_path` root or a directory below it (`details.unreadable_sources`), the walk hit its bound (`details.limit` files / `details.timeout_ms`), or, for a source whose slugs are pinned to its git root, where `local_path` sits in its git work tree (`details.git_root_skipped`) | fix the path or permissions (`gbrain sources status`); for a large source re-run with `GBRAIN_DRIFT_LIMIT= GBRAIN_DRIFT_TIMEOUT_MS= gbrain doctor`; for a git-root source, check that `git -C rev-parse --show-prefix` succeeds (a checkout git can read and that it trusts) | | `embed_staleness` | the stale-chunk count (the embed worker's own predicate) | the reason names the database error; re-run `gbrain doctor` once it is fixed | | `schema_pack_consistency`, `schema_pack_source_drift` | the pages or config query, or a source's active schema pack | `gbrain schema lint --with-db` runs the same classification locally; `gbrain schema active` debugs pack resolution | `bootstrap_push_health` and `gbrain bootstrap status` report only the push record of the workspace named by this machine's bootstrap receipt; another workspace's stale or failed push is listed as such (a warn naming that workspace), never as this workspace's state. `reranker_health` auth warnings are audit-log history (`details.live_probe_performed: false`), not a live check of the key. These warns can appear while every other check reads green; each one names a check that did not run. **`gbrain doctor` warns or fails `git_convergence`, or `bootstrap_push_health` warns that a workspace is ahead?** A Git checkout the brain syncs from (a source's `local_path` or `sync.repo_path`) has commits its upstream does not have, or uncommitted changes, so the system of record on the remote is behind the brain (#5063). Commits not on the upstream warn after 6 hours and fail after 24; uncommitted changes older than 6 hours warn. Doctor compares with the local `@{u}` ref and never fetches, so it is only as fresh as the last fetch; checkouts without an upstream and connector sources are skipped. Commit and push the checkout (`gbrain sources push --path ` for a bootstrap workspace, `git push` otherwise). `bootstrap_push_health` no longer reports `ok` on a recent successful push while the workspace still has commits not on origin. **`gbrain doctor` warns `tracked_ownership_marker`, or `gbrain sources push` / `bootstrap verify` refuses a `.gbrain-owner*.json` or `.gbrain-bootstrap.lock/` path?** An ownership marker (`.gbrain-owner.json`, the `.gbrain-owner-.json` reservation beside a root, a `.gbrain-owner.json..tmp` staging file) or the bootstrap lock directory was committed to a source checkout (#5186). These files are this machine's claim on the checkout, never brain content, and must not reach a remote. Since this fix gbrain appends their patterns to the checkout's `.git/info/exclude` whenever it stamps or adopts a root, the bootstrap `.gitignore` template lists them, and the push deny list refuses a tracked one while an untracked one is excluded from the commit (`excludedUntracked`). None of that un-tracks a file already in the index: run the `git -C rm --cached ` the check prints (the files stay on disk, gbrain keeps reading them), commit, push, then `gbrain doctor --only tracked_ownership_marker --json`. A root stamped before this fix gains the exclude rule on its next stamp or transfer; until then the deny list is the backstop, and `bootstrap verify`'s `deny_globs` treats a merely-on-disk marker as fine. If a marker already reached a remote, rotate it the same way you would any published credential: the token alone is not remote authority (the reservation that proves ownership lives outside the root), but treat the history as exposed. **`brain_score` shows a low "timeline density (entity and event pages)"?** The 15-point timeline component grades only linkable pages whose type's active-pack primitive is `entity` or `temporal` (people, companies, meetings, emails, events…); reference documents such as notes, writing and guides have no events and are not graded, so do not stamp "page created" rows onto them. Types the pack does not declare are still graded. Raise the score by giving those entity and event pages real timeline entries (`gbrain extract timeline`). **`gbrain doctor` warns `slug_collisions`, or sync prints slug collisions?** Two or more files in a source map to the same page slug (for example `notes/Foo Bar.md` and `notes/foo-bar.md`), and only one is indexed. Rename all but one file in each group, commit, then sync. **Doctor warns `persistence_session_timeouts` with `session_timeouts_not_applied`?** gbrain sets `statement_timeout` (`GBRAIN_STATEMENT_TIMEOUT`, default 5min) and `idle_in_transaction_session_timeout` as connection startup parameters, and a transaction-mode pooler (PgBouncer with `ignore_startup_parameters`, Supavisor) drops them: `SHOW statement_timeout` through the configured URL reads `0`, so a statement outside a transaction has no server-side bound and a relation lock held in another session (a migration's `ALTER TABLE pages` queued behind a long transaction, for example) holds it for as long as the lock lasts. Managed write *preparation* reads are not affected since v0.60.108: they run under a transaction-local `statement_timeout` of the write's remaining preparation budget, so a lock-blocked read ends at the budget and the member is released `preparation_deadline` with no statement left on the server (#6278). For every other autocommit statement, set the default on the role instead of the connection: `ALTER ROLE SET statement_timeout = '5min'` (read-only to check: `gbrain doctor --only persistence_session_timeouts --json`). Setting `GBRAIN_STATEMENT_TIMEOUT` has no effect through such a pooler. **The catch-up is parked (`sources status --json` says `data_moving: false`, doctor warns `managed_sync_not_moving`, `gbrain sources writer movement` exited 1), or the drain stopped `drain_stalled` with `cause: owner_wedged_here`?** Liveness is not movement: `/health` can be ok, the pid alive and `sync_running: true` while no page has committed for weeks. A managed source is *not moving* when an unfinished `managed-sync` cursor has had neither a committed `managed_sync_*` receipt nor a head step advance for longer than `persistence.preparation_ceiling_ms` (600 s) while a drain or a full consumer is live on the owner host; with nothing live it is `movement_state: parked` (a cursor left between cron runs; informational, not scored). Every surface routes to the same two read-only calls, so the agent reaches the owner, the step and the next action without SQL: 1. `gbrain sources status --json`: `data_moving`, `not_moving_since`, `movement_state` and, beside `data_moving: false`, the exact `gbrain sources writer status --source --json` to run next. Doctor's `managed_sync_not_moving` fix is that same command. 2. `gbrain sources writer status --source --json`: every running claim ends in one `claim.next` envelope (`code` `claim_running`, `claim_overdue` or `claim_lapsed`). `claim_running` carries `retry_after_ms` and a `run` of the same status command while the owner is inside its allowance (the owner's own budget will cut the write; nothing to do but rerun after the delay), `tell_user_to_run` naming the owner `kind` and pid past the ceiling or when the owner reads `restart_required` (restart that process), `run` the printed resume command for a lapsed claim. `host.consumers` lists the processes on this host holding a consumer (pid, kind, mode `full|waiter_only`, started_at, the pool gauge), so a second consumer is visible in the same call. 3. After the restart (or after waiting out the ceiling), `gbrain sources retry-held ` re-screens the entry the owner held and prints the sync to run; `gbrain sources writer movement ` then waits one window and exits 0 only when a page committed (`moved`) or nothing is pending (`current`). See the [transcript](#managed-sync-wedge-transcript) below for what each step prints. A `drain_stalled` stop with `cause: owner_wedged_here` is the same situation seen from the sync CLI: a live owner on this host holds the head write past the ceiling or its heartbeat row reads wedged. Before the ceiling the stop is loop-safe (`next.safe_to_loop: true`, `retry_after_ms` set); past it the stop names the pid to restart. Reference: [drain stops](write-refusals.md#drain-stalled). **Doctor warns `two_consumers_on_host`?** Two resident gbrain processes on this host (`serve`, `sync`, `jobs`, `autopilot`, `mcp`) have each run a *full* persistence consumer for more than 30 s, so both claim writes on the same roots. One full consumer per host is a preference, not a fenced role: a `serve` always starts full; every other resident kind probes the `persistence_consumers` heartbeat table at start and on every tick and runs *waiter-only* (it submits and waits, claims nothing) while a live, full, not-wedged consumer of this host exists, promoting itself when that row lapses (60 s), reads wedged or stops being full. Two processes that start within one renewal of each other can both start full, and a sync CLI that took its consumer before the `serve` started keeps it for that run; doctor lists both until one exits. Rows younger than 30 s and short-lived kinds (`put`, `import`, `dream`, `cycle`, `sources`, `cli`) are listed, never warned. `gbrain sources writer status --json` shows them under `host.consumers` (read-only); let the shorter-lived process finish, or restart it so it defers. With `persistence.single_consumer` off (`gbrain config set persistence.single_consumer false`, or `GBRAIN_SINGLE_CONSUMER=0`) every process keeps its own consumer as before and the warning describes the configured behavior. `gbrain sync --no-delegate` keeps that one run's own consumer on either engine. **Doctor warns `consumers_without_heartbeat`?** A running claim on this host is stamped by a process (`claim_phase.owner`: kind, pid, gbrain version) that has no `persistence_consumers` row. Every full consumer on this release renews its row every 10 s, so an owner without one runs a gbrain release from before the heartbeat table: it never defers to the resident consumer and the host runs two full consumers until it is upgraded. The check names the pid, kind and version from the stamp; upgrade and restart that process (after the user agrees). A new sync CLI beside an older `serve` says so on its start line (`owner row missing: older serve or no serve; running own consumer`). **Doctor warns `host_identity_mismatch`, or fence repair, chronicle and other maintenance jobs have reported `owner_unavailable` with reason `host_mismatch` for days?** Host identity is a file, `host.json` under the persistence home (`GBRAIN_HOME`, else `$HOME/.gbrain`). A job worker or container launched with a different `HOME` or without `GBRAIN_HOME` mints a new identity after every restart and sees the binding as another host's, so its writes never reach the owner while every health signal stays green. Doctor flags it when this process's identity differs from the binding owner's and the owner's checkout resolves on this filesystem (same machine id or volume), or, for a container without a stable machine id, when the checkout is present at the binding's path and this `host.json` was minted after the binding under a different `HOME`/`GBRAIN_HOME`. The check names both files, the environment each was minted under (`minted_under: {home, gbrain_home, hostname, machine_id}`; `unknown` for a file minted before it was recorded) and the one-line fix: set `GBRAIN_HOME` on that process's supervisor to the owner's home (the printed value; config appends `.gbrain` itself) and restart it. Nothing re-derives or rewrites an existing `host.json`: re-owning the binding would break every other host sharing the brain. Verify with `gbrain doctor --only host_identity_mismatch --json`. **What the agent sees while a catch-up is wedged, step by step.** The brain is `default`, the owner is a `gbrain serve` with pid 4121, the sync CLI runs beside it, and the head write of a 16-member bulk group has been `preparing` with nothing in flight at the database. Output is abbreviated to the fields that decide the next step. ```text $ gbrain sync --source default --no-pull --no-embed [sync] managed catch-up: 13627 entries frozen, 13412 remaining; publishing in bulk groups (each page keeps its own request). [sync] 215/13627 processed · stalled 160s on import_screen (waiting on db) · allowed 2m30s · owner serve pid 4121 [sync] 215/13627 processed · stalled 430s on import_screen (waiting on db) · allowed 2m30s · owner serve pid 4121 Managed sync blocked: 215 entries this run (215 written, 0 waived), 13412 remaining. Oldest unfinished request 01J9W6Y5D0X7K2Q8R4T1V3B9ZC (running), step=import_screen, waiting_on=db, cause=owner_wedged_here, owner serve pid 4121; claimable here: no. Next: gbrain sync --source default --no-pull --no-embed (safe to rerun in a loop) Why: The head write is held by gbrain serve (pid 4121) on this host, parked at step import_screen for 430s. Its budget frees the root by the 600 s ceiling (170 s from now) and the next pass holds the entry; rerun then. If it is still held after the ceiling, restart pid 4121. $ gbrain sources status default --json { "id": "default", "managed_sync": { "index": 215, "total": 13627, "held": 0 }, "sync_running": true, "data_moving": false, "not_moving_since": "2026-10-08T15:02:11Z", "movement_state": "not_moving", "next": { "command": "gbrain sources writer status --source default --json", "why": "No page of this source committed and the head write's step did not advance for 11 min while a consumer is live on its owner host." } } $ gbrain doctor --only managed_sync_not_moving --json { "name": "managed_sync_not_moving", "status": "warn", "message": "Source default has 13412 unfinished entries and a live consumer on this host, but no page committed and no step advanced for 11 min.", "details": { "source_id": "default", "data_moving": false, "not_moving_since": "2026-10-08T15:02:11Z", "movement_state": "not_moving" }, "fix": { "command": "gbrain sources writer status --source default --json", "next": "run", "actor": "agent", "consent": [] } } $ gbrain sources writer status --source default --json { "host": { "consumers": [ { "pid": 4121, "kind": "serve", "mode": "full", "started_at": "2026-10-08T14:50:03Z", "pool": { "checked_out": 10, "max": 10, "waiters": 3 } }, { "pid": 4388, "kind": "sync", "mode": "waiter_only", "started_at": "2026-10-08T14:51:20Z" } ] }, "blockers": [ { "request_id": "01J9W6Y5D0X7K2Q8R4T1V3B9ZC", "operation": "managed_sync_import", "state": "running", "claim": { "phase": "preparing", "step": "import_screen", "step_age_ms": 661000, "waiting_on": "pool", "owner": { "kind": "serve", "pid": 4121, "build": "0.60.113.0" }, "pool": { "checked_out": 10, "max": 10, "waiters": 3 }, "stall": { "reason": "preparation_overdue", "step": "import_screen", "budget_ms": 120000 } }, "next": { "code": "claim_overdue", "retry_after_ms": 0, "fix": { "next": "tell_user_to_run", "actor": "host_admin", "consent": [], "user_message": "gbrain serve (pid 4121) has held this write at import_screen for 11 min, past the 600 s ceiling, waiting on its connection pool (10 of 10 checked out, 3 waiting). Restart that serve.", "why": "The owner's budget cut the write at 120 s, but the preparation ignored cancellation and still holds its root barrier; only the owner process can release it now.", "verify": { "argv": ["gbrain", "sources", "writer", "status", "--source", "default", "--json"] } } } } ], "data_moving": false, "not_moving_since": "2026-10-08T15:02:11Z", "sync_running": true } # after the operator restarted pid 4121: $ gbrain sources retry-held default Re-screening 1 held entry of source default. Run: gbrain sync --source default --no-pull --no-embed $ gbrain sources writer movement default Judging movement for 300 s (preparation budget 120 s + 60 s, floored at 300 s)... default: moved (3 pages committed, head advanced import_screen -> knowledge_publication). Exit 0. ``` Two calls, no SQL: the doctor or `sources status` row names the command, and `writer status` ends in the action. A `data_moving: false` that returns after the restart means a different owner is wedged or the same one wedged again: the same two calls name it. **A managed write stays `running` and later writes queue behind it, or doctor warns `persistence_write_stall`?** The owner keeps renewing a hung claim; doctor warns past `persistence.max_claim_ms` (default 600000) and names the request, root, age, phase and, for a preparation, its step and what it waited on. A write stuck *preparing* is released at its deadline and finally held or failed `preparation_stalled` on its own: see [catch-up stuck](#catch-up-stuck). A write stuck *publishing*, or one whose owner predates the deadlines, still needs the owner restarted: run `gbrain sources writer status --source --json` (read-only), then restart the `gbrain serve` that owns the root and attach that output when reporting it. **`fence_repair`, chronicle, fact-fence adoption or another maintenance step refused `owner_unavailable`?** A managed source publishes only through the host that owns its checkout, and the refusal's `reason` says which condition this is; the fix is always the read-only `gbrain sources writer status --source --json`, and nothing claims or transfers ownership to get a step through. | `reason` | What it means | Next | |---|---|---| | `host_mismatch` | Another host id owns the checkout. The message prints both ids' first 8 characters and, to a local caller, the `host.json` this process read (`/.gbrain/persistence/host.json`). A maintenance worker (launchd, cron, autopilot) started with a different `GBRAIN_HOME` or user than the shell that ran the sync reads a different identity file and is, to the brain, a different host. `retryable: false`: a retry from the same environment cannot change the answer. | Run the step on the owner host writer status names, or make both environments use one `GBRAIN_HOME`. Never copy or regenerate `host.json`: a new identity orphans the ownership and pending writes recorded for the old one. | | `transfer_in_progress` | The worktree is `draining`: a writer transfer was prepared and not yet accepted or cancelled. | `fix.next` is `wait`: retry after 30 s; the status read shows the state. | | `clone_in_progress` | The worktree is `recovering`: a topology clone or reclone is in progress. | `wait`: retry after 30 s; it is `active` when the clone finishes. | | `binding_missing` | The source has no canonical worktree binding at all. | Claiming a checkout is the host administrator's decision. | | `incarnation_changed` | The binding belongs to an earlier incarnation of the source (removed and re-added, or restored). | Re-bind the checkout on the brain host. | | `local_path_missing` | This host owns the worktree but has no checkout path registered. | Register the checkout again on the brain host. | | `coordination_path_missing` | This host's registration has no coordination directory. | Repair the registration on the brain host. | | `not_sent`, `outcome_unknown` | The local owner's IPC lane did not take the write, or lost its acknowledgment. | `gbrain sources writer status --probe --json`; for `outcome_unknown` read the receipt (`gbrain write-request -- `) before resubmitting. | A running sync is not one of these: `gbrain repair fences` now runs during a catch-up for every candidate the sync does not still name ([fence repair during a sync](repair.md#fence-repair-during-a-sync)), and a candidate it does name is `sync_in_progress`, not `owner_unavailable`. **Managed writes refused with `queue_capacity`, or doctor warns `persistence_capacity`?** A per-principal or per-brain write-journal limit (`persistence.limits.*`) is full (`queue_capacity`), or doctor sees lifetime request IDs or receipt bytes at 80% or more (`persistence_capacity`). For the cumulative limits (lifetime request IDs and receipt bytes), the warning and the refusal print a `gbrain config set persistence.limits. ` sized for about one more year; run it on the brain host, then retry with the same request ID. For outstanding-request or queued-byte limits, let outstanding writes finish and check `gbrain sources writer status`. Receipt compaction age is `persistence.receipt_retention_days` (default 30). Limits and defaults: [bounded admission and retention](concurrent-writes.md#bounded-admission-and-retention). **Dream keeps skipping one transcript, or doctor warns `dream_paid_loop`?** A dream key died `dream.breaker.max_dead_submissions` times (default 3) in 24 hours and is refused so it stops billing you. Fix the cause, then `gbrain dream reset-key --list` and `gbrain dream reset-key ''`. See [the paid-loop breaker](../operations/spend-controls.md#dream-paid-loop-breaker-dreambreakermax_dead_submissions). **Hourly cron sync keeps timing out on a federated brain?** Switch your cron to a per-source loop with shell `timeout(1)` doing the OS-level kill and gbrain self-terminating gracefully half-a-minute earlier: ```bash gbrain sync --break-lock --all --max-age 1800 for src in $(gbrain sources list --json | jq -r '.[].id'); do timeout 600 gbrain sync --source "$src" --timeout 540 || true done ``` When `--timeout` fires mid-import, `gbrain sync` exits 0 with status `partial` and `last_commit` UNCHANGED — the next run re-walks the same diff and `content_hash` short-circuits already-imported files. The `--max-age 1800` first command self-heals any wedged-but-alive locks left by a hung previous run, keyed on the lock's last refresh time (NOT when it was acquired) so healthy long-running holders are safe by construction. Scope note: the extract + embed phases still run to completion once started; `--timeout` interrupts the import walk only. **Dream cycle silently losing wiki links on Supabase?** The engine self-retries every bulk batch write (`addLinksBatch` / `addTimelineEntriesBatch` / `upsertChunks`) on Supavisor pooler blips, with a 12s worst-case wait that covers the full 5-10s circuit-breaker recovery window. `gbrain doctor` surfaces incidents via the `batch_retry_health` check (reads the last 24h of `~/.gbrain/audit/batch-retry-YYYY-Www.jsonl`). To tune for an unusually slow pooler: ```bash # Defaults: 3 retries, base 1s, max 10s, decorrelated jitter. # Override per operator without a release: export GBRAIN_BULK_MAX_RETRIES=5 # int >= 0; 0 disables retries export GBRAIN_BULK_RETRY_BASE_MS=2000 # int > 0 export GBRAIN_BULK_RETRY_MAX_MS=15000 # int >= base ``` Bad values surface at `gbrain doctor` startup with a paste-ready fix (not at first-retry mid-cycle). PGLite-only installs pay zero cost — the retry wrap is engine-level, but PGLite has no pooler so retries never fire in practice. **Dream cycle losing ~150 link rows per run with `'No database connection: connect() has not been called'` errors in the log?** The retry layer self-heals on a nulled-out database singleton: a `reconnect` callback on `withRetry` rebuilds the connection between attempts, and `PostgresEngine.batchRetry` injects `() => this.reconnect()` so engine-level batch writes survive a mid-cycle disconnect by something else in the same process. `gbrain capture` does not trail a `'No database connection'` stderr line from a background facts:absorb worker firing after CLI exit, because op dispatch awaits `getFactsQueue().drainPending({timeout: 1000})` before `engine.disconnect()`. To find which code path is still calling disconnect mid-process, run `gbrain doctor --json | jq '.checks[] | select(.id=="batch_retry_health")'`; the check surfaces the 24h disconnect-call count and the most-recent caller frame from the `~/.gbrain/audit/db-disconnect-YYYY-Www.jsonl` audit. **`gbrain brainstorm` returning `judge_failed: true` with 0 scored ideas?** You are on an outdated build. Upgrade with `gbrain upgrade` (it also runs any pending migrations; no config change is needed), then confirm with `gbrain --version` and `gbrain doctor --json`. Current builds size the judge's output cap to the idea count instead of truncating mid-JSON past ~40 ideas, and slash-form model ids (`gbrain brainstorm --judge-model anthropic/claude-sonnet-4-6 --max-cost 5`) resolve pricing the same as the colon form instead of failing with `BudgetExhausted reason=no_pricing`. **`gbrain reindex --markdown` wiped your auto/dream/signal-detector tags?** Run `gbrain upgrade`. Tag reconciliation is add-only: re-import and `reindex --markdown` ADD current frontmatter tags and never delete, so enrichment tags written to the DB (auto-tag, dream synthesize, signal-detector) survive a re-chunk. The reindex DB-only fallback also reconstructs the full markdown (frontmatter + body + timeline) before re-chunking, so a page with no on-disk source keeps its frontmatter, title, and timeline instead of getting overwritten with empty frontmatter. Trade-off: removing a tag from a page's frontmatter does not remove it from the DB on the next sync (frontmatter-tag removal needs a provenance column, deferred). **`gbrain sync` wedges on a large brain (no progress, high CPU)?** Three tools. First, name the stalling file: ```bash GBRAIN_SYNC_TRACE=1 gbrain sync --no-pull --no-embed --yes ``` The last `[sync] begin import: ` line with no following completion is the file being processed when the hang hit. Second, if you suspect a schema-pack `inference.regex` with catastrophic backtracking, complete the sync with the pack disabled and re-run extraction later: ```bash gbrain sync --no-schema-pack --no-pull --no-embed --yes ``` `gbrain schema lint` warns on the classic nested-quantifier ReDoS shapes (`(a+)+`, `(a*)*`, …) in pack regexes, and the runtime caps inference-regex input length (override via `GBRAIN_MAX_REGEX_INPUT_CHARS`). Third, on a PGLite brain with a live `gbrain serve` (your agent's MCP server), `gbrain sync` delegates through authenticated local IPC to the owner, whether it serves HTTP or stdio. If the client exits, accepted page requests can finish; repeat the same options to resume the managed sync cursor. Embeds defer to the owner's background work. See [`docs/architecture/serve-sync-concurrency.md`](../architecture/serve-sync-concurrency.md) for supported flags, managed-mode limits and the full triage. **`gbrain init --migrate-only` / a schema migration fails on Windows with `getaddrinfo ENOTFOUND`?** Run `gbrain upgrade`. Schema bring-up runs its phases in-process rather than spawning a child `gbrain init --migrate-only` per phase; a spawned child is what dies on Windows + bun + Supabase pooler with a DNS-resolution failure even though the parent connects fine, and running in-process removes the spawn entirely. The grandfather migration runs as a chunked bulk SQL pass (keyed on the page PK, soft-delete-filtered, source-safe) and completes in seconds on an 80K-page PGLite brain. ## Hybrid search returns only keyword hits **Symptom.** On a large Postgres brain, `gbrain query` answers look keyword-only, search metadata carries `vector_candidates_incomplete`, or each query takes about 8 seconds. The vector arm ran out of its 8 s candidate budget, usually because the planner chose a sequential scan over the HNSW index (`idx_chunks_embedding`). **Say to your agent:** *"Check whether vector search is using its index"* — the agent runs `gbrain doctor` and reads the `vector_plan` check. `gbrain doctor` reports `vector_plan` (Postgres only): - `ok`: the statement vector search sends uses the HNSW index. - skipped: PGLite, a column wider than pgvector's HNSW cap (exact scan by design), or fewer than 10,000 embedded chunks (a sequential scan is right for a small brain). - warn, index unused: the message names the plan the planner chose and whether the HNSW index exists and is valid. Fix in this order: upgrade gbrain on the brain host (`gbrain upgrade`) and rerun `gbrain doctor`; if the index is missing or INVALID, run the `CREATE INDEX CONCURRENTLY` / `REINDEX INDEX CONCURRENTLY` command doctor prints. - warn, stale text above 5%: run `gbrain embed --stale` so chunks edited after embedding get fresh vectors. - warn, legacy guard: see below. **Legacy guard (temporary rollback).** Vector search checks content freshness outside the HNSW candidate scan. `search.vector_legacy_guard` restores the older statement, which checked freshness inside the scan. Use it only when vector search is slower or returns different results than before your upgrade, and report the regression: 1. The setting belongs to the process that runs searches on the brain host (`gbrain serve`, autopilot, job workers), never to a thin client. 2. Set it with `gbrain config set search.vector_legacy_guard true`, or export `GBRAIN_VECTOR_LEGACY_GUARD=1` in that service's environment (the variable wins over the config key). 3. Restart the owning service. It reads the setting once at its first search and prints `[gbrain] vector legacy guard active` to stderr. 4. Confirm with `gbrain doctor`: `vector_plan` warns "legacy guard configured … active after restarting the owning service". 5. Remove it once the regression is fixed: `gbrain config set search.vector_legacy_guard false` (or unset the variable) and restart again. **Filtered vector results changed (#6132).** HNSW scans use `relaxed_order`; restore with `gbrain config set search.hnsw_iterative_scan strict_order`, then restart serve. The guard is temporary. The release that retires it prints a one-time notice when the inert setting is still present. ## Global maintenance timeouts **Doctor warns `global_maintenance_timeouts`, or late maintenance phases (orphans, purge, the brain-wide embed) never seem to run?** On a large brain one `autopilot-global-maintenance` job may not fit every phase before its deadline (30 minutes by default). Each job stops starting phases that its deadline would cut off and the next job resumes at that phase, so one pass can span several jobs; the resume point is the config row `autopilot.global_maintenance.progress`. A phase that was running when a job died is skipped for the rest of that pass so the others still run, and doctor warns once it has killed three jobs in a row (or the last three jobs all died at the deadline). 1. Run the named phase in the foreground, without the job deadline: `gbrain dream --phase ` (for example `gbrain dream --phase embed`). 2. Give the job more time if your brain needs it: `gbrain config set autopilot.global_maintenance_timeout_ms 3600000`, or set `GBRAIN_GLOBAL_MAINTENANCE_TIMEOUT_MS` in the autopilot service environment (the variable wins over the config key; values below 60000 are ignored). Unset both to return to the default. 3. Confirm with `gbrain doctor`: `global_maintenance_timeouts` is `ok` after the next job finishes within its deadline. Doctor also warns when the last three dream `patterns` or `synthesize` subagent children died at their own subagent timeout, even though each maintenance job completed (the phase records a partial result). Raise that child's timeout with `gbrain config set dream..subagent_timeout_ms `; `gbrain dream --phase ` reruns it and spends model tokens, so ask the user first. The warning clears once that phase's next child completes. ## auto_chronicle has no effect **Say to your agent:** *"Why aren't my meetings showing up as timeline events?"* Automatic event extraction is on by default. See the [Life Chronicle guide](life-chronicle.md) for what qualifies, the cost, the three-step check, and the skip codes. `gbrain doctor` reports it as the `auto_chronicle` check. To turn it off, run `gbrain config set auto_chronicle false`.