# Config Module ## Overview The config package loads runtime configuration from `~/.kiro/crew/config.json` using stdlib dataclasses with sensible defaults. `config/sections.py` and `config/loader.py` are the two facades callers import. Each composes the owner modules below and re-exports their names as the same objects, so an existing `config.loader.X` or `config.sections.X` import keeps resolving. A patch reaches the code that looks the name up in the patched module, and only that code. Names read by code that stays in `loader.py` are call-time seams on the loader: `config_path`, `config_dir`, `config_local_path`, `env_path`, `workspace_root`, `_default_workspace_base`, `write_config_atomically`, `update_config_locked`, `atomic_write`, `_config_write_lock`, `_config_fingerprint`, `_validate_config_data`, `_persist_config_migration`, `_apply_document_migrations`, `_log_config_clamp_event`, `_DEFAULT_CHAT_TURN_TIMEOUT_SECS`, `DEFAULT_POOL_SIZE`, `unsandboxed_exec_declared`, `publish_config_timezone`, `record_adoptions` (the loader passes it to the migration transform at call time), each `_build_*` name as `build_config` calls it, and the published-snapshot globals. A helper or constant read INSIDE a relocated builder or migration rule is patched on the module that reads it: `config.section_builders` for the value coercers, the STT, computer-use and instance bounds, `coerce_runtime_ceiling` (which reads the monitoring bounds in `monitoring.limits`), the `_resolve_stub_*` roster readers and the section DTO classes a builder constructs (a replacement must be a dataclass, and its field defaults are what an omitted key reads); `config.migration` for `auto_adoptable`, `drop_drifted_keys`, `stored_value_or_none`, `superseded_default_drift` and `drift_summary`. The loader facade and `config.migration` hold the same warn-once set `_REPORTED_SUPERSEDED_KEYS`: clear it with `.clear()`, never rebind it. `test_config_refactor_contract.py` pins representative seams of each kind. | Owner | Owns | |---|---| | `config/fields.py` | `_meta` field metadata, the `_safe_*` value coercers every section shares, and `field_default`, the one reader of a DTO field's declared default (`SectionReader` reads through it). An integer coercer keeps a legacy numeric string or integral float already in a config file (`"5"` -> 5, `5.0` -> 5) and rejects booleans and non-integral, malformed or non-finite values in favour of the default. A leaf: it imports nothing from `kiro_crew`. | | `config/sections.py` | The DTOs other specs and tests anchor here: agent, crew record, workspace, session, dashboard (with `TailscaleConfig` and its parser), the messaging channels, `wakatime`, speech-to-text and its degradation rules, telemetry, decisions, resource limits, and the bounds constants. It is also the facade for the three section owners below. It also holds `SectionReader`, which resolves a builder's omitted key and coercer fallback to the DTO field's declared default. | | `config/memory_sections.py` | `memory`, `knowledge`, `skills`, `session_summary` and the named `memory_stores` records. | | `config/integration_sections.py` | `mcp`, `mcp_gateway` (with the MCP stub roster readers the gateway seed shares), `instances`, `tunnel`, `publish`, `computer_use` and the external app `registries`. | | `config/service_sections.py` | `taskrunner`, `messaging`, `cron_history`, `monitoring`, `heartbeat` and `watchdog`. | | `config/section_builders.py` | The `_build_*` helper of 27 sections, grouped by the module that owns each section's DTO. Each reads its section through `sections.SectionReader`, so an omitted key builds the DTO field's default and a `read()` coercer falls back to it; the values a builder keeps of its own and the deliberate departures are listed under "Defaults come from the DTO fields". Four `_build_*` helpers stay in the loader (agent, session, telemetry, dashboard). Sections with no helper are built inline in the loader's `build_config` (`heartbeat`, the external app `registries`, `memory_stores`, the `agents` crew roster, `workspaces`) or by their DTO (`DecisionsConfig.from_raw`, `ResourceLimitsConfig.from_raw`, `ChannelConfig.from_dict` for `slack_channels`). | | `config/migration.py` | The write-back migration ids, the document transform `apply_document_migrations`, the one-shot `connections_ui` marker name, the legacy `skills.lazy_load` cohort test, superseded-default reporting, and the in-memory half of an adoption. | | `config/resolution.py` | Raw overlay merging, top-level section classification, and degraded-input tracking. | | `config/validation.py`, `config/schema.py` | Schema validation with the validated-data cache, and the JSON schema and restart registry built from the DTOs. | | `config/paths.py`, `config/live.py`, `config/superseded_defaults.py` | Pure path primitives, the one live-config watcher and applier registry, and the superseded-default registry with its acknowledgment ledger. | | `config/loader.py` | `KiroCrewConfig` (load, serialize, save, model and provider resolution) and the residual core below. | Imports run one way: `fields`, then the three section owners, then `sections`, then `section_builders` and `migration`, then `loader`. Each module imports only modules earlier in that order plus the existing leaves (`sections` imports `resolution`, `migration` imports `superseded_defaults`, and neither leaf imports back), and none of the owners imports `loader`, `schema` or `validation`; `test_config_module_boundaries.py` pins the graph. Every module split out of the loader logs as `kiro_crew.config.loader`, so a relocated warning keeps the record name operators filter on. `sections.py` keeps its DTOs because other specs name the file for them: [messaging](messaging.md) (the channel restart flags), [slack-gateway](slack-gateway.md), [crew-mode](crew-mode.md), [history](history.md), [stt-streaming](stt-streaming.md), [metrics](metrics.md), [decisions](decisions.md), [security](security.md), [model-selection](../common/model-selection.md), [model-fallback](model-fallback.md), [subagent](subagent.md), [acp-client](acp-client.md) and [crew-log-projection](crew-log-projection.md). Four source scans also allow a construct only in that file: `ResourceLimitsConfig.from_raw`, `_tailscale_config_from`, the `_AVATAR_MOTIONS` literal and the `AgentConfig.acp_backend` declaration. `loader.py` groups each residual responsibility into one bannered section. Each stays in that file because something outside this package names the file, or because its readers look up a name the loader's callers and tests patch there: | Kept in `loader.py` | Held there by | |---|---| | Credential keys, the `.env` reader, the dashboard port | [code-style](../common/code-style.md) names `config/loader.py` for `CRED_*` and `_DEFAULT_PORT`. | | Data-home and workspace path helpers | Tests and callers patch `config_path`, `config_dir`, `env_path`, `workspace_root` and `_default_workspace_base` on this module. The instance-pairing and redactor-registry scans key `read_local_secret` and `credential_redaction_path` to this file. | | The unsandboxed-exec platform policy | [security](security.md) places that resolution in the loader, where the raw document is read. `unsandboxed_exec_declared` reads the patched `config_path`/`config_local_path` and is itself patched on this module. | | Document I/O: `_raw_config`, `read_config_for_update`, `write_config_atomically`, `update_config_locked`, the meta stamp | The config-writer scans in `test_config_rmw_preserves_settings.py` exempt only `loader.py`, and the writers read this module's patched path and `atomic_write` names. | | Write-back persistence (`_persist_config_migration`, the backup) and the `_apply_document_migrations` seam | It rewrites `config.json` under the same writer exemption, and passes this module's `record_adoptions` to the transform as the adoption-ledger writer. | | The validated-document cache fingerprint, its overlay sidecar and invalidation | `_config_fingerprint` reads the patched `config_path`/`config_local_path` and is itself patched on this module; `save()` and the write-back call `_invalidate_config_cache` beside it. | | `KiroCrewConfig.load`, the load pipeline (`read_config_document`, `build_config`, `persist_write_back`) and `save` | They read `config_path`, `config_local_path`, `_config_fingerprint`, `_validate_config_data`, `_persist_config_migration` and `write_config_atomically` by name, all patched on this module. `test_config_section_construction.py` pins `build_config`'s assembly shape. | | The security clamp and its SEL event | [security](security.md), [sel](sel.md) and [resource-protection](../../architecture/resource-protection.md) name the loader. | | The loop-stall and managed-launch readers | `load_loop_stall_exit_after` reads the loader's patchable `KiroCrewConfig`; `resolve_loop_stall_exit_after` and `consume_managed_service_launch_environment` are its two halves, and the dashboard server imports all three from here. The consume takes the marker out of `os.environ` so descendants do not inherit it, and hands it to `platform_compat.keep_for_reexec`, so the gateway's own exec successor (an in-app restart) is still a managed launch; `launched_as_managed_service` answers from either place, so it does not change once the dashboard has started. | | The agent, session, telemetry and dashboard builders | The harness-parity review scope; the patchable `DEFAULT_POOL_SIZE` fallback; [metrics](metrics.md) naming the loader as the telemetry parser; the feature map naming the loader's `folder_sort` read. | | Published snapshots: materialized agents, the alias table, the compaction threshold, the timezone | Tests rebind this module's snapshot state, and the second-boot witness in `test/integration/test_boot_smoke.py` keys the counters to `kiro_crew.config.loader`. | | Agent resolution and the provider factory | [crew-mode](crew-mode.md) and [context-management](../../architecture/context-management.md) name the loader for `resolve_agent_bindings` and `resolve_effective_model`. The agent-spec read inventory keys its call sites to this file, the ACP import is a baselined agent-SDK edge, and the blocking harness-parity and memory-store review rules cover `config/loader.py`. | New section constants, including local speech's automatic-language default, are read from `config.sections` directly; they do not expand that historical facade. New work goes to the owner of its responsibility, not to a facade: a field to the module that owns its section's DTO, a new section's DTO to `memory_sections.py`, `integration_sections.py` or `service_sections.py` by domain, a shared coercer to `config/fields.py`, a section's `_build_*` helper to `config/section_builders.py`, a migration rule to `config/migration.py`, and overlay, validation, schema, live-applier and path work to its owner in the table above. `config/loader.py` takes only work inside one of its residual rows, and `sections.py` gains a new DTO only when another spec anchors it there. A feature whose section spends tokens on the user's behalf defaults to off and documents its knobs in its own spec — `session_summary` is the current example (see [session-summary.md](session-summary.md)), following the shape `SkillsConfig` established: every field carries `_meta` label/help for the config surfaces, out-of-range values are clamped with a warning rather than raising, and a malformed section degrades to defaults so a hand-edited file cannot prevent the gateway from starting. `dashboard.dynamic_dashboard_cards` follows the same cost rule: default false, hot-applied through the live watcher, with a native control on Dynamic Dashboard surfaces. It enables event-driven per-session HTML cards without enabling `session_summary`. Runtime status and native decisions remain available when it is off. Route and watcher registration do not import or construct the optional producer. Both server entrypoints activate it in the deferred post-listen watcher task when enabled, or on the first live enable; later toggles reuse that instance and its charged budgets. Fixed call, byte, cache and iframe limits and the disable/cancellation contract are documented in [learn-cron-dashboard](learn-cron-dashboard.md#automatic-session-status-cards). ## Orchestration prompt contract `config/prompt.md` guides direct work and delegation using a concrete-value policy. Parent-plus-child parallelism depends on the spawn receipt's delivery capability. Runtime checks and compatibility are owned by [subagent.md](subagent.md), not inferred from prompt wording. ## Embedding rebuild request publication `memory.embed_rebuild_generation` is an explicit-apply request identity, not a model label or a completion counter. Model apply commits it with the validated model settings before invalidating vectors. Managed vector publications hold the same config sidecar lock while checking that request and committing their SQLite write. They cannot publish an old result after a newly committed request, even when the store has not yet been reconciled. Store-local signatures and handled requests remain checked inside SQLite write admission. Conditional model rollback preserves the request and unrelated edits, and refuses a competing model/request change. An untouched upgrade retains the empty request default. This field is install-local runtime obligation state even though it is published atomically in the model-settings transaction in `config.json`. It is not a portable preference or a completed-work flag. Back up and restore the model settings, this request, and memory databases together. Copying a non-empty request to a different installation intentionally invalidates stores that have not acknowledged that request, including aligned ones; do not distribute it in a fleet configuration template. Clearing/resetting it while keeping existing memory databases loses the outstanding same-label rebuild obligation. The ordinary model-space signature check remains, but cannot replace that lost request. Such a reset is not a supported way to cancel or complete rebuilding: after a reset or mismatched restore, explicitly apply the intended model again (`Rebuild memory vectors` for an unchanged path) before relying on vector search. That publishes a fresh obligation for open and later-opened stores. This design accepts config/state coupling to retain one atomic publication point; it does not claim to recover obligations that an operator deletes out of band. Managed vector publication resolves a symlinked config through `_lock_target`, exactly as the config writer does, before taking its sidecar lock. Within a store operation the order is config sidecar, the store's process-local lock, then SQLite write admission. A writer waiting for the config sidecar therefore holds no database lock that could stall a concurrent reader. Both locks remain held through vector validation, commit or rollback. Model apply releases its config mutation before aligning stores; it never holds that sidecar while waiting for a store lock. Native inference runs before those publication locks. Store close releases its SQLite and lifetime handles without saving a vector index or acquiring config admission again, including when rollback failure closes an uncertain connection. ## Data Home Location Kiro Crew's data root nests **under kiro-cli's own `~/.kiro/` base** so all Kiro-family apps share a single directory a user can secure. `config_dir()` (in `kiro_crew/config/paths.py`, re-exported from `kiro_crew/config/loader.py`) is the single accessor and resolves to: 1. `$KIROCREW_HOME` when set (expanded and resolved; refuses any filesystem or drive root and system directories such as `/usr`, `/System`, `/etc` and the macOS-resolved `/private/etc`), else 2. `~/.kiro/crew` (the default). **No migration — net-new users only.** All supported installs start directly in `~/.kiro/crew`; there is no `~/.kirocrew` to relocate, so `config_dir()` simply resolves and `mkdir`s the home above. The one-time `~/.kirocrew` → `~/.kiro/crew` data-home migration that earlier releases carried has been **removed** (see `docs/system-specs/post-launch-removals.md`). A leftover top-level `~/.kirocrew` from an old install is never read, migrated, or deleted; it is left in place — still credential-gated by the `.kirocrew` security-path spelling — and `kirocrew doctor` reports it, warning rather than advising deletion when it still holds a virtual environment (`venv`/`.venv`/`venvs`), since that may be the running interpreter. The same section warns, without failing, when the workspace root or `KIROCREW_PROJECT_DIR` is the data home or one of its parents (`$HOME` is the usual case): that puts the managed browser launcher inside the agent-writable tree, where the launcher fence refuses it, so the one-click browser install refuses up front (see the browser module). **Repository-controlled uninstall contract.** Every uninstall path owned by this repository preserves the Kiro Crew data home by default. `kirocrew service uninstall` removes only its service definition; the Python/npm packages define no uninstall lifecycle hook; and the desktop shell's generated NSIS uninstaller removes only installed program state: its install directory, shortcuts, channel-scoped updater cache, and any legacy “start with Windows” registry entry (`deleteAppDataOnUninstall` stays false), without resolving or removing the Kiro Crew home. App Kit uninstall also preserves the app's `data/` subtree unless the dedicated `purge_data=true` API action (CLI `--purge-data`, or an explicit dashboard choice) is supplied. The API checks for the literal boolean `true`; absent, legacy, or malformed values fail closed to preservation. A whole-home purge is never coupled to uninstall. **Uninstaller consideration (external dependency).** Because the data home now lives under `~/.kiro/`, a hypothetical Kiro-family uninstaller that removes `~/.kiro/` would also remove `~/.kiro/crew` and take Kiro Crew's data — config, credentials, memory DB, session history, and the SEL audit chain — with it. This is a persisted-data one-way door, and — unlike when an archived rollback copy existed — there is now no `~/.kirocrew.archived` fallback for ANY install (upgrader or fresh), so such a wipe is unrecoverable total data loss. Any Kiro-family uninstaller spec **MUST** either explicitly exclude `~/.kiro/crew` from a `~/.kiro/`-wide wipe, or prompt before deleting it. Independently, a user who wants the data home entirely outside `~/.kiro/` can set `KIROCREW_HOME` to relocate it. **Technical hedge — recovery-pointer breadcrumb.** `config_dir()` writes a small, non-secret `~/.kirocrew.breadcrumb` pointer file at the top-level home (`RECOVERY_BREADCRUMB_NAME`), deliberately **outside** `~/.kiro/`, recording the data-home path (see `_write_recovery_breadcrumb`). It is idempotent where the platform can check safely (on POSIX the prior content is read via `O_NOFOLLOW` and rewritten only when the recorded path changes; where that flag is missing — Windows — the check is skipped and the file is atomically rewritten once per process), best-effort (never blocks startup), and written only on the default path (a `KIROCREW_HOME` override carries no `~/.kiro/` wipe risk). It is **not a backup** — just a durable signpost that survives a `~/.kiro/`-wide uninstaller wipe so a user or support script can find any surviving data or understand what was removed. This narrows, but does not eliminate, the one-way-door risk above; the release gate still stands. > **Release gate (UNINSTALLER-EXCLUDE-CREW).** This is a pre-release, > human-sign-off dependency, NOT a code change in this repo: the code cannot > constrain another product's uninstaller. For every release that ships data > under `~/.kiro/`, the Kiro Crew product owner MUST confirm the > Kiro-family uninstaller either excludes `~/.kiro/crew` or prompts — because > there is no `~/.kirocrew.archived` fallback for any install, so a > `~/.kiro/`-wide wipe would be unrecoverable total data loss. Until confirmed, > the placement decision is acknowledged-but-owned here under this name so it is > not lost. **Tracked as release-blocking in > [issue #355](https://github.com/kirodotdev/KiroCrew/issues/355)**; the sign-off > must be recorded there. **Paths are resolved per call, never captured at import.** Because `config_dir()` re-reads `$KIROCREW_HOME` on every call and resolution/maintenance is deliberately lazy, the resolved value is only correct at the moment it is needed. Modules therefore MUST NOT bind a path factory result to a module-level constant: ```python _SOME_DIR = config_dir() / "some" # WRONG -- frozen at import ``` An import-time binding captures whatever home was active when the module was first imported, which breaks two things at once: pod isolation (a pod exports its own `KIROCREW_HOME`) and test isolation — `conftest.py`'s autouse `_isolate_kirocrew_home` fixture runs *after* collection has already imported the module under test, so it cannot reach a frozen constant, so a test run could write fixture rows into an operator's real usage store. The required shape keeps the module-level name as an explicit opt-in override (`None` = resolve live), so existing `monkeypatch.setattr(mod, "_SOME_DIR", tmp)` call sites keep working: ```python _SOME_DIR: Path | None = None def _some_dir() -> Path: return _SOME_DIR if _SOME_DIR is not None else config_dir() / "some" ``` Annotating the override as `Path | None` is load-bearing: any consumer that still reads the constant directly becomes a **mypy error** rather than a silent `None` at runtime. This is enforced repo-wide by `test/test_lazy_data_home_paths.py`, which walks the AST of `src/kiro_crew` for module-level assignments calling any factory declared in `config/paths.py` and fails on every hit. The factory list is derived from `paths.py` itself, so a newly added factory is covered without editing the test. **`config_dir()` maintains; `data_home()` only resolves.** `config_dir()` is *resolve + maintain*: besides resolving the home it `mkdir`s it and refreshes the recovery breadcrumb (a stat + a read). That work belongs to process start — `ensure_data_home()` is the startup hook — and the distinction did not matter while callers froze the result in a module constant, because the maintenance then ran exactly once, at import. Resolving per call makes it load-bearing: a request handler would otherwise refresh the breadcrumb **on the event loop** as a side effect of asking where a directory is. So the accessors above call **`data_home()`**: | branch | behaviour | | --- | --- | | a **valid** `KIROCREW_HOME` override | delegates to `config_dir()` every call, so an override set *after* import is honoured. That branch performs no breadcrumb refresh — only a cheap `mkdir`. | | default home already resolved | returns the cached `_resolved_home` directly — no `mkdir`, no breadcrumb. | | not yet resolved | delegates to `config_dir()`, so the **first** resolution in a process creates the home and refreshes the breadcrumb once. | The first row tests `_valid_override_home()` — the **same predicate `config_dir()` gates on**, not merely "is the env var set". An override naming a system directory (`/`, `/usr`, …) is rejected there and resolution falls through to the default home, so gating on the raw env var would send every call down the maintenance path for anyone with a bad override. The two predicates must not drift apart; a regression test pins both directions. `data_home()` keeps no cache of its own — the override branch must stay live, and the cached branch reads the same `_resolved_home` that `config_dir()` populates, so there is one source of truth for the location. Direct `config_dir()` callers keep the maintenance behaviour; a hot or async path should call `data_home()` instead. ## Workspace Root `workspace_root()` returns the base directory for all LLM working directories (kiro-cli cwd, task runner output, etc.): Resolution order: 1. `KIROCREW_WORKSPACE` env var — no `kirocrew-workspace` subdirectory appended 2. Saved path in `~/.kiro/crew/workspace_dir` (written by `kirocrew setup`; re-running setup preserves the existing value as the prompt default) 3. Platform default: | Platform | Path | |----------|------| | macOS | `/Volumes/workplace/kirocrew-workspace` (falls back to `~/workplace/kirocrew-workspace` if `/Volumes/workplace` doesn't exist) | | Linux | `~/workplace/kirocrew-workspace` | | Windows | `~/workplace/kirocrew-workspace` | Values from (1) and (2) lose one surrounding quote pair and have `~` expanded; a non-absolute result is logged and replaced by (3). Each session/task gets an isolated subdirectory under this root via `_session_work_dir(key)`: - Chat sessions: `kirocrew-workspace/cli_chat`, `kirocrew-workspace/{thread_ts}` - Background: `kirocrew-workspace/_bg` - Cron: `kirocrew-workspace/cron_{job_id}` - TaskRunner: `kirocrew-workspace/taskrunner_main` - One-run directories for a subagent and a stateless cron run: `kirocrew-workspace/subagent_` and `kirocrew-workspace/cron__` (`session_work_dir._DERIVED_SESSION_SHAPES`). These carry a provenance marker and are reclaimed when only residue is left, at shutdown and by the predecessor sweep; the contract is in [acp-client](acp-client.md). `workspace_root()` creates the selected root on first call and realpath-normalizes it before session subdirectories are derived; `workspace_root(create=False)` resolves it without creating it (the doctor's census uses that). ## Project Directory Resolution `KIROCREW_PROJECT_DIR` env var controls where agent config and skills are loaded from: 1. Env var `KIROCREW_PROJECT_DIR` (if set and valid) 2. CWD walk-up — CLI walks up from CWD looking for `skills/` + `src/kiro_crew/`. Agent config ships in `src/kiro_crew/config/`; when a user creates `/agents/defaults.json` or `/agents/prompt.md`, `agent._shipped_defaults` / `_shipped_prompt` prefer those over the bundled copies 3. Saved path in `~/.kiro/crew/project_dir` (written by `kirocrew setup`) 4. Bundled fallback — `config/defaults.json` and `builtin_skills/` inside the package The CLI (`cli.py:main()`) auto-detects and sets the env var at startup. ## Folder steering directories (`dashboard/chat_folders.py`) A chat folder record in `folders.json` may carry an optional `steering_dirs` list of absolute (or `~`-prefixed) directories, alongside its existing `project_dir`/`default_agent`/`color`/`tags`. Every chat in the folder's subtree loads each directory's `**/*.md` files whose inclusion is `always` or absent as steering (`manual`, `auto` and `fileMatch` files are skipped), in addition to global and project steering. The key is omitted when empty (like `color`/`tags`), so "absent means none" is the single on-disk representation and a PATCH with `[]` clears it. - **Validation** (`dashboard/chat_folders.py` `_validate_steering_dirs`) checks each entry: at most `MAX_FOLDER_STEERING_DIR_LEN` (4096) characters before and after canonicalization; a UNC path refused before resolution; refused outright on a platform without a pinned tree walk (`pinned_fs.supports_pinned_tree_walk`); canonicalized and screened by `hooks.validate_file_path` (a refusal is SEL-logged); refused when the memory-silo fence cannot be computed or the path crosses a memory silo; and proven to exist by a pinned, no-follow open rather than `isdir`. On top of that come a `MAX_FOLDER_STEERING_DIRS` (16) list cap and rejection of duplicates within one folder (compared by resolved path). The fence, platform and inclusion rules are specified in [providers](providers.md). Failures return `400` with code `steering_dirs_invalid`. - **Principal gate** (`_refuse_principal_steering_dirs`) runs at both write sites BEFORE validation touches any path: only the person may declare a non-empty list. A `POST`/`PATCH` carrying one from an app or crew-member principal is refused `403` with code `steering_dirs_forbidden` and SEL-logged as denied, because folder permission is not host-file permission -- the gateway reads these files unsandboxed on the folder's behalf, and an app that could point its own folder at an arbitrary readable Markdown tree would have that read laundered into its own model session with no tool grant and no signal. Clearing to `[]` stays allowed for every principal (it only removes reads). The person may still declare steering on a folder an app or member owns; delivery then routes it to that principal's chats as described below. - **Resolution** (`_resolve_folder_steering_dirs`) is ACCUMULATIVE up the `parent_id` chain (root ancestor first, then descendants), unlike the nearest-wins `project_dir` resolver: an org-standards folder above a per-repo folder contributes both sets. The walk is cycle-guarded, re-validates each stored path (never trusting `folders.json`, which can list a directory since moved or made sensitive), and dedups by resolved realpath keeping the first occurrence. - **Live resolution, no slot field.** The effective list is never cached on the chat slot: `_ChatSlot` carries no steering field, and neither slot create nor agent switch resolves one. The dashboard turn path (`dashboard/chat_runner.py`) resolves it live from the slot's current `folder_id` against the committed folder tree, on every fresh session and every post-compaction reinjection turn; a V2 member chat resolves it on every turn, because its essentials envelope is rebuilt each turn (`dashboard/chat_turn/turn_context.py` `_folder_steering_turn` is the gate). So editing a folder's `steering_dirs` reaches a template chat already filed in it at its next fresh session and a member chat on its next turn, with no cache to invalidate; a slot with no `folder_id` contributes nothing. A resolver error logs a warning naming the slot and the turn proceeds with no folder steering, and a directory that has since disappeared is skipped at read time rather than failing the turn. Delivery is performed once by the context builder (`kiro_crew.context.ContextBuilder`) reading through `kiro_crew.folder_steering` — the one layer every backend passes through, so no provider can silently drop it. See [providers](providers.md). ## Named Memory Stores (`memory_stores.py`) The reserved `agents.default` assistant uses the existing Global Memory **V1**. Explicit creation of a new Crew Member allocates one **V2** memory store. Automatic discovery and existing V1 members retain their V1 bindings. Changing `default_agent` selects the member with its existing memory version and binding; it never converts member memory to Global. A materialized provider template which is not a Crew Member continues to use V1. ### Separate files preserve V1 | Resolver | Global `default` | Member V2 store | |---|---|---| | `memory_store_dir_for` | `/workspace/` | `/memory_stores//` | | `resolve_store_path` | `/memory.db` | `/memory_stores//memory.db` | | `memory_index_path_for` | `/memory_index.db` | `/memory_stores//memory.db` | V1 retains its markdown, history and separate index layout. Explicit V2 creation adds only manual `memory/preferences.md` and `memory/projects.md` beside its single SQLite database; learned history and search use that database. No path is renamed and no V1 data is migrated, copied or algorithmically converted on member creation. Member stores begin empty. Explicit selected-content inheritance and its provenance are owned by [memory-skills-hooks](memory-skills-hooks.md). ### Member identity and creation A V2 member has an immutable persisted `member_id`, independent of its editable config/display label. `MemoryStoreConfig.owner_member_id` and the database's `member_database` singleton record identify the same owner and store ID. `memory_version: 2` selects V2; legacy declarations default to V1. `owner_member` is descriptive display metadata, never execution authority. Templates and projects cannot select a member's memory. The template-picker roster withholds `member_id`; the full configuration view applies the same credential and exfiltration redaction to this hand-editable string as other member fields. Stored identity values remain unchanged. `provision_member_memory(config, member)` allocates an exclusive random directory and creates a new SQLite database through `create_member_database`, plus the explicit manual preferences/projects documents. The member/config binding is published only after initialization succeeds. It never adopts an existing unidentified database, resets a store, or copies earlier V1/V2 data. An existing V2 member must retain its ID and store. No manifest, retirement registry, private payload copy, OS admission gate, or memory provisioning feature flag is involved. `POST /api/agents` and `kirocrew agent create` allocate new stores automatically. A supplied named store is rejected. Owner edits may echo the current store but cannot silently rebind it. Legacy members keep their explicit Global or named V1 binding. Member updates and automatic discovery never initialize a V2 database; only explicit creation of a new member may allocate one. Existing V1 memory and transcripts remain unchanged. `persist_member_config` publishes the member and store under `update_config_locked`, rechecking duplicate creation, immutable member ID, expected binding, and exclusive store ownership while preserving unrelated concurrent edits. When creation fails or is cancelled after allocating a store, the creation paths call `retire_unpublished_allocation`: it re-reads `config.json` under the same lock and removes the fresh directory only when no member and no store entry references it, then restores the in-memory binding so a retry allocates fresh. A publication that landed is kept; pre-existing bindings and stores are never deleted or retired. A create also purges the member name from the crew-teams record inside the same locked mutation (`crew_teams.release_for_create`); a purge that cannot be made (`TeamsUnavailable`) refuses the creation, 409 `teams_unavailable` on the dashboard and exit 1 from the CLI. Degraded or unreadable configuration refuses creation instead of guessing defaults. Non-string `member_id` or `owner_member_id` values refuse allocation with `UnknownMemoryStore` before forming the reserved-ID set, preserving the raw values, existing databases and V1 bindings instead of replacing damaged identities. A `memory_version: 2` record whose `owner_member_id` is empty is a member store written before identities existed; the loader keeps it verbatim and every resolver refuses it. `repair_legacy_member_stores()` runs in the `cli.main` prologue after `boot_platform` for every CLI subcommand except `gateway`, `doctor` and the `mcp-*` servers, and for the gateway in its post-readiness memory preparation worker (never on the boot path), so both surfaces publish the missing `member_id` / `owner_member_id` pair before any member is resolved. It writes through `update_config_locked`, skips a config whose memory section degraded, and never raises. Criteria, refusals and the database half are owned by [memory-skills-hooks](memory-skills-hooks.md#pre-identity-member-stores-are-upgraded-at-start). ### Exact resolution and explicit failures Admission resolves one frozen `ExecutionContext` containing member ID, store ID, selection namespace, template, app attribution, and privacy mode. The snapshot is part of the owning session/run/job record, captured before awaiting asynchronous work. Background workers, continuations, schedules and reruns inherit it directly; closing the originating chat or editing a display name does not retarget execution. An explicit `target_member` selects an existing configured member under ordinary execution permissions. There is no automatic provisioning or fallback to Global. `resolve_declared_store` and `require_member_memory_store` reject malformed, missing and contradictory declarations. `resolve_agent_bindings(..., validate_memory_files=False)` captures configuration identity without opening learned memory. Persona/project/manual context can therefore load while the learned database is unavailable. Actual memory operations validate the captured store's database identity with `open_member_database`; an absent, corrupt or wrong member database is an explicit error and is never recreated at read time. Store names use lowercase letters, digits and hyphens, 1–80 characters, with no leading/trailing hyphen, path separators or Windows device basename. Managed paths must remain under `memory_stores/` and may not redirect to another store. `memory_store_version` reads the exact configured declaration; it does not infer ownership from labels, paths or database contents. V1 opens refuse a database bearing the explicit V2 identity table rather than treating it as Global memory. Incognito and Temporary modes are inherited monotonically. New restricted sessions keep their canonical record in memory and suppress Crew body/checkpoint writes. Tightening an existing persisted execution updates only retention metadata so a restart cannot broaden it. Ordinary transport authentication, app/owner permissions, enterprise governance, host sandbox and credential protection remain independent of memory routing. Same-host arbitrary-code confidentiality between member stores is not a goal. See [memory-skills-hooks](memory-skills-hooks.md) and [security](security.md) for the storage and ordinary host-security contracts. ## Workspace fall-through is logged `workspace_dir_for(name)` also reads the LOADED config's `workspaces` table rather than the raw bytes, so it and `resolve_agent_bindings` cannot answer the same question two ways. An unmapped name falls back to `default_workspace` and then to `WorkspaceConfig().dir`, and the fall-through is logged: two DISTINCT names both resolving to `/workspace` warns, because that is how a workspace split becomes a shared tree nobody notices. An install that simply has no `workspaces` section is the ordinary fresh state and logs at debug. It **never raises** — `default_project_dir` and the workspace-identity block are built on it, so a raise would break a fresh install and take both with it. A legacy FLAT `{"name": "dir"}` workspaces entry is a schema type mismatch that the validator keeps (`_is_flat_workspace_string`): `_migrate_workspaces` turns it into `WorkspaceConfig(dir=...)` in the loaded table, and the `MIGRATE_WORKSPACES` write-back persists it as `{"dir": ...}`. Both `resolve_agent_bindings` and `workspace_dir_for` answer from that table. ## Superseded Defaults (reported; a named few adopt themselves once) `config.json` is a full materialization of the schema -- every field is written to disk, including fields the operator never set -- and each field is resolved as `data.get(key, DEFAULT)`. A stored value therefore always beats the dataclass default, so **changing a shipped default reaches only installs created after the change**; a pre-existing install keeps whatever value was materialized last. `config/superseded_defaults.py` holds an append-only registry (`SUPERSEDED_DEFAULTS`) of default changes existing installs should be told about, each entry naming the dotted key, the old default, the new default, and the release that changed it. `superseded_default_drift(base_data)` returns the entries whose stored value equals the old default, comparing type as well as value so a stored `0` is not read as `False`. Every registered row is listed once, in the table under *Auto-adoption* below, with its change and the reason it adopts or only reports. **Three** carry `auto_adopt` -- the agent timeout budgets and the spawn memory floor -- and every other row is report-only. `test_the_spec_table_lists_every_registered_row` keeps that table equal to the registry. A row may also carry a `note`: one sentence that `doctor`, `kirocrew config defaults` and the `--keep` confirmation append to it. It exists when choosing between `--adopt` and `--keep` needs a fact the old and new values do not carry. `skills.lazy_load`'s says what `false` means now (below). `decisions.history_budget_chars` says adopting is bounded by the consented history ceiling. The two tool-stall windows say they are coupled (see the table). A row marked `meaning_moved` also puts its note on the one-line load-path warning. That mark is for a stored value whose MEANING moved along with the default, and `skills.lazy_load` is the one such row. On 0.6.0 and earlier `false` was the default full skills dump, and today it selects the shorter entry naming only the eight hottest skills. A 0.6.0 install may hold a materialized `false`, and an operator who chose `false` there chose the full dump. The row stays report-only because `false` is also the supported switch to the short entry. Its note says what `false` means now, on every surface, so keeping it is never mistaken for keeping the 0.6.0 behaviour, and `--adopt` takes the index. A row may also carry `applies`, a predicate that says where the old and new defaults behave differently; elsewhere the row is not drift. `stt.language_code` uses it: only the local recogniser auto-detects, and every other provider resolves `"auto"` to `en-US`, so a stored `en-US` there already behaves like the default and is not reported. The row's own value is still detected on the base document alone, but the predicate reads the EFFECTIVE `stt.provider`, with `config.local.json` merged over the base, because the recogniser that runs is the effective one. The subagent turn budget follows 1000 automatically when its key is absent, including in an existing installation after an update. Every valid stored value is retained, even 100: a materialized old default and an explicitly selected 100-turn cap are indistinguishable without historical per-key provenance. The defaults report exposes the change and the existing adopt/keep commands; it does not claim that ambiguous legacy files can be upgraded without risking an operator's deliberate cap. The three-hour execution timeout remains separate. ### Auto-adoption, and the line it does not cross Reporting is the right answer only while the two readings of a stored value are indistinguishable AND holding the old value is survivable. On the two agent timeout budgets neither holds: an install carrying `agent.subagent_timeout_secs: 1800` reaps every subagent at 30 minutes on a build whose default is 10800, and its operator sees timeouts instead of results having never chosen 1800. Nor on the spawn memory floor: a materialized `agent.spawn_min_memory_gb: 4.0` keeps 4 GB free after every start, which a 16 GB laptop rarely has, so its subagents wait in the queue and essentially never start. The existing mechanism's only answer was a CLI command they have no reason to know exists. So `SupersededDefault.auto_adopt` opts ONE entry into a one-shot rewrite. What keeps the set small is **not** a judgment about how wide the value's range is. That criterion was tried and is wrong: `instances.warm_set_cap` is numeric with a range, and an operator running five crews who types 5 stores exactly the old default. The line that holds is whether the old value is something an operator sets on purpose. For four rows another suite pins it as a supported configuration, and the table names that test; every other report-only row gives its plain reason: | Key | Change | | Pinned by, or why | |---|---|---|---| | `agent.subagent_timeout_secs` | 1800 -> 10800, #8891 | adopts | -- | | `agent.chat_turn_timeout_secs` | 7200 -> 14400, #8949 | adopts | -- | | `agent.spawn_min_memory_gb` | 4.0 -> 2.0, #15890 | adopts | -- (the 4.0 inputs in the admission tests set a floor, they do not pin a stored 4.0 as supported; the opt-out is `0`, not the old default) | | `session.autocompact_pct` | 90.0 -> 70.0, #4388 | reports | `test_a_persisted_ceiling_value_is_left_alone` | | `dashboard.loop_stall_exit_after_secs` | 25 -> unset, #6651 | reports | `test_explicit_desktop_default_is_preserved_for_managed_service` | | `stt.streaming` | false -> true, 0.5.0 | reports | `test_put_persists_streaming` | | `mcp_gateway.forward_declared_env` | false -> true, #4566 | reports | `test_a_real_false_still_turns_it_off` | | `stt.model` | "turbo" -> "base", 0.5.0 | reports | a picker value; adopting changes transcription accuracy | | `instances.warm_set_cap` | 5 -> 0, #7248 | reports | 5 is an ordinary deliberate cap | | `agent.subagent_max_turns` | 100 -> 1000, #12203 | reports | an explicitly stored 100 is a supported cost/turn cap | | `agent.session_control` | false -> true, #8375 | reports | false is the supported global withdrawal of peer-session tools | | `skills.lazy_load` | false -> true, #12131 | reports | false is the supported switch to the short skill entry; its `note` reaches the load line too (`meaning_moved`, above) | | `agent.subagent_spawn_stagger_secs` | 2.0 -> 0.25, #12203 | reports | the knob to raise when the host or provider is the bottleneck | | `session.watchdog_rss_max_mb` | 1536 -> 0, #16393 | reports | 1536 is an ordinary deliberate ceiling; holding it only recycles idle sessions, which keep their history | | `stt.language_code` | "en-US" -> "auto", #9246 | reports | a locale picked on purpose; adopting changes what the recogniser listens for. Only where the effective provider, overlay included, is local (`applies`): elsewhere "auto" resolves to en-US | | `decisions.history_budget_chars` | 0 -> 2000, #12928 | reports | 0 is a supported setting below the consented history ceiling; adopting raises it only up to that ceiling (carries a `note`) | | `watchdog.stale_window_secs` | 300.0 -> 600.0, #8949 | reports | a tuning knob; holding it probes a live think sooner, which regenerates it | | `watchdog.tool_stall_suspect_secs` | 3600.0 -> 5400.0, #8949 | reports | a tuning knob; holding it cancels a tool the oracle cannot attest after an hour, with no re-run, so that tool's work is lost. Coupled with the hard cap (`note`): the window is the smaller of the two, so both must be adopted to get 5400 | | `watchdog.tool_stall_hard_cap_secs` | 3600.0 -> 7200.0, #8949 | reports | a tuning knob; holding it caps the same forbearance at an hour, with the same cost. Coupled with the suspect window (`note`): adopting either one alone leaves the window at an hour | | `watchdog.model_silent_probe_secs` | 900.0 -> 1800.0, #8949 | reports | a tuning knob; holding it probes a long silent think sooner, which regenerates it | | `agent.session_start_concurrency` | 2 -> "auto", #17055 | reports | the knob to lower when the host or provider is the bottleneck; holding 2 only serialises a burst of starts | A row whose old value is a supported configuration is not stale noise, whatever its type. `test_only_unpinned_broken_budgets_adopt_themselves` pins both sets by name -- the rows that adopt, and the rows that only report -- so a new row has to be placed in one of them on purpose, and moving a row across edits that test in the same change. `auto_adopt` defaults to False, so a new row is report-only until someone states otherwise. Two further properties make the rewrite safe on the rows that remain, without the per-key provenance the config layer still lacks: - **One-shot.** `auto_adoptable()` excludes any key already in the sidecar's `adopted` map, so a key is adopted at most once per install and a value the operator sets back afterwards is theirs forever. Without that record the loader would re-remove a restored value on every load -- the one behaviour worse than saying nothing. - **Marker first, chosen for its worst case.** `record_adoptions` writes the ledger entry BEFORE the removal, inside the config write lock, and a failing record aborts the whole migration write. `config.json` and the sidecar are two files with no shared transaction, so exactly one of two windows exists and the ordering decides which one: | Ordering | Window | Consequence | |---|---|---| | marker first (this) | durable marker, failed removal | the operator KEEPS their value and the adoption is not retried. A **missed improvement**, recoverable with `kirocrew config defaults --adopt` and still named by the startup line. | | removal first | landed removal, lost marker | nothing suppresses a later adoption, so a value the operator deliberately RESTORES is deleted a second time. A **destroyed choice**, unrecoverable. | No retry, rollback, compensating write or two-phase pending/committed scheme removes the window -- each only moves it onto another write that can fail the same way, and a rollback that fails recreates the hazard it exists to prevent. So the marker goes first and the failure lands on the recoverable side. `test_a_failed_write_keeps_the_value_and_does_not_re_adopt_later` pins that path end to end, including that a later load does not take the value on a retry, and `test_the_unapplied_adoption_is_still_reported_so_it_is_recoverable` pins that the marker suppresses the retry but not the report. - **An adoption that did not reach disk drops the validated-data cache.** Only a load that READS the base document can decide an adoption (`adoptable` is empty on a cache hit, by design), so a read-and-skip that left its document cached would have every later load serve the stale value and never retry. A contended lock and an exception caught by the best-effort handler both skip the write, so `persist_write_back` tracks `adoption_landed` separately from `persisted` (which starts True so the `connections_ui` marker still lands on a load that needed no migration) and invalidates in a `finally` both share. Both variables are bound before the `try`, or an early exception would turn a logged write-back failure into a `NameError` out of `load()`. The **degraded-sections branch is deliberately excluded**: its retry condition is "the operator fixes the file and restarts the gateway" (a degradation observation is sticky for the life of a process), not "the next load", so an invalidation there would only re-read and re-parse `config.json` on every load for as long as a malformed section coexists with a stored stale timeout. After the restart the fixed file's fingerprint misses the cache and the adoption retries on that first load (`test_a_degraded_load_keeps_its_document_cached_instead_of_re_reading_forever`). - **An unreadable ledger adopts nothing.** `_read_ack_document_status` returns `(document, readable)`, and `auto_adoptable` returns `[]` when a sidecar exists but cannot be parsed: reading it as empty would re-arm the one-shot over a value the operator restored. The ACK half stays fail-soft, because a missed ack costs one report line rather than a deleted setting. `record_adoptions` takes the sidecar lock **single-shot** (`wait_for_lock=False`), because the config load path runs on the asyncio event-loop thread in places and a blocking acquire there stalls the gateway for as long as another writer holds it. A contended acquire raises `BlockingIOError`, which the migration treats as "defer to the next load" -- the same deferral `_persist_config_migration` already takes on a contended config lock. A CLI caller keeps the wait: no loop to stall, no later retry. `--keep` still wins: an acknowledged value is not drift, so affirming a key before it is adopted keeps it, which is the answer for an operator who did choose 1800. Adoption is an **un-materialization**, not a write of the new number: `drop_drifted_keys` removes the stored key, so the field resolves through `data.get(key, DEFAULT)`. That holds only until the next FULL rewrite of `config.json` -- any settings save re-materializes the current number, as `drop_drifted_keys`' own docstring says -- so a later default move on the same key still needs its own registry row and does not ride along. It is applied in memory as well, because the gateway reads these budgets once at startup -- a disk-only fix would leave the run that performed it still holding the old value, which is exactly the "upgraded and nothing changed" complaint. Two guards on that half: `_adopt_in_memory` reads the field's OWN dataclass default rather than the row's `new_default` (on a key whose default moved twice, the matching row's `new_default` is an intermediate value), and it replaces the value only when the parsed field still EQUALS `old_default`, so a value the loader clamped or coerced keeps the loader's correction. A key the `config.local.json` overlay supplies is cleared on disk but left alone in memory: the overlay is the operator's live choice. After the config write succeeds, the loader warns at the default log level for each adopted key, naming the removed value and the `kirocrew config set` command that restores it. A deferred or failed write emits no adoption notice. The warning describes the stored value without claiming to know whether the operator chose it. That warning is one line in one gateway log, so the same two facts are replayed on demand: `kirocrew doctor`'s `Stored Defaults` section and a bare `kirocrew config defaults` both render the sidecar's `adopted` map through `adoption_summary` -- one `adopted:` line per key, naming the value removed from `config.json` and the exact restore command -- and both render it BEFORE opening `config.json`, so a missing or unreadable config does not hide what an earlier load removed from it. The line says "removed from `config.json`", not "the default now applies", because `config.local.json` may still carry the key; and it allows for the marker-first window (an entry whose config write failed describes a value that is still stored and still listed as drift). Both fields come from the sidecar, a file the agent sandbox can write, so they are untrusted output: every character is rendered terminal-safe (control characters escaped, never executed), and the pasteable restore command is built only from `SUPERSEDED_DEFAULTS` literals after matching the entry by key and exact value -- no quoting scheme is portable across every shell an operator might paste into, so an entry the registry does not vouch for is shown, escaped, with no command. An adopted key holds no stored value any more, so it is neither drift nor an `--adopt`/`--keep` target; naming one there is refused like any other non-drifted key. **Downgrade residual.** A build older than the adoption ledger serializes the sidecar as `{"acked": ...}` only. Running that build's `--keep` or `--adopt` after a downgrade therefore rewrites the file WITHOUT the `adopted` map, which re-arms the one-shot: on the next upgrade a value the operator restored to the old default is adopted a second time. Current builds carry both maps through every write (`_update_map`) and refuse to rewrite a sidecar they cannot parse, so the window exists only across that specific downgrade-then-write sequence, and the second adoption still announces itself at WARNING with the restore command. `stt.provider` is deliberately absent from `SUPERSEDED_DEFAULTS` even though its default moved to `local`: `_validated_stt_provider` coerces a retired value at parse time, so the stored value never wins and there is no *default* for an operator to adopt. It is instead a **coerced value**, tracked separately in `COERCED_VALUES` — see below. Both sides of an entry are **history**, so both are literals: a later change to the same key APPENDS a new entry rather than editing an existing one, which keeps the older row a true record of the change it describes. What must stay current is the END of each key's chain -- `test_every_registered_key_ends_at_the_live_default` asserts the newest entry per key names the default the loader actually applies, so moving a default without appending a row fails rather than leaving the report telling operators to adopt a value that no longer exists. Two surfaces render it. Neither writes config; the load path's adoption above is the only thing that does, and a key it adopts is excluded from the warning rather than pointing the operator at a command for something already fixed: - The load path emits **one** warning naming every drifted key plus the command to resolve it, evaluated on the **stored base document before the `config.local.json` merge** -- an overlay value is the operator's live choice and says nothing about what the base materialized, so a base drift is still reported when an overlay masks it, and an overlay-only value is not reported at all. One line rather than one per key: the registry is append-only, so a per-key line grows without bound on exactly the long-lived installs with the most real drift, and it lands on every short-lived `kirocrew` invocation, where the once-per-process guard buys nothing because there the process IS the invocation. The per-key text is still emitted at debug, so `-vv` keeps it in the log. The one per-key text the line does carry is the note of a `meaning_moved` row, because that stored value now selects a different behaviour from the one its operator chose, and the line is what someone who never opens `config defaults` reads before running `--keep`. - `kirocrew doctor` prints a `Stored Defaults` section reading `config.json` directly. Drift is informational and does NOT become an issue; an unreadable or malformed config does. ### The legacy `skills.lazy_load` rewrite 0.6.x and earlier materialized `skills.lazy_load: false` into every `config.json` they saved, and `false` was then the default full skills listing. Since 0.7.0 `false` selects the short skill entry and the default is `true`, so an upgraded install silently runs the narrowest mode. This is NOT an `auto_adopt` row: on a 0.7+ install `false` is the supported switch to the short entry, so value equality cannot tell noise from a choice. What can is WHEN the value was written, the same boundary the `connections_ui` launch migration uses. `migration.legacy_lazy_load_rewrite_due` decides on the BASE document, on a load that read it (never on a cache hit). It returns the writer's stamp only when all of these hold; every other case leaves the value exactly as stored: - the base stores exactly the boolean `false` (an explicit `true`, `0` or a string is never touched); - `meta.lastTouchedVersion` parses as `major.minor.patch` with at most a short suffix and names 0.6.x or older. An absent, non-object or unparsable `meta` is an unknown writer, not a provable one, and declines; - `connections_ui_migrated.json` does not exist. It first shipped in 0.7.0-insider.1 and every clean later load writes it, so its absence proves no 0.7 build has loaded this home; - the adoption ledger is readable and does not name the key. This keeps the rewrite one-shot even if the marker is later deleted; an unreadable ledger declines. **What stays unmigrated.** The registry's report-only `skills.lazy_load` row (`meaning_moved`) reports a declined `false`; the loader skips that key in the superseded-default line only in the load that removes it. The marker predates the meaning change: 0.7.0-insider.1 to insider.5 wrote it while `false` still meant the full listing, so an install that ran one of those keeps its materialized `false`. That is the conservative side of the boundary: the proof cannot tell those installs from a 0.7 operator who chose the short entry. The stamp is the proof, so nothing may re-stamp the document before the rewrite uses it. The load decides before it writes, and `refresh_config_meta_stamp` (the gateway's post-bind refresh) holds its refresh while the rewrite is still due, so a degraded first load leaves the proof for the next clean one. A writer that runs before any load and re-stamps -- `kirocrew config set --file`, which replaces the whole document with the operator's own -- ends the cohort, which is correct for a document the operator just supplied. So does any other real write by this build during a degraded session (a settings save, `config set`): those writes are this build's bytes, and holding the stamp across them would let the rewrite undo a value set on this build, which the in-lock re-check exists to prevent. The rewrite then rides `MIGRATE_SKILLS_LAZY_LOAD` through the same write-back as an adoption: re-detected inside the config write lock (still an exact `false`, still stamped 0.6.x or older; `kirocrew config set` and every settings save re-stamp, so a value set since the load's read is never undone), recorded in the adoption ledger as `{"skills.lazy_load": false}` BEFORE the key is removed, in the SAME record as any superseded-default adoption of that pass (two records would let a failed second one strand the first as adopted with nothing removed), the key un-materialized rather than written as `true`, the in-memory value moved only once the write is confirmed and only where `config.local.json` does not supply the key, the validated-data cache dropped when the write did not land, and the `connections_ui` marker deferred with it. One WARNING per rewrite names the key, the writer's version, why, which value now applies (the default, or the overlay's when `config.local.json` sets the key), and the `kirocrew config set skills.lazy_load false` that chooses the short entry. The `skills.lazy_load` registry row has the ledger entry's key and old value (`superseded_defaults.LEGACY_LAZY_LOAD_ADOPTION`), so `adoption_summary` vouches for it and `doctor` and `kirocrew config defaults` replay it with a restore command (a bool spelled as JSON); the marker-first residual leaves a stored value that row still lists as drift. A load that declines on an unreadable ledger or stamp still writes the marker, so that install keeps its value for good; a degraded or deferred load writes neither and the next clean load decides again. `test_config_lazy_load_legacy_migration.py` pins the truth table. ## Acknowledging a superseded default Value equality alone cannot falsify a report, so before #7559 an operator who deliberately chose a value equal to a superseded default was told about it on every load forever, with no way to answer -- and that unanswerable line competed for attention with the genuine drift on the same install. `kirocrew config defaults` is the surface that resolves the ambiguity the load path must not resolve for anyone: - no flag lists each drifted key with its stored value, the current default, and the release that changed it, marking anything already affirmed; - `--adopt [KEY...]` REMOVES the stored keys, so `data.get(key, DEFAULT)` resolves the current default from the next load and the next full rewrite materializes it. Rewriting is safe here where it is not on the load path because the operator asked by name, and only a key whose stored value IS the superseded default is ever removed. Detection runs again inside the write lock, so a value changed since it was listed is left alone. It prints each key it removed, and asks for a gateway restart only for the keys `requires_restart` marks, the schema's one statement of which fields a running gateway cannot adopt (see *`restart=True` is the single source of restart truth*). That makes a row whose consumer reads its key only at boot a row that must carry the mark: `dashboard.loop_stall_exit_after_secs` does, because the gateway sizes its loop-stall watchdog from it once at start; - `--keep [KEY...]` records the stored values as intentional, which suppresses the load-path line for exactly those values. A kept row that carries a `note` prints it again in the confirmation. An acknowledgment records ` -> the acked VALUE`, not the key alone, so it covers the choice rather than the key: change the value later and the report returns. Acks live in `~/.kiro/crew/superseded_acked.json` (`{"acked": {...}}`), **not** in `config.json` -- a `to_dict()` rewrite carries only schema fields, so the same materialization behaviour this whole mechanism reports on would silently drop an ack stored in the config document. Three properties of that file matter: - **Reads cannot block indefinitely.** The read runs on the config-load path, which is an event-loop path, and the file sits at a path the agent can name -- where `open()` on a FIFO waits for a writer forever and would wedge the gateway rather than merely delay it. `_read_ack_document` therefore `lstat`s and refuses anything that is not a REGULAR file (links included) or is over `ACK_MAX_BYTES` (64 KiB), opens with `O_NONBLOCK | O_NOFOLLOW` where the platform has them so a leaf swapped after the `lstat` fails instead of waiting, re-checks the OPENED object with `fstat`, then finishes with one capped `os.read`. - **Every refusal fails soft.** Missing, non-regular, oversized, unreadable, malformed, not an object, a non-string key: an ack suppresses one line and changes no runtime behaviour, so the worst consequence of ignoring a broken file is being told again. That soft read is also why the file carries no schema version -- any shape it cannot understand is already handled, so the field would have no reader. - **Writes never RESOLVE the leaf.** The config writers deliberately FOLLOW a link, because symlinking `config.json` into a dotfiles repo is a supported setup; here that would let a link planted at this path redirect the write onto an arbitrary file. `_update_acked` refuses a link it can see AND writes through `atomic_write`, which renames a fresh temp file OVER the leaf -- so a link swapped in after the check is replaced rather than followed. The check reports the condition; the rename is what makes it unexploitable. - **Every mutation is a locked read-modify-write.** `_update_acked` holds the ack file's own lock across read, merge and write, so two concurrent `--keep` calls cannot both read the same map and have the second replacement drop the first operator's acknowledgment. - **`record_acks` re-reads the config under its lock**, rather than trusting the caller's snapshot, and re-checks that each key is still drifted. A value changed between the listing and the call would otherwise be acknowledged at its superseded snapshot, which then suppresses the report for a value the operator never affirmed. The ack write happens inside that same config lock hold; lock order is config-then-ack wherever they nest. A settings save is itself an answer, exactly as `--keep` is. When a config PATCH writes the old default of an `auto_adopt` row, `api_kirocrew_config_patch` calls `ack_explicit_write` under the config write lock and BEFORE the write (the same config-then-ack order), so the value is acked and the next load does not un-materialize the choice the operator just made. A refused ack raises `OSError` and aborts the write, rather than persisting a value the next load would delete. Only `auto_adopt` rows are acked this way; a report-only row loses nothing on load and is left untouched. `--adopt` also drops the ack for a key it removed, since the acked value is no longer stored and keeping it would silence a genuinely deliberate choice made later. When `config.local.json` also carries an adopted key the report says the overlay still overrides it, because the EFFECTIVE value did not change. Every filesystem refusal on these paths surfaces as a controlled non-zero CLI error, never a traceback. `doctor` LISTS an acknowledged entry rather than hiding it: an ack answers the unsolicited load-path line, while `doctor` answers "what does this install still hold?", and hiding an affirmed value would make that answer wrong. ## Coerced values (removable, never affirmable) `COERCED_VALUES` in the same module tracks a second, distinct kind: a stored value the loader **replaces** at parse time rather than merely overriding. The difference decides what an operator may do about it. A superseded default still wins, so it may be a deliberate choice and must not be rewritten. A coerced value cannot win, so there is nothing to preserve — which makes removing it unambiguously safe and makes affirming it meaningless, and `--keep` refuses it by name rather than promising a setting that never takes effect. Left in place it is inert bytes that cost a warning on every load, forever, because a load never writes. One entry today: `stt.provider`. Its retired values (`whisper`, `mlx`, `parakeet`, `faster`) degrade to `local`; any other unknown value degrades to `off`, so a value nobody can account for never selects the in-process native recogniser (kirodotdev/KiroCrew#13179). Because the two resolutions differ, `resolves_to` is a callable of the stored value (delegating to `sections.stt_provider_resolution`, the loader's own pure resolver) and the entry carries the section `default`. `--adopt` uses both: it REWRITES a coerced key as what it resolves to, or deletes it only when that resolution IS the default (an absent key resolves to the default). So a retired name is adopted by removal and runs as `local` before and after, and an unknown name is adopted as the literal `"off"` and runs as `off` before and after. Adoption never moves the effective provider — deleting an unknown value would have made the default `local` apply, handing the incident's user back the engine they were escaping — and `coercion_summary` names the resolution instead of claiming removal changes nothing. The `is_coerced` predicate rides on the ENTRY, not in the detector's loop, so appending a retirement is genuinely sufficient — a detector switching on `dotted_key` would leave an appended entry silently unreported, with no test red and an operator stuck with a warning nothing can clear. That predicate delegates to `sections.stt_provider_is_coerced()` rather than restating a provider list, so the surface offering to remove a value cannot come to disagree with the loader about which ones are dispatchable. The retirement notice in `_validated_stt_provider` names the command, for the same reason the drift line does. **Why nothing is corrected automatically.** At least one registered key also has a documented escape hatch (`mcp_gateway.forward_declared_env`, whose stored `false` is pinned as honoured by `test_a_real_false_still_turns_it_off`). On disk that escape hatch and a stale materialized default are the same bytes, so a rewrite cannot correct one without overriding the other. Telling them apart needs per-key provenance -- a record of which keys the operator actually set -- which this layer does not have. ## Config Overlay (config.local.json) User overrides can be placed in `~/.kiro/crew/config.local.json`. This file is deep-merged on top of `config.json` at load time and is never touched by `kirocrew setup` or package upgrades. Because `save()` keeps an overlay-owned leaf OUT of `config.json`, such a value exists only here, so the overlay rides the dashboard export and the snapshot `config` component next to `config.json` (see [Settings import](#settings-import-dashboard-merge)). `config.json` is the persistent settings file, not a generated one: no upgrade or restart regenerates or resets it. The routine writers (dashboard PUTs, keyed `kirocrew config set`, setup, boot migrations) are locked delta read-modify-writes of the keys they own (`update_config_locked`), and the whole-document `KiroCrewConfig.save()` refuses to publish a snapshot of a file it could not read (see its API entry). Two explicit, user-invoked paths replace the document instead: `kirocrew config set --file` installs the given file whole (under the same lock, with `on_corrupt="reset"`, so it also overwrites an unreadable file), and the dashboard import's **Replace** installs the archive's `config.json`. Its **Merge** never overwrites a settings document this install has (see [Settings import](#settings-import-dashboard-merge)). The overlay is for a value you want PINNED above whatever those writers later put in the base. Resolution order: 1. Load `config.json` (the persistent settings file every writer updates) 2. Deep-merge `config.local.json` on top (user-owned, never touched by setup/migration) 3. Return merged result ### CLI Usage ```bash # Pin a setting in config.local.json (wins over config.json): kirocrew config set --local agent.log_level debug # Save to config.json (the persistent settings file): kirocrew config set agent.log_level debug ``` `config set --local` does not validate the dotted key: it warns when the top-level section is unknown and writes any leaf into `config.local.json`, which outranks `config.json`. Only the plain `config set` answers `Unknown key`. The enum, beacon and tailnet gates run before that split, so they apply to both. A list-typed key (`agent.apps_trusted`, `slack.allowed_users`, ...) takes a JSON array (`kirocrew config set agent.apps_trusted '["a","b"]'`); anything else is refused with nothing written. Details: [cli](cli.md) (`config set`). ### `config_local_path() -> Path` Returns `~/.kiro/crew/config.local.json` (or `$KIROCREW_HOME/config.local.json`). ### `_deep_merge(base: dict, overlay: dict) -> dict` Recursively merges overlay into base. Dict values merge recursively; all other types in overlay replace base values. ## Browser UI preferences (ui-prefs.json) `~/.kiro/crew/ui-prefs.json` is a backup of the dashboard settings that live in the renderer's `localStorage`, not in `config.json`. It exists because `localStorage` is keyed by ORIGIN and, in the desktop app, stored inside Electron's `userData` directory, so a moved dashboard port, a relocated `userData` directory, a switch between the stable and nightly builds, or a browser storage eviction wipes every setting the user chose — and reads to the user as "the upgrade lost my settings". Owned by `kiro_crew/ui_prefs.py`, served by `GET`/`PUT /api/ui-prefs`, and consumed by `website/src/lib/uiPrefs.ts`. Deliberately NOT a section of `config.json`: - `KiroCrewConfig.save()` re-emits the whole dataclass, so a key the running build does not model is dropped on the next save. A bag of client-owned UI keys is exactly the shape that loses that fight. - `config.json` is operator-facing; renderer layout keys do not belong in it. Contract: - Values are opaque strings the server never parses. For most keys the value is exactly what `localStorage` holds (a UTF-8 string). ONE durable key — `mc-chat-config`, a JSON object of ~20 independent chat settings — is the exception: it is NOT stored whole. Each of its fields travels under its own wire key `mc-chat-config.`, whose value is that one field's JSON fragment, so the server's existing per-KEY merge becomes a per-FIELD merge and a profile uploads only the fields it actually changed (issue #15236; before this, a second origin holding one stale field uploaded the whole blob and overwrote fields it never touched). The field name is hex-encoded (`[0-9a-f]`, 4 digits per code unit) for two reasons: a raw name could contain the `.` separator, and `showContextTokens` — a real field — contains `token`, which the credential denylist below would reject, 400-ing every flush. Hex provably contains no denied substring and no separator and reverses exactly. The server stays oblivious: it just holds more, smaller opaque keys. The file is `{"prefs": {...}}` and nothing else: an earlier revision carried a `version` and an `updated_at` that no code read, and the loader is tolerant of any shape it does not recognize, so a future reshape needs no version field to be safe. - `PUT` is a merge patch; a `null` value deletes its key. BOTH methods are owner-gated: the write so a viewer cannot overwrite the owner's settings, and the read because some values name real paths on the host (the file explorer's saved state, the cloud launch defaults). Gating the read costs nothing, since a non-owner can never have written a backup. - Keys whose name looks like a credential (`token`, `secret`, `password`, `credential`, `api_key`, `apikey`) are refused on write and filtered on read, so the dashboard bearer token can never land here. - Bounds: 200 keys, 128-char keys, 64 KiB per value, 512 KiB total. A patch that would breach them is rejected WHOLE; nothing partial is written. The client answers a rejection by retrying the patch one key at a time, so one unstorable value cannot discard the valid changes bundled with it. - An unreadable or malformed file means "no backup", never an error: the client falls back to whatever `localStorage` holds. - The client reads the backup when this profile has never successfully reached the host (`mc-ui-prefs-synced` absent) and only fills keys that are absent, so it can never clobber a value the running profile already has. Keyed on never-synced rather than no-settings-present so a boot whose fetch failed retries on the next one instead of forfeiting the restore. Otherwise the client is write-only: the backup is a backup, not a live cross-tab sync channel. - When the restore actually wrote something the page RELOADS instead of rendering. Restoring before the first render is not sufficient on its own: static imports are evaluated before the entry module's first statement, so a store that reads its key at module scope (`hooks/useBottomTerminal.ts`) has already captured the pre-restore value and its first write would persist that stale copy back over the restored one. The reload is correct for every module-scope reader without a per-store re-init hook, costs one extra load on a fresh profile, and cannot loop because the synced marker is written first. - Every key the host holds is baselined with whatever is in `localStorage` for it after the restore — including a local value that DIFFERS from the host's and was kept. The hydrating origin is by definition the one that has not been syncing, so uploading its value on first flush would overwrite the newer backup with a possibly months-stale one; baselined, it stays in use locally and is uploaded the moment the user changes it. A value the quota-safe writer had to drop is NOT baselined, or the first flush would read it as a deletion and null out a good host backup. - A FAILED first restore writes `mc-ui-prefs-hydrate-pending` holding the list of durable keys the profile held AT THAT MOMENT. On the next successful restore a key in that list — and any change the user made to it since — is the user's, so local wins as usual; a key NOT in the list was written after the failure by a page that rendered without its settings (a login screen counts; mount-time hooks persist defaults), so for it the HOST wins — treating those defaults as the user's choice would upload them over the real backup. A repeat failure never widens the list. A never-failed first restore keeps the normal rule. Letting the host win for every key instead had the mirror-image defect: a returning user whose GET failed once and then changed a preference saw the stale host value overwrite the change. The marker is cleared only after the synced marker is written, so a crash between the two leaves the profile pending rather than synced-with-untrusted-locals. - `mc-ui-prefs-synced` also holds the NAMES this profile last synced, which is how a deletion made before a reload is still reported as a `null` while a key this profile never synced is never nulled — that is what stops a second browser from deleting the first one's settings. - Which keys are durable is the client's decision (`DURABLE_PREF_KEYS`). Session-scoped and derived state (height caches, panel tabs, drafts, touched files) is excluded, as are the settings `config.json` already owns (theme mode/colour, language, onboarding flags) and keys a migration deliberately deletes (`mc-zoom`, `mc-font-scale`), so no setting has two homes and nothing resurrects a key a migration removed. The per-surface prefs that used to silently reset across origins -- notification sound (`mc-notification-sound`), interface mode (`mc-ui`), reading width (`mc-reading-width`) -- are durable. - A SETTINGS IMPORT rewrites this file under a profile that has already synced, which the cold-profile rule above would never re-read. The client therefore pauses the sync before sending the import (`pauseUiPrefsSync`: no flush, including the `pagehide` one, may land after the import and put this page's values back), and when the response says `ui_prefs_restored` it calls `adoptHostUiPrefsOnNextLoad` and reloads: that removes `mc-ui-prefs-synced` and writes an EMPTY `mc-ui-prefs-hydrate-pending`, so the next load cold-hydrates and the HOST wins for every key it holds (by the failed-restore rule, a key not in the list is the host's). A key the host does not hold keeps its local value. Server side a Merge installs the archive's copy only where the host has no `ui-prefs.json` (`ui_prefs.install_imported_ui_prefs`, which decides under the module lock that serializes it against the PUT handler); a host that keeps one keeps it whole, since the browser adopts the host copy on its next load. Either mode holds each entry to the same deny-list and bounds as a PUT (an unstorable entry is dropped and counted, not fatal). - GROWING the durable set is guarded by a reconcile pass (growth-gap issue 9491). A warm profile never runs the cold restore, so a key added to the allowlist by an upgrade would otherwise be flushed at its local DEFAULT -- often written by a hook on mount (`mc-ui` is the live example) -- overwriting the value another origin backed up. Before the first flush after an upgrade that added keys, the client reads the host copy once for the keys this profile has never synced: a key absent locally adopts the host value (with the same reload-if-restored rule as the cold path), a key present locally is baselined so the first flush does not upload it (it goes up when the user next changes it), and a key the host does not hold is seeded from local. The reconciled roster is recorded as a reserved entry inside `mc-ui-prefs-synced` -- deliberately not its own key, so a downgraded (pre-roster) build's next fingerprint rewrite sheds it and a re-upgrade reconciles again instead of trusting a stale roster -- making the pass one GET per allowlist growth, not per boot. A profile with no roster predates the mechanism and is baselined against the frozen pre-mechanism allowlist, so only genuinely new keys are reconciled: a key the profile merely never held is NOT bulk-imported from another origin. A failed reconcile suppresses the sync for the session -- flushing unreconciled keys is the exact clobber the pass exists to prevent -- and the next boot retries; the failed-restore marker applies as on the cold path, so a default written by a settings-less render between the failure and the retry loses to the host. - The composite (`mc-chat-config`) rides that SAME reconcile and failed-restore machinery at per-FIELD granularity, with four wrinkles worth knowing before touching `uiPrefs.ts`: - MIGRATION from a host backup written by a pre-split build (one whole-blob `mc-chat-config` key): the whole blob is expanded into child wire keys once, on both the cold restore and the warm reconcile, so an existing user's chat settings are not lost on the first load after upgrade. A field the host already holds as a child key wins over the same field inside the legacy blob; neither copy is retired (that is more machinery than the issue needs, and a child key is authoritative the moment any fixed build writes it). - CLEARING a reconciled child: a plain key clears when it is in the synced fingerprints OR the reconciled roster. A child gets BOTH — a child the host held is fingerprinted; a child the host did NOT hold has nothing to fingerprint, so the reconcile records it in the roster regardless of outcome. Without the roster entry such a child would read as unreconciled on every boot — withheld from flush forever while the first flush of any other key fingerprinted it from a value never sent, silently dropping the user's choice. The reconcile trigger therefore fires only while a child lacks BOTH a fingerprint and a roster entry (or a legacy whole-blob fingerprint is still present), so a fully reconciled profile pays no per-boot reconcile GET and a transient GET failure cannot cost a whole session's backup. - DOWNGRADE: the failed-restore marker records per-field child entries PLUS the parent `mc-chat-config` key. A pre-split build iterates the whole-blob allowlist and reads ownership by the parent, so the parent entry keeps its stale host blob from overwriting local chat settings. A split-aware build does NOT treat the parent as a blanket grant over its children — it honours the parent only on a legacy marker that carries no child entries at all; when child entries are present the exact child key is required, so a default a settings-less render wrote AFTER the failure cannot masquerade as owned. - DIRTY marker: `saveChatConfig` records the fields the user edited in `mc-chat-config-dirty` (`markCompositeFieldsDirty`). A dirty child is never withheld and always uploads, even at its default value; it counts as owned during restore and reconcile; its marker clears only for a child the host acknowledged and only once the synced fingerprints persisted; and host adoption (`adoptHostUiPrefsOnNextLoad`) removes the key. - Also excluded: any value that GATES A SAFETY CONFIRMATION. `mc-yolo-ack` is the instance — its presence makes the approval-mode picker skip the confirmation and enable full auto-approval — and the reason is that this file sits in the agent-writable data home, so a restorable ack is an ack an agent can forge for the user's next fresh origin. Convenience does not outrank a human gate. ## Settings import (dashboard Merge) **HTTP refusals** of `/api/portability/*` (`dashboard/handlers/portability.py`) each carry a machine-readable `code`: `file_field_required` (400), `invalid_import_mode` (400), `import_archive_invalid` (400, with the validator's detail), `invalid_components` (400, a `?components=` value other than `memory`), `memory_only_merge_only` (400, a Replace asked for memory only or given a memory bundle), `import_archive_too_large` (413), `auth_required` (401), a 409 with the refusal's own code, and `export_failed` / `import_failed` / `preview_failed` (500, opaque prose). A 4xx carries an actionable detail; a coded 5xx does not, so the dashboard's `PortabilityTab` `refusalText` shows its own localized fallback for a coded 5xx and the server's text otherwise. **Memory only.** `GET /api/portability/export?components=memory` builds an archive holding only the snapshot `memory` component (`memory.db`, `memory_index.db`, `workspace/memory/`, `workspace/knowledge/`, `memory_stores/`), never chats. Its manifest declares `components: ["memory"]` and carries `portability.MEMORY_EXPORT_MANIFEST_VERSION`, which lies outside the `1..EXPORT_MANIFEST_VERSION` range a whole-install import accepts, so a reader that does not know memory bundles refuses one rather than applying it as a whole install. `POST /api/portability/import?mode=merge&components=memory` extracts and merges only the memory members of any archive; an archive whose manifest declares memory only is imported that way without the parameter. The manifest is the one member at `/MANIFEST.json` (`portability._top_manifest`); both the validator and the import read only that one, so a deeper file with that name is ordinary content. No top-level manifest, or more than one, is refused. Memory is not redacted: the dashboard shows a confirm step listing what the file holds before it builds the export. **Member names.** Because the memory filter reads member NAMES, both `validate_import_zip` and `apply_import_zip` refuse a name an extractor could place elsewhere (`portability._unsafe_member_name`): `..`, an absolute name, a drive prefix (`C:/x`, which Windows `zipfile` strips), and, only on a host where a backslash is a path separator, a backslash-spelled form of those. On POSIX a backslash is an ordinary filename character and is kept. After each extract the import also refuses a member that did not land under its own top directory. The manifest is capped at `_MAX_SETTINGS_DOCUMENT_BYTES` (8 MiB, the budget the export sizes its own manifest to, session times included) before it is decoded. The dashboard export (`portability.create_export_zip`) carries every document a Settings choice is persisted in: `config.json`, `config.local.json`, `ui-prefs.json` and `notification_settings.json`. The snapshot `config` component carries the same set. Every archive settings document is vetted before either mode applies one (`portability._vet_archive_settings`): a document that is not its reader's shape (a config that is not a JSON object, a `ui-prefs.json` that is not `{"prefs": {...}}`, a `channel_settings` that is not an object) is refused, removed from the extraction so neither mode can install it, reported `