# dsh-smart-restart English | [中文](README.zh.md) ![dsh-smart-restart demo — real-time restart](assets/demo.gif) *A real-time restart: the agent restarts the service (canary pre-flight passing, `Canary: passed — restarting…`), and the plugin's boot notification brings the very same session right back — the conversation continues automatically, no user prompt needed.* A **DeepSeek Harness (DSH) host plugin** that keeps the main agent aware of service restarts — **without the user having to prompt it**. On every boot it detects that a new process has taken over, wakes the target agent with a short "Smart-restart" notice (boot time, previous boot, downtime). New in **v0.2.0** it added the `smart_restart` tool to **restart DSH itself** and return the notice to the exact session that asked; new in **v0.3.0** it **auto-detects** when the service was stopped while an agent session was active, so even a *plain* `systemctl restart` the agent ran notifies that session at boot; new in **v0.4.0** it can **validate the launch first** with an optional canary pre-restart gate that aborts the restart when an ephemeral boot fails. New in **v0.5.0** it **auto-detects the systemd unit** from `/proc/self/cgroup`, so the tool and the canary work with zero config on systemd-managed installs. [![CI](https://github.com/edusrez/dsh-smart-restart/actions/workflows/ci.yml/badge.svg?style=flat-square)](https://github.com/edusrez/dsh-smart-restart/actions/workflows/ci.yml) [![npm](https://img.shields.io/npm/v/dsh-smart-restart?style=flat-square&logo=npm)](https://www.npmjs.com/package/dsh-smart-restart) [![license](https://img.shields.io/npm/l/dsh-smart-restart?style=flat-square)](LICENSE) [![stars](https://img.shields.io/github/stars/edusrez/dsh-smart-restart?style=flat-square)](https://github.com/edusrez/dsh-smart-restart) [![last commit](https://img.shields.io/github/last-commit/edusrez/dsh-smart-restart?style=flat-square)](https://github.com/edusrez/dsh-smart-restart) ## Table of contents - [Overview](#overview) - [How it works](#how-it-works) - [The `smart_restart` tool](#the-smart_restart-tool) - [Canary pre-restart validation](#canary-pre-restart-validation) - [Requirements](#requirements) - [Install](#install) - [Configuration](#configuration) - [Behavior & lifecycle](#behavior--lifecycle) - [Limitations](#limitations) - [Development](#development) - [License](#license) ## Overview A long-lived DSH instance restarts for many reasons: the agent installs or reconfigures a plugin and triggers a reboot, the host machine reboots, or the service is restarted under **systemd**. After any such restart the agent is up and running again, but it has **no idea that anything happened** — the previous session is gone. Today, the only way to get the agent to pick back up is for the user to tell it, usually with something like *"you restarted"*. `dsh-smart-restart` closes that gap. It detects the restart itself and wakes the main agent at boot with a short, self-contained notice describing what happened, so the agent decides whether to resume interrupted work, note the downtime, or simply acknowledge — no user prompt needed. **New in v0.3.0**, the plugin also covers the *unplanned* restart of an **active agent**. It tracks the last active session and, on shutdown (`SIGTERM`/`SIGINT`), persists a `shutdown-notice.json`; on boot it pins the notice back to that session (when it was active within `shutdownGraceMs`). So a plain `systemctl restart` the agent ran — or a restart that happened while an agent was mid-task — is auto-notified without the user prompting. v0.2.0 let the main agent *cause* restarts: the `smart_restart` tool records the calling session and restarts DSH through systemd, returning the notice to that exact session. - **Zero prompting** — the user never has to tell the agent "you restarted". - **Self-restart** — the main agent can restart DSH itself and resume its task automatically afterwards. - **Automatic wake** — an idle main agent is woken and handed the notice on its own. - **Targeted delivery** — the post-restart notice returns to the session that requested it; otherwise the `target` config decides where it lands. - **Read-before-edit guard** — `smart_restart` reads the LIVE agent registry and refuses to restart while other sessions are mid-turn (unless `force: true`), so a blind interruption is structurally impossible. - **Interrupted-head resume** — heads whose turns were cut by a restart each get an automatic resume notice at boot, so the organization never hangs idle post-restart without knowing. - **Precise context** — the notice carries the boot time, the previous boot time, the downtime, and an optional reason. - **Canary safety** — an optional pre-restart gate boots an ephemeral DSH instance and aborts the restart when it fails (see below). - **Self-contained** — a single host bundle; nothing to run, no external service. ## How it works On every boot the plugin: 1. **Reads the previous marker** — a durable `marker.json` (`{lastBootAt, pid, dshVersion?}`) stored under `//`. 2. **Detects a restart** — a previous marker with a **different `pid`** than the current `process.pid` means a *new process* has started, i.e. the service restarted; downtime is `now − previous lastBootAt`, clamped to `≥ 0`. 3. **Writes a new marker** immediately, so the next boot can be compared against this one. 4. **Checks for a pending notice** — if the `smart_restart` tool left a `pending-notice.json` before the previous restart, this boot **pins** delivery to that exact session (**highest priority**; the file is consumed once read). 5. **Checks for a shutdown notice** — if there was **no** pending notice, the plugin reads `shutdown-notice.json` (written by the previous process on `SIGTERM`/`SIGINT`, recording the last active session). If that session was active within `shutdownGraceMs`, this boot **pins** delivery to it (**second priority**); otherwise the pin is skipped and delivery falls back to `target`. The file is consumed once read. 6. **Delivers the notice** — once per boot via the source-agnostic `agent/session-start` hook (a pinned session is matched regardless of publication source, so a session that **resumes** with `source: 'resume'` is caught too), backed by a **bounded poll** (750ms tick, ~15s cap) that picks up a pinned session which resumes lazily late. There are **three delivery paths**, depending on how the restart happened: ### (a) Agent-initiated restart (`smart_restart` tool) The main agent calls `smart_restart(reason?)`. The tool: - validates the configured `restartUnit` token, - **persists the pending notice synchronously** (calling session + optional reason) to `pending-notice.json` under the state dir, *before* any spawn so it survives the imminent service kill, - spawns a **detached** `setsid bash` process (a ~1s delay lets the tool's response be written) that runs `systemctl restart ` and survives this process being killed by systemd, - returns `{ok: true, restarting: true, ...}`. On the next boot, step 4 above reads the pending notice, **pins** it to the calling session (**highest priority**), and the notice returns there — **regardless of source** — so the agent that asked resumes its interrupted task. ### (b) Smart shutdown auto-detection (plain restart while an agent was active) If there is **no** pending notice, the plugin looks for a `shutdown-notice.json`. This file is written synchronously by the previous process on `SIGTERM`/`SIGINT`, recording the **last active session** (tracked on `agent/session-start` and `agent/pre-step`) and its timestamp. On boot the plugin **pins** the notice to that session *only if* it was active within `shutdownGraceMs` of shutdown (recent activity ⇒ the user restarted while the agent was mid-task, so auto-notify). If the session was idle well before shutdown (the user probably restarted while idle), the pin is skipped and delivery falls back to `target`. This is what makes a **plain `systemctl restart`** that the agent ran — or a restart that happened while an agent was active — auto-notify that session at boot, with no user prompt and no need to have called `smart_restart`. Since **v0.6.0** the boot also auto-notifies the **interrupted heads**: every session of the recorded interrupted set that matches `ignoredSessionPrefixes` (Deepartments `head-` heads by default) receives its OWN resume notice once it comes live at boot — so after a restart that cut an active head turn, the organization is never left hanging in idle with no notice (fb-46). Heads are still never *pinned* as the single-session last-active target (the pin and the `shutdownTarget` gate keep ignoring `head-*`, preserving the v0.3.1 anti-spurious-notice behavior); only a head whose turn was genuinely cut is notified. ### (c) External restart (systemd, host reboot, dev tools) — no active session There is **no pending notice** and **no usable shutdown notice** (either none was written, or the last activity predates `shutdownGraceMs`). The boot falls back to the `target` config (`primary` | `all` | ``) to decide which agent(s) get the notice. The `source === 'startup'` gate applies only to this non-pinned path. ### Restart vs. HMR semantics The detection is intentionally precise: | Situation | Detected as a restart? | | --------- | ---------------------- | | New OS process, previous marker exists | **Yes** — service (re)started. | | Same OS process (in-process HMR / hot-reload) | **No** — ignored. | | First-ever boot (no marker) | **No** — nothing to compare against. | | Marker present but stale/corrupt `lastBootAt` | **Yes** — downtime reported as 0. | Only a genuinely **new process** counts as a restart. An in-process hot-reload keeps the same `pid`, so it is not treated as a service restart. Delivery happens **once per boot**: a `deliveredIds` set plus a `primary` guard (for the `primary` target) prevent the startup event and the fallback poll from double-sending. ### Notice The default notice (English) reads: > Smart-restart: the DSH service restarted at ``. Previous boot: `` (downtime ~2m 9s). If a task was in progress, resume it; otherwise reply with a one-line acknowledgment. When a `reason` was recorded by `smart_restart`, it is appended (`… reason: .`) before the resume instruction. The previous-boot/downtime segment is omitted when no previous boot time is known, and the whole text is fully customizable via the `notice` config (see below). ## The `smart_restart` tool Registered via `ctx.tools.register` in `apply` (so it is available to agent sessions) when `toolEnabled` is true (the default). It lets the main agent restart DSH itself through systemd. **Parameters** | Param | Type | Required | Description | | -------- | ------ | -------- | ----------- | | `reason` | string | no | Optional human-readable note, e.g. `"installed dshmarket in stable+dev"`. Included in the post-restart notice. | | `canary` | boolean | no | Optional canary pre-restart validation for THIS call: boots an ephemeral DSH instance and aborts the restart on failure (see [Canary pre-restart validation](#canary-pre-restart-validation)). Overrides the configured `canary` for this call. | | `force` | boolean | no | Explicit override of the [read-before-edit guard](#read-before-edit-guard): restart even when OTHER sessions are mid-turn. The tool refuses (and returns the in-flight session list) unless this is `true`; the override and the list are ALWAYS visible in the log. Use only after explicitly confirming the in-flight work is safe to interrupt. WINS over `wait`: with `force: true` the tool never waits. | | `wait` | boolean | no | DEFER the restart instead of refusing it: re-read the LIVE agent registry (the same in-process snapshot the guard uses — it sends no message and wakes nobody) until every OTHER session is idle, then restart. Bounded by `waitMaxMs`; once the cap is spent the tool returns the SAME loud in-flight refusal and states that the wait expired (nothing is spawned, nothing is persisted). | | `waitMaxMs` | number | no | Hard cap in milliseconds on a `wait: true` deferral. Default `120000` (2 minutes); an absent or invalid value falls back to the default, so a wait is never unbounded. Ignored without `wait: true`. It caps the **whole** deferral: the wait at the guard AND, when a canary gate runs, its post-canary re-check draw from the same budget — the re-check gets only what is left and refuses immediately once nothing is left, so the deferral can never last 2x this cap (`waitedMs` reports the accumulated total). | **Behavior** - Validates the configured `restartUnit` — a single systemd unit token (`/^[A-Za-z0-9_.@-]+$/`, no spaces/slashes) to prevent shell injection into the detached command. - **Fails fast** when `restartUnit` is not configured **and** auto-detection from `/proc/self/cgroup` finds no unit either (`ok: false`, error `restartUnit not configured`) rather than guessing a unit — an empty `restartUnit` is auto-detected first, and only a failed detection produces that error. - **Refuses to restart while other sessions are mid-turn** (read-before-edit guard, see below) — the live registry check runs before anything is persisted or spawned, and runs AGAIN after a canary pass, right before the spawn. - Persists `pending-notice.json` **synchronously** (before any spawn) so it survives the service kill and targets the restarting session. - Restarts via a **detached** `setsid bash` process (`sleep 1 && systemctl restart `) that outlives this process, then unrefs it. - Returns `{ok: true, restarting: true, sessionId, reason}` on success, or `{ok: false, restarting: false, error}` when it fails (a guard-blocked refusal also carries `inFlight: string[]`). ### Read-before-edit guard Since **v0.6.0**, `smart_restart` implements the **read-before-edit** pattern (fb-168): BEFORE persisting the pending notice or spawning the restart, the tool reads the **live agent registry** inside the DSH process — `ctx.agents`, a session whose `status === 'running'` has an active turn in flight (the same in-process signal Deepartments derives its "running" state from, and the marker of the exact work a restart would cut). This registry is updated synchronously by the harness on every `agent/status` transition, so the check **can never go stale** the way file-based state (e.g. a previously taken `posts.json` snapshot) could. - If **any session other than the calling session** is mid-turn, the tool returns `{ok: false, restarting: false, error: 'refusing to restart: N other session(s) mid-turn (…)' , inFlight: []}` — the caller is excluded because it restarts itself intentionally and is pinned for resume. - An explicit **`force: true`** is the only way through; the override and the in-flight list are **always written to the log** (`[smart-restart] smart_restart: FORCE override — restarting with N other session(s) mid-turn: …`). - Alternatively, **`wait: true`** inverts the refusal into a bounded deferral: the tool re-reads the SAME live registry snapshot in a sleep loop (default poll 1s) until every other session is idle, then restarts — it creates no turns and wakes nobody. The wait is capped by **`waitMaxMs`** (default 120000); at the cap the tool returns the same refusal with `waitTimedOut: true` and an `error` stating that the wait expired, so a spent wait is never confused with the immediate block. The deferral runs at the guard point, BEFORE the calling-session flush and BEFORE the canary gate, and `force: true` always wins. - **`waitMaxMs` is ONE budget for the whole deferral.** The canary window is exactly when a session starts a turn again (the window closes *while* it is being checked), so the post-canary re-check does not get a fresh cap: it runs the same `waitForIdle` with `max(0, waitMaxMs − spent)`, where `spent` is the deferral window since the wait opened (the guard-stage wait + the calling-session flush + the canary window). When nothing is left it refuses **immediately** with the spent total — never a zero-length wait. `waitedMs` in the result is the **accumulated** wait of both stages (guard + re-check), so the published figure is the real input of that arithmetic: either the restart happens, or there is a loud refusal — never a wait spent in silence. - The check runs at call entry **and again after a canary pass** (the canary window can be tens of seconds — the final gate reflects the current state, not the state at call entry). - A blind interruption is thus structurally impossible: the tool cannot spawn while other agents are running unless the caller explicitly confirms it. **Intended agent flow** ``` install/change a plugin → call smart_restart(reason) # e.g. "installed dshmarket in stable+dev" → DSH restarts (detached, ~1s) → after boot, the notice returns to THIS session (pinned) → the task continues automatically — no user prompt needed ``` **Safety notes** - The unit token is validated against a strict regex to block shell injection through `restartUnit` into the detached shell command. - The tool refuses to interrupt other sessions' live turns unless `force: true` is passed explicitly — a blind restart while agents are mid-turn is structurally impossible. - The tool targets a **systemd-managed** DSH install (`setsid` / `systemctl`); it does not apply to a bare process without a systemd unit. ## Canary pre-restart validation > **New in v0.4.0.** Optional, opt-in, and fully generic — it works on any DSH > install: `restartUnit` is the only tool config you may need to set (auto-detected since **v0.5.0** when empty on systemd installs). When enabled, the `smart_restart` tool can validate the launch **before** anything is persisted or restarted: it boots an **ephemeral DSH instance** from the same binary and profile as the systemd unit, verifies it starts, and only then proceeds with the real `systemctl restart`. A failed canary **aborts the restart** — no pending notice is written, nothing is spawned, and the calling session is alerted live (same plugin-source notice channel). **What it does, step by step** 1. **Resolves the dsh binary/profile** — an explicit `canaryBinary` / `canaryProfile` wins; otherwise both are derived from `systemctl show -p ExecStart ` (binary falls back to `dsh` on PATH). The `--profile` flag is omitted when no profile resolves. 2. **Creates a temp state dir and a temp patch overlay** (`dsh --patch /canary.patch.yml`, applied after the profile layer): the `smart-restart` row itself is disabled in the canary (`enabled: false`) and every `canaryStateDirOverrides` entry gets its `stateDir` redirected into the temp dir — so the canary never writes a marker, notice, or live board state. 3. **Pre-flights the launch** with `--dump-config` (20s timeout; a compose failure aborts the restart) and **checks the dump is coherent** — the composed tree must be a valid entry list and the `smart-restart` row must be DISABLED in the canary (the patch applied; an enabled row would let the ephemeral write into the live state dir). 4. **Boots the ephemeral instance** detached with the same overlay on an auto-picked free port (`canaryPort` when set) and **polls** `http://127.0.0.1:/` until HTTP 200 or `canaryTimeoutMs` elapses (per-attempt 800ms; a refused connection or non-200 is "not yet"). 5. **Post-boot validation** — with the instance up, the canary checks: the **client boot graph** (default ON — every `/plugins//client.js` must register its graph row id); **agent liveness** (default ON — every non-retired member of the deepartments catalog must appear alive in the runtime's live agent registry); **pooler health** (default ON — `/v1/models`, `/usage` and `/__keypool/status`; a missing endpoint, e.g. the fb-75 pooler-capacity deploy pending, is a graceful skip) and the **R8/R9 runtime markers** (default ON — `presence.json` + `toolset-audit.jsonl` well-formed). 6. **Stops the ephemeral** (process-group kill) and returns: `passed` → the restart proceeds; `failed` → abort + alert the caller; `skipped` → the restart proceeds (a skip is not a failure). **When it skips (never blocks)** — if the dsh binary/profile cannot be derived (no `systemctl` lookup result AND no explicit `canaryBinary`/`canaryProfile`), the canary reports `skipped` and the restart proceeds unchanged. Generic installs are therefore always safe: a canary failure only ever aborts a restart when *you* opted in with an actual, resolvable launch target. **How to enable** - **Per profile**: add `canary: true` (plus any overrides) to the `smart-restart` row. The per-profile patch row REPLACES the row's whole config, so restate your existing fields (see the example in [Configuration](#configuration)). - **Per call** (agent-facing): call `smart_restart(canary: true, reason)`, which overrides the configured default for that single call. **Config fields** | Key | Type | Default | Description | | --- | --- | --- | --- | | `canary` | boolean | `false` | Master opt-in for the canary gate. | | `canaryTimeoutMs` | number | `45000` | Hard window (ms) for the canary boot liveness probe; a timeout is a canary failure and aborts the restart. | | `canaryPort` | number | `0` | HTTP port for the ephemeral canary instance; `0` auto-picks a free port. | | `canaryProfile` | string | `''` | Explicit dsh profile for the canary launch; empty derives from the unit's `ExecStart` (`--profile`). | | `canaryBinary` | string | `''` | Explicit dsh binary for the canary launch; empty derives from the unit's `ExecStart`, else `dsh` on PATH. | | `canaryStateDirOverrides` | object | `{}` | Plugin-row id → temp dir: those rows get their `stateDir` redirected in the canary patch (e.g. `deepartments: ''` keeps the canary off live board state). Relative or empty values resolve under the canary temp dir; absolute values are used verbatim. | > **Note on overrides and row replacement.** Patch rows replace their row's > WHOLE config (no merge) — and the canary is a throwaway instance, so a > listed row gets only its `stateDir` overridden and its remaining config > reverts to that plugin's own defaults. Use `canaryStateDirOverrides` only > for rows that boot fine with defaulted config (the live profile's row is > untouched). Rows that don't exist in the canary's composed tree are skipped > with a loader warning. **Output** — when a canary ran, the tool result carries `canary` (`passed` / `skipped` / `failed`) plus `canaryDetail`, and the rendered response gains a `Canary: …` line (`Canary: passed — restarting…`, `Canary: failed — restart ABORTED: `, or `Canary: skipped — `). ## Requirements - **DSH `0.1.0-rc.7+` / `0.1.1-rc.x`** — a **long-lived instance with a live main-agent session** (the web / GUI profile). This plugin is designed for a continuously-running service whose main agent stays resident; it is **not** aimed at the one-shot headless CLI. - **systemd-managed DSH install** for the `smart_restart` tool — v0.5.0 auto-detects the service unit from `/proc/self/cgroup`; set `restartUnit` explicitly when the unit differs or on non-systemd (where the tool fails safe). - **Node.js / pnpm** — the usual DSH toolchain for building and installing host bundles. ## Install `dsh-smart-restart` is a **DSH host bundle**: `package.json` carries `dsh.bundle.patch = ./cordis.patch.yml`, so installing the package lets the plugin layer auto-join the profile's `dsh.profile.bundles`. ```bash # From a registry (npm) dsh plugin --profile add dsh-smart-restart # Or link a local checkout while developing dsh plugin --profile add /path/to/dsh-smart-restart ``` **A restart is required after `add`** — which is exactly the scenario this plugin exists to surface. The bundled patch inserts the layer into the profile's layer stack: ```yaml # cordis.patch.yml (bundled with this package) - insert: - id: smart-restart name: dsh-smart-restart config: enabled: true stateDir: .smart-restart target: primary wakeup: true notice: '' restartUnit: '' # OPTIONAL since v0.5.0 — auto-detected on systemd installs; set explicitly when the unit differs toolEnabled: true ``` > **`restartUnit` is OPTIONAL since v0.5.0 on systemd installs.** The plugin > auto-detects its **own systemd unit** from `/proc/self/cgroup`, so the > `smart_restart` tool (and the canary's `ExecStart` derivation) work with > **zero config** on a systemd-managed DSH install. Set `restartUnit` > explicitly when the unit name differs from the auto-detected one or when > detection is unavailable (e.g. a bare non-systemd process). If you DO set > it, add a `cordis.patch.yml` override on the `smart-restart` row — > **restating the full config** (a partial override would drop the other keys) > — with the unit for that profile. For example, for the stable instance and > the dev instance respectively: ```yaml # Override on the smart-restart row — profile "stable" (dsh.service) - insert: - id: smart-restart name: dsh-smart-restart config: enabled: true stateDir: .smart-restart target: primary wakeup: true notice: '' restartUnit: dsh.service toolEnabled: true # Override on the smart-restart row — profile "deepartments-dev" (dsh-deepartments-dev.service) - insert: - id: smart-restart name: dsh-smart-restart config: enabled: true stateDir: .smart-restart target: primary wakeup: true notice: '' restartUnit: dsh-deepartments-dev.service toolEnabled: true ``` Because the bundle declares `dsh.bundle`, the layer auto-joins `dsh.profile.bundles` on install — no manual profile edit required. `restartUnit` is **usually auto-detected** (v0.5.0); add the explicit override above only when the unit name differs or detection is unavailable. ## Configuration All behavior is controlled through the plugin row's `config`: | Key | Type | Default | Description | | ------------- | ------- | ----------------- | ----------- | | `enabled` | boolean | `true` | Master switch; `false` skips all processing. | | `stateDir` | string | `.smart-restart` | Sub-directory under `` where `marker.json`, `pending-notice.json` and `shutdown-notice.json` are written. | | `target` | string | `primary` | Which agent(s) to notify when there is **no** pending/shutdown notice: `primary` \| `all` \| ``. | | `wakeup` | boolean | `true` | `true` → `agent.followup()` wakes the agent and delivers; `false` → `agent.inject()` queues model-facing context only (no wake). | | `notice` | string | `''` | Optional custom notice text; returned verbatim when non-empty, else the default. | | `restartUnit` | string | `''` | Systemd unit to restart when `smart_restart` is invoked (e.g. `dsh.service` or `dsh-deepartments-dev.service`). Since **v0.5.0** it is **auto-detected from `/proc/self/cgroup`** (the plugin's own unit) when empty; an explicit value always wins. Empty with no detectable unit → the tool fails safe with a clear error. | | `toolEnabled` | boolean | `true` | Whether the `smart_restart` tool is registered (available to agent sessions). | | `shutdownGraceMs` | number | `600000` | Grace window (ms, default 10 minutes) before shutdown within which last agent activity counts as "agent-involved" for the smart shutdown auto-notification. If the last-active session was idle beyond this window on shutdown, the pin is skipped and delivery falls back to `target`. | | `ignoredSessionPrefixes` | string[] | `['head-']` | Session-id prefixes that must never be selected as the single-session "last active" PIN for the smart-shutdown auto-notification — Deepartments department-head sessions (`head-`) never get a spurious pinned notice (the v0.3.1 regression). Since v0.6.0 the same prefixes identify the **interrupted-head resume recipients**: a head whose turn was genuinely cut by a restart still receives its own notice at boot (never a pin). Configurable list; default ON (heads skipped from the pin). | | `canary` | boolean | `false` | Opt-in [canary pre-restart validation](#canary-pre-restart-validation): boot an ephemeral DSH instance and abort the restart on failure. The per-call `canary` tool parameter overrides this for one call. | | `canaryTimeoutMs` | number | `45000` | Hard window (ms) for the canary boot liveness probe (default 45s); a timeout is a canary failure and aborts the restart. | | `canaryPort` | number | `0` | HTTP port for the ephemeral canary instance; `0` auto-picks a free port. | | `canaryProfile` | string | `''` | Explicit dsh profile for the canary launch; empty derives it from the unit's `ExecStart` (`--profile`). | | `canaryBinary` | string | `''` | Explicit dsh binary for the canary launch; empty derives it from the unit's `ExecStart`, else `dsh` on PATH. | | `canaryStateDirOverrides` | object | `{}` | Plugin-row id → temp dir; those rows get their `stateDir` redirected in the canary patch so the ephemeral never writes live state (e.g. `deepartments: ''` keeps the canary off live board state). Relative or empty values resolve under the canary temp dir; absolute values are used verbatim. | | `canaryClientCheck` | boolean | `true` | Post-boot client-graph validation (P1 lesson): after liveness, parse `__DSH_BOOT__` from the served page and verify every row's `/plugins//client.js` bundle registers that row's id. A boot with no `__DSH_BOOT__` (non-web surface) passes trivially. | | `canaryClientTimeoutMs` | number | `15000` | Whole-phase budget (ms) for the client-graph validation. | | `canaryAgentCheck` | boolean | `true` | Post-boot AGENT-LIVENESS check (R8 liveness family): every NON-RETIRED member of the deepartments catalog (`posts.json`) must appear alive in the runtime's live agent registry before the restart proceeds. A member registered but missing from the live registry fails the canary (a restart that lands with heads/workers missing hangs the org). No catalog (generic install) → skip. | | `canaryCatalogPath` | string | `'/.deepartments/posts.json'` | Catalog path the agent-liveness check reads (the deepartments runtime's durable registry). | | `canaryRuntimeStateDir` | string | `'/.deepartments'` | Runtime stateDir whose R8/R9 marker files the markers check reads. | | `canaryPoolerCheck` | boolean | `true` | Post-boot POOLER-HEALTH check: probes `/v1/models`, `/usage` and `/__keypool/status` on the ephemeral web port. A missing endpoint (HTTP 404/405 — e.g. the fb-75 pooler-capacity deploy pending) is a graceful skip, never a failure; a 5xx or unreachable endpoint fails. | | `canaryPoolerTimeoutMs` | number | `5000` | Whole-phase budget (ms) for the pooler-health probes. | | `canaryMarkersCheck` | boolean | `true` | Post-boot RUNTIME-MARKERS check: the R8 presence cache (`presence.json`) and the R9 toolset-audit sidecar (`toolset-audit.jsonl`) must exist and be well-formed. Absent files (generic install) skip; a malformed file fails. | `target` semantics (fallback path only — a pending notice or a usable shutdown notice overrides `target` for that boot): - `primary` — the first root agent to start (the main agent). Delivery is guarded so exactly one primary is notified. - `all` — every root agent. - `` — an exact session id, pinned to one specific agent. Full example patch row, restating every key with a custom notice and an explicit `restartUnit`: ```yaml - insert: - id: smart-restart name: dsh-smart-restart config: enabled: true stateDir: .smart-restart target: primary wakeup: true notice: "The DSH service restarted. Please check for interrupted work and report your status in one line." restartUnit: dsh.service toolEnabled: true shutdownGraceMs: 600000 ignoredSessionPrefixes: - head- canary: true canaryTimeoutMs: 45000 canaryPort: 0 canaryProfile: '' canaryBinary: '' canaryStateDirOverrides: deepartments: '' ``` > **Single delivery channel.** When `wakeup` is enabled the notice is delivered via `agent.followup()`; when disabled, via `agent.inject()`. It is **never** both with the same message — a followup that queues into the inbox and a parallel inject of the same message would collide with Inbox's "already pending" validation. ## Behavior & lifecycle - **When woken**, the main agent receives a plugin-source user message (`source.kind: 'plugin'`, `form: 'notice'`) and typically acknowledges with a one-liner or resumes any interrupted task. - **On success**, the plugin logs `[smart-restart] notice delivered to `; when a pending notice is pinned on boot it logs `[smart-restart] pinned restart notice to session `, and when a smart shutdown notice is pinned it logs `[smart-restart] pinned restart notice to last-active session `. Interrupted heads are logged too — `[smart-restart] interrupted sessions at shutdown: …`, `[smart-restart] interrupted heads (resume recipients): …` and `[smart-restart] resume notice delivered to interrupted head ` — all observable boot evidence in the journal. - **Once per boot** — the startup-event delivery and the bounded poll cannot both fire, so a restart produces exactly one notice. - **Pinning priority** — (1) a tool-caller `pending-notice.json` wins; (2) a smart shutdown `shutdown-notice.json` pins to the last-active session when it was active within `shutdownGraceMs`; (3) otherwise `target` decides. - **Reversible lifecycle** — the event listeners, poll timer, tool registration, and `SIGTERM`/`SIGINT` handlers are reversible via `ctx.effect` (dropped on plugin unload / HMR). The only intentional exceptions are the marker and the `pending-notice.json`/`shutdown-notice.json` files, which must survive the restart they document. ## Limitations Be honest about what this plugin does not do: - **Agent-side awareness only.** There is no desktop or browser toast — DSH currently has no notification service, so the notice surfaces only in the agent's own context (visible in the GUI session, not as an OS/browser notification). - **Not for the one-shot headless CLI.** A boot-time wake may not exit cleanly in a single-shot headless run; this plugin targets long-lived GUI instances. The notice is still delivered and committed, but for headless one-shots it is of little use. - **Per-DSH-home marker.** The marker lives under a single ``, so separate homes (e.g. your stable vs. dev instance) are tracked independently — a restart of one does not notify agents in the other. - **Tool requires a systemd unit.** Since **v0.5.0** the `smart_restart` tool **auto-detects the plugin's own unit from `/proc/self/cgroup`** on systemd-managed installs — no `restartUnit` config needed. On a host where no unit is detectable (a bare non-systemd process), the tool still fails safe: set `restartUnit` explicitly to override the auto-detected unit or to enable the tool there. Non-tool restarts are still auto-detected when an agent was active within `shutdownGraceMs`, otherwise they fall back to `target`. - **Canary adds latency.** A canary-gated call blocks the tool for up to `canaryTimeoutMs` (default 45s) while the ephemeral instance boots and is probed; disable the canary (or lower the timeout) for fast, low-risk restarts. The canary validates config/compose + boot health, not the `systemctl restart` command itself (the restart remains fire-and-forget). - **Existing sessions may lack the tool.** A session whose toolset was created **before** the plugin was installed won't have `smart_restart` — start a new chat after installing to pick it up. - **rc-era API.** The plugin targets DSH `0.1.0-rc.7+` / `0.1.1-rc.x`; pre-1.0 APIs (events, session ids, message forms) may change in later releases. ## Development ``` src/ index.ts — apply() wiring: marker + pending/shutdown-notice I/O, restart detection, activity tracking + SIGTERM/SIGINT hook, smart_restart tool (incl. the read-before-edit active-agent guard (fb-168), the optional canary gate + abort alert, restartUnit auto-detection — resolveRestartUnit reads /proc/self/cgroup when config empty), delivery (followup/inject), interrupted-head resume notices (fb-168) boot.ts — pure, deterministic restart + notice logic (I/O-free, unit-testable), incl. parseShutdownNotice / shutdownTarget / activeAgentGuard / guardRefusalMessage / waitForIdle / interruptedHeads canary.ts — optional canary pre-restart validation: ExecStart derivation, temp patch build, free-port pick, dump-config pre-flight, boot + liveness probe (all IO injectable via CanaryHooks) test/ marker.test.js — detectRestart / parsePendingNotice / parseShutdownNotice / shutdownTarget / activeAgentGuard / interruptedHeads / parseCgroupUnit / selectsAgent / targetsAgent / compiled exports notice.test.js — buildNotice / humanizeDowntime canary.test.js — deriveExecStartParams / buildPatchContent / pickFreePort / probeStatusHealthy / resolveExecTarget / runCanary with injected hooks wait.test.js — waitForIdle / guardRefusalMessage (scripted registry + virtual clock) + the smart_restart wait-path harness (fixtures/smart-restart-wait-topology.js) ``` - `pnpm install` — install dependencies. - `pnpm build` — compile `src/` to `lib/` with `tsc`. - `pnpm test` — run the unit tests in `test/` (`node:test`) against the built `lib/`. The unit tests cover **pure logic only** (restart detection, pending-notice parsing, targeting, humanized downtime, notice building) plus a small check of the compiled plugin exports. A real reboot smoke — install into an isolated development profile, trigger a service restart, and confirm the notice is delivered — is performed against a dev profile, since a true process restart cannot be exercised inside a unit-test process. ## License MIT