# App lifecycle What Maison actually *does* to an app — install, start, update, save, uninstall — and in what order. This document is authoritative for the sequences and their failure semantics. Its two companions: - [`app-model.md`](./app-model.md) — where an app **lives** on disk and how its tile state is derived. Read that first; this document is what happens *to* that layout. - [`x-compose-app.md`](./x-compose-app.md) — the **declaration** of `folders` and `hooks`. This document is what Maison does with them. [`backup.md`](./backup.md) is the **design** for offsite backup and disaster recovery — not yet implemented. It builds on the stop and restart sequences below rather than replacing them; the local archive behaviour documented here is what exists today. --- ## The one rule > **Every `docker compose up` Maison runs goes through `internal/stackup`.** There is no other place in the codebase that starts an app's stack. This is the whole reason the package exists: "create this folder before the stack comes up" has to be true when the app is installed, when it is started from the tile a month later, when a store update recreates it, and when the operator saves a config change. Five call sites, one guarantee. ``` ┌─────────────────────────────────────────┐ install ─────────────┤ │ start (tile) ───────┤ stackup.Up(project, files) │ store update ────────┤ │ save config ────────┤ converge: │ save web-UI ────────┤ folders → secrets → variables │ │ → init(pre_up) → seed → files │ │ → pre_up │ │ → compose up -d --remove-orphans │ │ → init(post_up) → post_up │ └─────────────────────────────────────────┘ ``` **If you add a sixth thing that starts a stack, route it through `stackup.Up`.** Calling `composecmd.Up` directly means the app starts without its directories, and the bug will only show up on someone's second boot. --- ## The up sequence `stackup.Up` is the primitive. It is **idempotent** — it starts a stopped stack, recreates a removed one, and re-applies a changed compose, all with the same call. | Step | What happens | On failure | |---|---|---| | **1. Resolve the spec** | Read `x-compose-app` (`folders`, `secrets`, `variables`, `files`, `init`, `hooks`) from base + override, with `x-casaos` `pre-install-cmd` / `post-install-cmd` as the fallback for the install hooks. Later files win, key by key. | — | | **2. Ensure folders** | Create every folder declared under `folders`; apply user/group/mode; walk the tree when `recursive`. Declared folders are the *only* directories Maison creates — it never infers them from `volumes:`. | **Fatal** — a declared folder is the author's contract. | | **3. Secrets** | Generate each `secrets` value the app's `.env` does not already carry, and write it there. A key already holding a value is **reused, never regenerated**. | **Fatal.** | | **4. Variables** | Render each `variables` template and refresh it in `.env`, so a derived value follows the deployment instead of freezing. | **Fatal.** | | **5. `init` (pre_up)** | Run each one-shot container whose `when:` guard says it is due; bind `capture:` output for the renderers below. | **Fatal** — a stack must not start on a store that was never seeded. | | **6. Seed** | Mirror the app's `.seed` tree into its folder: `.tmpl` rendered, everything else copied, **create-if-absent**. Paths a `files` entry claims are left to it. | **Fatal** — including an unresolved `${VAR}` in a template. | | **7. Files** | Write each `files` entry: `ensure: once` skips an existing file, `ensure: always` re-renders it. | **Fatal.** | | **8. `pre_up`** | Run the hook. | **Fatal** — a precondition that doesn't hold must not start the stack. | | **9. `docker compose up -d --remove-orphans`** | Base + override, with the app's `.env` and interpolation variables. `--remove-orphans` removes the containers of services the file no longer declares — a store update that renames a service would otherwise leave the old container holding its `container_name`, and the new service comes up as `_`. | **Fatal.** | | **10. `init` (post_up)** | Run each `phase: post_up` step — a seeder that needs the app's own network or a running service. | **Logged and swallowed.** | | **11. `post_up`** | Run the hook. | **Logged and swallowed** — the stack is already running; tearing a healthy app back down over a failed after-the-fact tweak is worse than the failed tweak. | Steps 2–7 are the **converge**: everything the app declared, brought into being before anything runs. Every one of them is idempotent, and every one of them is fatal, because the failure they replace was not — a shell hook whose `openssl` was missing wrote an empty secret and exited 0. The asymmetry in 8 vs 11 is deliberate and worth internalising: **pre-hooks gate, post-hooks garnish.** Anything flaky in a `pre_up` blocks the app on *every* start. A hook is also failed when it calls a command outside the set Maison makes available to hooks — even if the script itself exited 0, which is the usual case when the command sat inside a `"$(...)"`. That verdict feeds the same table above, so it gates a `pre_*` and is only logged for a `post_*`. See [`x-compose-app.md`](./x-compose-app.md) § *The command set* for the list and for the sanctioned way to reach the host. --- ## Install `Installer.Install` — the only operation that is not just "an up". It runs the install-only hooks around the ordinary up sequence, because it is the only caller that knows the app is being installed for the *first* time. ``` 1. fetch the app's compose from the store 2. restore backup (only when installing from an archive — see app-model.md) 3. write docker-compose.yml (the store's bytes, unchanged — overwritten on every install/update, never otherwise) 4. write .env (prefilled, and NEVER clobbered if one already exists) 5. write the update reference into the override's x-compose-app 6. copy the store's seed/ tree to .seed (the app folder stands on its own after) 7. ensure folders ← early, because pre_install seeds files into them 8. pull images (Download progress bar, 0 → 100) 9. pre_install hook ← fatal on failure 10. ┌ stackup.Up ─────────────────────────────────────┐ │ converge (again, idempotent): │ (Start progress bar) │ folders → secrets → variables → init │ │ → seed → files │ │ pre_up → compose up -d → init(post_up) → post_up│ └─────────────────────────────────────────────────┘ 11. post_install hook ← logged, not fatal 12. await readiness (poll Docker until running + healthy, ~30s) ``` Folders are ensured **twice** — at step 7 and again inside step 10. That is not redundancy to clean up: step 7 is what makes them exist before the `pre_install` hook and the image pull, and step 10 is what makes them exist for every *later* start, when there is no installer in the picture at all. Step 6 copies, it does not write: the store's `seed/` becomes the app's own `.seed/`, and the files themselves are written by the converge in step 10 — so the same code path seeds a fresh install and re-seeds every later start. Steps 4–5 are the non-destructive contract that makes "install from backup" work without special-casing: the strict base is meant to be replaceable, an existing `.env` is meant to be kept, and app data is never touched. ### Progress The install emits `Event`s on two independent tracks, which the UI shows one at a time on a **single bar**: **Download** (image pull, real per-layer progress, blue) and then **Start** (bringing the stack up, driven by Docker's live running/healthy fractions rather than a guess, green). The bar's colour is what says which step is running — see `appProgress()` in `web/src/lib/stores/apps.ts`, the one place that turns these fields into a bar, for both the tile and the store's install pill. Progress rides the live app list, so the tile keeps advancing after the store panel is closed. A failed install **stays visible** on the tile until it is retried or dismissed. --- ## Start · Stop · Restart | Operation | Managed app (Maison wrote its folder) | Unmanaged app (a stack Maison merely discovered) | |---|---|---| | **Start** | `stackup.Up` — so a fully-down stack whose containers were removed is *recreated*, folders and hooks included. | `docker start` on the existing containers. There are no compose files, so there is nothing to declare. | | **Stop** | `docker stop` on the project's containers. No hooks. The folder stays. | Same. | | **Restart** | `docker restart`. **No hooks, no folders, no compose** — it is a container-level bounce, not an up. | Same. | An app that declares `lifecycle.stoppable: false` is refused **Stop** (`403`; the menu withholds it) — see [`x-compose-app.md`](./x-compose-app.md#lifecycle-what-the-operator-may-do-to-the-app). Start and Restart stay available. Restart deliberately does *not* run the up sequence. If you want folders and `pre_up` re-applied, that is a **Start** (or a config save), not a restart. While any of these run the tile is **busy**: greyed with a `…` overlay and no burger menu (see `app-model.md`). --- ## Update `Installer.ApplyUpdate`, driven by the update reference recorded in the override at install time (`store-ref`) and editable from the Update tab (`Installer.SetUpdateRef`). ``` 1. fetch the store's current compose for the app store-ref names 2. equal to what's on disk, byte for byte? → nothing to do, report "up to date" room for the rollback point on the local disk? no → REFUSE, nothing changed 3. pull the new version's images ← while the old version is still serving 4. back up the app ← the rollback point, taken before anything is written (failed? → REFUSE, nothing changed) 5. stop the old version (not for an app declaring lifecycle.stoppable: false) 6. overwrite docker-compose.yml (the strict base only) 7. refresh .seed from the same store sync as the compose above 8. stackup.Up → converge (folders, secrets, variables, init, seed, files — including anything the new version introduces) → pre_up → up → post_up 9. Up failed? → restore the rollback point, start it, watch it, and report ``` The rollback point (step 4) is **always the local engine**, whatever engine is configured for scheduled backups. A rollback happens in the seconds after an update broke something, so it has to be a rename; restoring from a repository is a download, and the app would be broken for the duration. These are ordinary local archives, so the nightly run prunes them under the local engine's own retention like any other — there is no separate retention for them, and `retention.Plan` never drops the newest, so a rollback point cannot be expired out from under an update. If the rollback point cannot be taken, **the update is refused** and nothing is changed (`installer.ErrNoRollback`; `409` with `no_rollback: true` from `POST /api/apps/{id}/update`, `no_rollback` on the run item). Room is checked first, before the image pull: the same stat-only measurement as the backup dialog (`EstimateBackup` against the local engine: the folder minus its `backup.exclude`, times the copy's headroom, against free disk). A copy that fails anyway is a refusal too. Each refusal raises an `app.update:` warning that says how to go ahead. Going ahead is the **owner's choice, per app**: `{"noBackup": true}` on `POST /api/apps/{id}/update`, or on `POST /api/updates/run` naming exactly one app. The Update tab and the refused row on Settings → Updates offer it as *Update without backup*. It skips the room check and the rollback point, and the response carries a `warning` that the update cannot be undone. "Update all" never asks for it. Why refuse rather than go ahead: the old behaviour updated anyway so the largest apps were never pinned on old versions. But it made the one destructive change Maison makes irreversible without anyone having chosen that, and it ran the copy until the disk was full — failing not just the update but every app still writing to that disk. The explicit per-app option keeps the large apps updatable without either. The copy itself writes through `apps.writebackFile`, which flushes and evicts each 8 MiB as it goes. Page cache is charged to Maison's cgroup, and an unbounded copy of a multi-GB file filled a 512 MiB limit with dirty pages and got Maison OOM-killed mid-update (watch.nsl.sh, 2026-09-30). The old version is **stopped before anything of the new one runs** (step 5). The new version's `init` steps run in `pre_up`, against the app's data, and the old containers still have that data open — taking the rollback point restarts them on its way out. A database that takes an exclusive lock fails every such step: FileBrowser's bolt database answers `timeout` after a second, and the update is rolled back for nothing. The images are pulled first (step 3) so the stop costs the swap rather than the download. A system app is not stopped: stopping the dashboard, or the gateway in front of it, takes down the process doing the update. The restore (step 9) replaces the whole folder, so it takes the old compose with it: the app returns to the state the rollback point captured rather than to a new compose running against old data. A rolled-back update still reports as a **failure** — the app is running the old version, and rendering that as success would be a lie. Put back is not running, so the rollback does two more things. It **starts** the app — `Restore` restarts only an app it found running, and step 5 stopped it — and an app that does not start is a rollback that failed. Then it **watches** it: two looks at its containers 30 seconds apart (`installer.Steady`). A container restarting, one whose restart counter moved, or one that was running at the first look and is not at the second means the previous version is not staying up, whatever the restore said. Every failed update leaves an `app.update:` incident, not just a tile: a warning when the previous version is back and running, **critical** when it is not, when the rollback itself failed, or when there was no rollback point to undo it. See `incidents.md`. The override and `.env` are never touched — that is the entire point of keeping the base byte-identical to the store. `pre_install` / `post_install` do **not** re-run; `pre_up` / `post_up` do, because an update is an up. ### Update all Settings → Updates lists every app by what can be done about it: an update available, a failed check or a failed update (its `app.update:` incident), no store reference (*untracked*), or not Maison's at all (*unmanaged* — a stack it only discovered). The check behind it is `Installer.CheckAll`: the same byte comparison as the Update tab, but each store is synced **once** however many apps follow it, and a store that cannot be reached marks its apps as failed, never as up to date. A check runs ten minutes after boot and daily at 04:00, after the store's own refresh. **It only ever checks** — nothing is updated without somebody pressing a button. **Update all** (`POST /api/updates/run`) queues every app with an update available and runs the queue on a background context (`server/updates.go`), so closing the page does not stop it. Each app goes through exactly the sequence above, under the tile's busy overlay. Three rules: - **Strictly sequential**, like the nightly backup: every update takes a local rollback point, a full copy of the app, and several at once is how a data disk fills. The confirmation dialog asks `GET /api/updates/preflight` first, which names the apps whose rollback point will not fit counting the ones taken before them — those will be refused, untouched, and can then be updated one at a time without a backup. - **Refusals say why.** A refused app carries a reason — `no_room` (with the bytes needed and free), `backup_timeout` or `backup_failed` — on its run item, on the 409 of the per-app endpoint and in its `app.update` incident's args; rollback incidents carry `rolled_back`, `not_running`, `rollback_failed` or `broken`. Settings → Updates turns the code into a sentence and keeps the raw error behind *Details*. Every app appears in exactly one row; confirmations open in the row that asked, and a plain single-app update asks nothing (it is backed up and rolled back on failure). - **A failure does not stop the run.** The failed app has already been put back and has raised its incident; the next app's update has nothing to do with it. - **Apps declaring `lifecycle.stoppable: false` are left out.** Updating the dashboard, or the gateway in front of it, can take down the process running the queue. Such an app is updated on its own, by naming it alone; a run naming it among others is refused. An untracked app is offered a **suggestion**, from two kinds of evidence: the project name an install of a store app would have created, and the image of the app's main service (never a sidecar — store apps share those). Both together wins, then the image alone, then the name alone, and nothing when that still leaves a choice. Accepting it is an ordinary `PUT /api/apps/{id}/update/ref` behind a confirmation, one app at a time: linking an app to the wrong store app replaces it in place on the next update (see `app-model.md`). --- ## Save config / Save web UI `SetConfig` writes the override (after validating that it parses — a typo must not leave an app whose only repair path is the config window that broke it), then `stackup.Up`. `SetWebUI` merges the `webui-*` keys into the override and does the same. So **saving a config re-runs the up sequence**, hooks and folders included. An override that adds a `folders` entry gets its directory created on save, not on the next restart. `SetTips` is the exception: tips never affect the running container, so saving them writes the override and stops there. No Docker call at all. --- ## Uninstall **Back the app up through the default backup engine, then remove it.** Nothing is deleted, and no hooks run — Maison has no `pre_uninstall` / `post_uninstall`, on purpose: a hook that fires while the app is being taken away is a hook that can fail and leave the operator unable to uninstall. The backup is the safety net instead. See `backup.md` §Uninstalling an app for the engine seam, and `app-model.md` for the archive format and the restore path. ``` 1. stop containers (Backup progress bar — stopped, not removed, so every failure path below can simply start the app again) 2. back the app up (Backup progress bar — a rename on the local engine, an upload on a remote one; the long step) 3. finalise (Archive progress bar — the commit point; instant unless it is a zip, which is metered by bytes) 4. remove containers (Remove progress bar — one tick per container, because a and the folder single stop can block for the whole stop-grace period) ``` **The order is the contract.** Nothing is destroyed until step 3 returns, so a repository that cannot be reached fails the uninstall and leaves the app installed and running rather than leaving its data nowhere. Which also means an uninstall on a box backing up offsite puts that app's data offsite — it did not use to, and the settings page said it did. On the local engine the backup is still a single rename of the app folder into `maison/.backups//`, so an uninstall stays instant and free whatever the app's size (`SnapshotOpts.Consume`). `zip` is a local-engine option only, and the dialog hides it against a remote engine, where a zip would defeat deduplication. ### Detached, like an install `DELETE /api/apps/{id}` **starts** the uninstall and returns `202 Accepted`; only an up-front refusal (an app declaring `lifecycle.uninstallable: false` → `403`) is answered synchronously. The work runs on a background context, so it survives the request, and the confirmation dialog closes at once instead of blocking the dashboard on a zip that can take minutes. Progress rides the live app list exactly the way an install's does — the tracker lives in `apps.Registry` (`StartUninstall` / `Uninstalls` / `ClearUninstall`), and `server.overlayUninstalls` stamps it onto the app's tile. The tile renders the *same single bar as an install, in red*: **Backup**, then **Archive**, then **Remove**. The bar is keyed on the phase rather than on which counter is still moving — an uninstall now opens with a step that can run for minutes, and reporting that as "Removing" would name the one thing that has definitely not happened yet. There is never a placeholder tile to append (unlike an install): the folder is what makes the tile, and it only disappears at the last step. A failed uninstall **stays visible** as a red `!` on the tile, with the error as its tooltip, until it is retried or dismissed (`POST /api/apps/{id}/dismiss`, which also clears a failed install or backup). Every failure path leaves the app folder in place, so there is always a tile for the error to land on. --- ## Backup Back up the app folder **without** uninstalling. The only operation that stops a running app on purpose, so the sequence exists to keep that window short: ``` 1. pass 1 the engine captures AppData/ while the app is still up — costs no downtime, and warms the engine's incremental state 2. stop (skipped when the app was already stopped) 3. pass 2 the engine captures it again: only what changed during pass 1. This is the entire downtime, and it is bounded by a timeout 4. commit the engine makes the backup real — this is the commit point 5. start deferred, so it runs even if a later step fails ``` The two passes are the registry's, not the engine's — along with the per-app lock, the deferred restart, and the tracked progress. **An engine owns exactly one thing: getting bytes to durable storage and back** (`internal/apps.Provider`). That boundary is why no engine can lengthen an app's downtime by restructuring the sequence, or produce an inconsistent snapshot by choosing when to read. What each pass *does* depends on the engine: | | `local` (built in, always available) | `kopia` (and any later remote engine) | |---|---|---| | A pass | mirrors into `maison/.backups//.staging-` | snapshots `AppData/` straight into the repository | | Commit | renames staging → ``, or zips it | drops the torn pass-1 snapshot | | Needs free disk | yes, a full second copy | **no** | | Survives losing the box | no — same disk as the app | yes | The local mirror is plain Go (`apps.mirror`), not rsync: the runtime image carries no rsync, and the incremental test — same size *and* same mtime — is the one rsync makes by default. Irregular files (sockets, fifos) are skipped rather than opened, so an app that leaves a socket in its folder is still backupable. **The no-staging shape is what makes a large app backupable at all.** A 300 GB app on a 400 GB disk needs 330 GB free for the local engine's copy, so `Estimate.Enough` is false and the backup is refused — that is current behaviour, not a hypothetical. An engine that streams to a repository has no such requirement, and `EstimateBackup` skips the guard entirely for one (`Estimate.Streamed`). What it costs: **downtime is no longer engine-independent.** A hung repository would extend an outage rather than merely failing a backup. Two things bound that — the restart is deferred, so a failure anywhere after the stop still brings the app up, and the stopped window has a timeout (`Registry.StoppedPassTimeout`, 15 minutes by default) after which the engine's container is killed *and removed*. Killing the `docker` client alone would leave the engine running and still holding the app's files, which is why `internal/engine` removes the container by name. `POST /api/apps/{id}/backup?zip=` **starts** it and returns `202 Accepted`. Only the up-front refusals are synchronous: an unknown app, or — for an engine that needs local space — not enough of it. Progress rides the live app list through `apps.Registry.StartBackup` / `Backups` / `ClearBackup` and `server.overlayBackups`, on the same single tile bar as an install or an uninstall, in amber: **Copy**, then **Sync**, then **Compress**. Nothing is listable until the commit, so a crash mid-backup leaves only a `.staging-…` folder, a `.partial` zip, or a snapshot tagged as the throwaway first pass — none of which is ever offered for restore. ## Restore Which of three paths a restore takes depends on **where the backup is** and whether there is room — never on which engine is currently selected. That last part is the rule that keeps a user's older backups reachable after they switch engines. ``` on disk 1. stop (if running) 2. archive rename AppData/ → the archive tree ← instant, free 3. restore folder archive renamed back; zip extracted 4. start deferred remote, room 0. fetch the engine downloads it into the archive tree … then exactly the above remote, no room 1. stop 2. undo the engine snapshots the current state — and if that fails the restore is REFUSED 3. restore written over the live folder, deleting files the backup does not have 4. start deferred, and refused while the marker below exists ``` The first two are atomic at their commit point and reversible by a rename. **The third is neither.** An interruption leaves the folder holding neither the old state nor the new one, and the only way back is a remote snapshot — so a restore is reversible only while the repository is reachable. That is the price of restoring an app too large to hold two copies of, and it is why the undo snapshot is mandatory rather than best-effort. While an in-place restore is running, `//.restoring` exists. It lives *outside* the folder being written (a delete-extra restore would remove it from inside) — which is also why the one app whose folder *contains* the archive tree cannot be restored at all; see [`app-model.md`](./app-model.md) and its name cannot parse as a stamp, so no lister mistakes it for an archive. **It gates `EnsureStarted`:** an app whose restore was cut short is not started, because it would initialise over the gap — fresh database, default config — and that invented state would become the next backup. `POST /api/apps/{id}/restore {"name": …}` for a live app, or `POST /api/backups/{app}/restore` for one that has been uninstalled — the same detached path, which simply finds nothing to do at the stop, archive and start steps. ## Scheduled backups A nightly run (`internal/backup.Scheduler`) backs up every app and, if the engine can, the user-data set — everything under the data root that is **not** `AppData`. It is Maison's own scheduler and cannot be delegated to the engine's at any price: a consistent app snapshot needs containers stopped, which no backup tool can do. Four properties that are not obvious from "run it daily": - **Strictly sequential.** The per-app lock protects one app; nothing else would stop a run taking six down at once. - **Skip, don't queue.** A run still going at the next window is skipped — waiting behind itself only compounds the delay. - **Jitter.** A fleet all firing at 03:30 is a thundering herd against one bucket. The offset is derived from the data path, so it is stable per box. - **Two kinds of app are never targets:** Maison's own state directory (stopping it kills the process running the backup), and any app declaring `backup.skip`. The grid an app sits in plays no part — a `view: system` app without `skip` is backed up like any other, which means it is stopped for the stopped pass. The platform stacks declare `skip` for that reason; backing them up *without* stopping them is a shape the app path does not have yet. --- ## What a hook sees Hooks run through `/bin/bash -c` **inside the Maison container**, with the working directory set to the app's folder, but they act on the **host** Docker daemon. | Variable | Value | |---|---| | `AppID` | The compose project name. | | `APP_DIR` | The app's directory, as the **host** sees it. | | `DOCKER_HOST` | `unix:///var/run/docker.sock` — the host daemon. | | `DATA_ROOT`, `PUID`, `PGID`, `TZ`, `REF_*`, … | The app's base interpolation variables. | | everything in the app's `.env` | So a hook sees the same values its compose does. | Because they target the host daemon, `/DATA` and `${DATA_ROOT}` **inside a hook's script text** are rewritten to host paths — a `docker run -v` in a hook must name a path the host daemon can resolve. That rewrite is also the one trap worth knowing: > A hook that just wants a directory to exist must **not** `mkdir` it. The path it > writes is a *host* path, but the `mkdir` runs *in the Maison container* — > creating the wrong directory in the wrong place, and leaving the app with an empty > mount. **Declare it under `folders`** instead: those are created through Maison's > data mount and are correct on both sides of the socket. Hooks are for Docker-level work (pulling a sidecar image, priming a volume with `docker run`, poking another stack). Directories are what `folders` is for. ### Paths, in code Two mappings, easy to confuse, both in `internal/envinject`: | Function | Use on | Does | |---|---|---| | `ContainerPath` | a **real path** | host spelling (`/DATA/…`, `${DATA_ROOT}/…`, the literal host path) → this container's data mount. This is how folders get created. | | `HostPath` | a **real path** | the inverse: container path → host path. This is how `APP_DIR` is built. | | `RewriteToHostPath` | **script text** | rewrites the `/DATA` / `${DATA_ROOT}` *spellings* a hook author wrote. Single-pass — a host path normally ends in `/DATA`, so rewriting twice would yield `/opt/maison/opt/maison/DATA`. | --- ## Failure semantics, in one table | Step | Fails the operation? | Why | |---|---|---| | Declared `folders` entry (bad path, unknown user, unquoted mode) | **Yes** | It is the author's explicit contract, and starting without it gives the app an unwritable mount — a confusing "permission denied" instead of a clear error. | | `pre_install`, `pre_up` | **Yes** | Preconditions. | | `post_install`, `post_up` | No (logged) | The stack is already up. | | `docker compose up` | **Yes** | Obviously. | | Ownership / mode (`chown`, `chmod`) on a folder | No (logged) | Not every filesystem supports it, and that should not block an otherwise healthy start. | | Free-space check before a backup | **Yes**, before anything is copied | Filling the data disk breaks every app on the box, not just this one. Skipped entirely for an engine that streams to a repository — it needs no room. | | Either pass of a backup | **Yes** | A backup that is missing files is not a backup, and silently keeping it would be worse than failing. | | The stopped pass exceeding its timeout | **Yes**, and the engine's container is killed *and removed* | Otherwise a hung repository is an outage rather than a failed backup — and killing only the `docker` client would leave the engine running, still holding the app's files, while Maison restarts the app and reports success. | | Committing a backup | **Yes** | Nothing is listable before the commit, so a failure leaves a `.staging-…`, a `.partial`, or a snapshot tagged as the throwaway first pass — none of which is ever offered for restore. | | Dropping the throwaway first-pass snapshot | No (logged) | The real backup already exists; refusing to commit it because a cleanup failed is the worse outcome. The orphan is invisible to listing and is swept later. | | Restarting the app after a backup or restore | No (logged) | The restart is deferred so it runs even when the operation failed — an app left down is a worse outcome than a missing backup. | | Restarting an app whose in-place restore was interrupted | **Yes**, it stays down | It holds neither the old state nor the new one. Starting it would initialise over the gap and that invented state would become the next backup. | | The undo snapshot before an in-place restore | **Yes**, the restore is refused | An in-place restore is not atomic and has no local undo. An unrecoverable overwrite is worse than a restore that did not happen. | | One target of a scheduled run | No | The other targets still run. A broken app must not cost the user every other backup that night; the failures are collected into one summary. | | Sending a failure notification | No (logged) | A broken SMTP configuration must never turn a successful backup into a failed one. | An install that fails leaves the app's folder **in place**, half-configured — which is correct: the folder is the tile, the failure is visible on it, and a retry is a plain re-install over what is already there.