# Backup and recovery > **Status: mostly implemented.** The engine seam, the `local` and `kopia` engines, > the two-pass app backup, the user-data set, all three restore paths, the nightly > schedule, retention and failure notifications are in the tree and tested. Listing, > restore and delete all read through the engine set, so a remote backup is a > first-class one in the UI (§[One surface, one mechanism](#one-surface-one-mechanism)). > > The user-data set is listable and restorable from the Backups page — both modes, with > the guards described in [Restore](#the-user-data-set). > > The adapter boundary ([The engine adapter](#the-engine-adapter)) is built and is how > every engine but the local one now arrives. **Maison no longer contains an engine.** > `internal/backup/kopia` is gone, and with it the last line of this binary that knew what > kopia is: a box's engines are exactly the adapter descriptors the host side wrote, and an > engine the configuration names and the box does not have is reported as > `backup.missing:` rather than quietly demoted (`server.checkBackupDestinations`). > The host-side pair that provisions a repository — `ensure-backup-credentials.sh` and > `ensure-backup-config.sh` in `template-root` — is described under > [What Maison consumes from the PCS](#what-maison-consumes-from-the-pcs); a > hand-connected repository against MinIO or a filesystem path remains how this is > developed and tested. > > **Not yet built:** disaster recovery / recovery mode ([below](#disaster-recovery)). Maison > also still falls back to a one-shot container when the resident engine is unreachable, > instead of reporting it. > > Two things changed during implementation and are corrected in place below: there is > **no local staging copy** on the remote path (§[Why there is no local staging > copy](#why-there-is-no-local-staging-copy)), and restoring an app too large to hold > two copies of writes **in place**, which is not atomic (§[Restore](#restore)). > > A third correction is larger, and reverses what this document used to say: an > uninstall now backs the app up **through the default engine** instead of renaming its > folder aside and calling that an archive (§[Uninstalling an > app](#uninstalling-an-app)). The store's install-from-backup picker reads through the > engine set with it, so a backup that exists only in a repository can be reinstalled > from. Its companions: - [`app-model.md`](./app-model.md) — where an app **lives** on disk. Backup is defined entirely in terms of that layout; read it first. - [`lifecycle.md`](./lifecycle.md) — install / start / update / uninstall. Backup hangs off the stop and restart sequences documented there, and does not invent its own. The server-side half — how a PCS is issued storage credentials in the first place — is out of scope here and is treated as a contract Maison consumes. See [What Maison consumes from the PCS](#what-maison-consumes-from-the-pcs). --- ## The two sets Maison backs up **two independent things**, and almost every design consequence in this document follows from them being different. | | Apps | User data | |---|---|---| | Source | `${DATA_ROOT}/AppData/`, one source per app | `${DATA_ROOT}` minus `AppData/` | | Contents | compose files, `.env`, the app's own data | Documents, Downloads, Media, whatever else the user drops at the data root | | Consistency | containers stopped — a real point-in-time snapshot | none; a live filesystem, read as-is | | Orchestration | stop → snapshot → start, per app | none — nothing to stop | | Granularity | restore one app without touching others | one set, or named top-level folders | | Restore | rename / materialise / in place, per app | copy into a new folder, or in place entry by entry | An app that declares `backup.skip` is in **neither** set: it is not a target for the app set, and `AppData/` is excluded from the user-data one. That is what "not backed up" means here, and it is deliberate for the app it was written for — see [Skipping an app outright](#skipping-an-app-outright). **User data is the simpler path and the right thing to build first.** Nothing is stopped, no staging copy exists, and its source path never changes. Every piece of the provider layer can be exercised against it before app orchestration enters the picture. **The user-data set is remote-only.** The local engine deliberately does not implement it and must not: its archives live under `AppData/maison/.backups/`, on the same disk the set would be copied from, so the copy protects against nothing while doubling the space used. The scheduler therefore does not offer it as a target for an engine that cannot do it — otherwise a default install (local engine, user data on) reports a failed backup every night and mails its owner about it while nothing is wrong — and the page says why rather than showing an empty list. **A database must never live under the user-data set.** It has no consistency guarantee: a file being written while the engine reads it is captured mid-write. That is normal and accepted for documents and media, and it is precisely why apps get the stop treatment and user data does not. --- ## Where things live on disk ``` ${DATA_ROOT}/ AppData/ / an app — the backup source for the app set /engine/ the engine's own state (Config.EngineDir) repository.config endpoint, region, bucket, prefix repository.password 0600, generated on this PCS cache/ engine cache — never backed up logs/ never backed up maison/ Maison's own state (Config.StateDir) .backups/// local archives (Config.BackupsDir) AppDataShared/ BEING RETIRED — the same tree, at backup// Documents/ Downloads/ Media/ … the user-data set ``` > **The engine directory is mid-move.** It is going from > `AppDataShared/backup//` into the engine's own app folder, so that the > engine is an app like any other and `AppDataShared` can be deleted outright. Maison > only ever *reads* that directory — the host-side `ensure-backup-config.sh` owns > writing it — so the two halves ship as separate releases and this build reads > **either** layout, preferring the app folder when both exist > (`Config.BackupEngineDir`, `adapter.Discover`). The resolution is done per call, not > once at start, so a box flips over the moment the host script moves the files. > > The engine's app folder is then kept out of the backup by its own `backup.skip` > declaration rather than by any of the machinery below — see > [Skipping an app outright](#skipping-an-app-outright). The property being given up > is the one the next paragraph describes, and it is smaller than it looks: reading > any backup at all requires the repository password, so a box that could use the > carried configuration has already recovered without it. Everything from here to the end of this section describes the layout being retired. `AppDataShared/` is **deliberately outside `AppData/`**, which means it falls inside the user-data set and therefore gets backed up. That is the point: on a box running more than one engine, each engine's backup carries the other's configuration, so recovering either one returns the rest. The operator has to have kept **one** set of credentials, not one per engine. Two exclusions make that safe, and they are matched **by pattern** (`**/cache/`, `**/logs/`) rather than by a fixed list — otherwise the third engine someone adds later silently ships its cache offsite forever: - **`cache/` must never be backed up.** It is multi-gigabyte, it turns over between runs, and it dedups poorly because its contents are already compressed and encrypted. Backing it up inflates snapshot time and inflates metered usage against the storage quota, nightly, for data that is rebuilt on demand. - **`logs/`** — same churn, no value. `repository.password` riding along inside the backup is **harmless**, and worth stating because it looks alarming: reading it requires the password already. It is not a leak, it is merely useless — the recovery path runs on the emailed copy, never on the repo. `repository.config` riding along is mildly useful; a restored box gets its endpoint, bucket and prefix back without re-deriving them. The archive tree needs no exclusion. It lives inside `AppData/` — two levels down, in Maison's own folder — and is therefore already outside the user-data set. It **is** excluded from the *app* set, and there it is not free: `AppData/maison` is an app folder like any other, so backing it up would copy the archive tree into a staging directory inside itself. `Registry.exclusionsFor` adds the exclusion for whichever app holds the tree, and restoring or uninstalling that app is refused outright — a restore renames the folder into its own subtree, and the in-place path deletes what the backup does not have, which here is every archive on the box. --- ## The engine is pluggable The engine is a **user-facing choice** — kopia first, then restic, then custom — set globally rather than per app. The seam is the deliverable; the first engine is just the first tenant of it. ```go type Provider interface { ID() string // "local" | "kopia" | "restic" | … Caps() Caps // LocalSpace, Instant, Retention, Offsite Snapshot(ctx, app, stamp string, emit func(Event)) (Ref, error) List(ctx, app string) ([]Backup, error) ListAll(ctx) (map[string][]Backup, error) // every app, in one call Materialize(ctx, app, stamp string, emit func(Event)) error Delete(ctx, app, stamp string) error } ``` That sketch is illustrative and has drifted — `internal/apps/provider.go` is authoritative, and carries `Commit`, `Abort`, `RestoreInPlace` and the `Caps` fields that grew since. How an engine is *delivered* against this interface is [The engine adapter](#the-engine-adapter). `ListAll` is the bulk shape of `List` and exists for a cost reason, not a convenience one: for a remote engine every call is a subprocess against the repository, and the global page shows every app at once. Asking "which apps do you have" and then listing each would make that page cost one subprocess per installed app. It also cannot be answered from what is installed — on a rebuilt box the repository is the only thing that still knows the apps existed, which is exactly when the page matters. ### What the registry owns, and what a provider owns The registry keeps everything that is not storage: the per-app `enter`/`leave` lock, stop → snapshot → **deferred** restart, the two-pass structure, tracked `BackupState`, idempotency, the sticky error on the tile, and archiving a live folder before restoring over it. **A provider owns exactly one thing: getting bytes to durable storage and back.** That boundary is what keeps a plugin from being able to extend an app's downtime or produce an inconsistent snapshot. ### Progress: engines report, Maison derives > **An engine reports what it observed. Everything computed from that is computed in > one place, the same way, for every engine.** A provider fills in as much of `apps.Event` as it happens to know — a percentage, byte counts, a message, or nothing but a message. It is never asked to declare which of those it can do, because the answer varies *within* one operation rather than between engines: kopia reports nothing but a message while it is still estimating the tree, then bytes and a percentage once it knows. `apps.Tracker` turns that stream into a transfer rate and a time remaining: | The engine reports | What the user gets | |---|---| | Bytes done and expected | Bar, rate, ETA | | A percentage only | Bar, ETA (from elapsed × (100−p)/p) | | Neither | Indeterminate bar, elapsed time | The percentage-only row is the one that makes this generic rather than kopia-shaped: a new engine that can only say how far along it is still gets a usable estimate, and the local engine — which knows its total exactly, because `mirror` walks the tree first — has the most trustworthy numbers on the box. **An engine's own ETA is deliberately not used**, though kopia prints one. It is engine-specific (the next engine's would have different semantics and different smoothing, so the number under the bar would mean something different depending on where the backup was going) and it is not available uniformly (kopia reports none while estimating; the local engine reports none ever), so the fallback has to exist regardless. Parsing an engine's output stops at the adapter that ships with it — nothing downstream of `apps.Event`, and nothing in Maison at all, knows which engine it is talking to. Three rules the tracker exists to enforce, all of them about not lying: - **Estimates are per phase, and reset at every boundary.** The tracks are cumulative on the wire, so an estimate carried from the live pass into the stopped pass would describe work that finished a minute ago. The reset also produces the most useful number on the screen for free: the stopped pass's ETA is *how much longer the app is down*. - **Nothing is shown until it is worth showing** — five seconds of samples and at least 2% done. An ETA offered in the first second is always wrong, and being wrong early is how a progress indicator teaches people to ignore it. - **A total is allowed to move.** An engine discovering the tree as it walks it revises its estimate upward, so a count can go backwards relative to it. That is normal, not a fault. The whole-box run reports through the same path: each target's progress is mirrored into `backup.RunState` from the very events that drive the app's tile, so the two can never disagree about what an app is doing. Targets are enumerated *before* the first one starts, which is what lets the page say "3 of 9" and show what is still to come. A target the run deliberately did not attempt is **skipped**, not failed — the user-data set while a restore is rewriting it, or an app somebody is already backing up by hand. The distinction is not cosmetic: a failure mails the operator "backups are failing on your server", and sending that for a box where the right thing happened is how a useful alert becomes a filter rule. ### The rule that survives a provider switch > **The active provider governs writes only.** Listing is the **union** across all known providers, and restore and delete dispatch on *where the backup actually is* — not on which provider is currently selected. Without this, switching from kopia to restic orphans every kopia snapshot and every local zip. Two consequences: - **`ID()` values are permanent once shipped.** They are how an existing backup finds its way home. - **A provider removed from the picker stays registered read-only.** You can stop offering an engine; you cannot stop reading what it wrote. ### One surface, one mechanism > **There is no such thing as a "cloud backup" feature.** There is backup, and the > engine is a setting it reads. This is a UI rule as much as a code one, and it was broken first in the UI: the engine, schedule and retention lived on a *Cloud backup* settings page while a separate *Backups* page listed archives. Two pages, two URLs one keystroke apart (`/settings/backup`, `/settings/backups`), two API prefixes likewise (`/api/backup`, `/api/backups`) — and the implication that backing up to a repository is a different feature from backing up to disk. It is not: everything that triggers a backup calls `Registry.engine()` and knows nothing else about where the bytes go. They are now one page — destination and schedule above, every backup below — and the word "cloud" does not appear in it. `local` and `kopia` are two values of one setting, and every backup is listed under the engine that holds it. What made the split more than cosmetic is that the two halves had drifted apart in the code. Writes went through the engine set; **reads did not.** Both listing endpoints walked the local archive tree directly, and `StartRestore` gated on a file being there before dispatching to the engine-aware path underneath. So on a box configured for kopia the Backups page was empty while the repository held everything, and a remote-only backup could not be restored from the UI at all — the code to do it was reachable only from the scheduler. Every read now goes through the set: | Path | Reads through | |---|---| | Per-app tab, global page | `Set.List`, `Set.ListAll` | | Store's install-from-backup picker | `Set.ListIn` — a group per engine, like the page's tabs | | Restore (precondition *and* execution) | `Set.LocateIn` — the engine the row belongs to | | Delete | `Set.Delete` — from **that one engine**, never across them | The store's install-from-backup path was the last one to be moved, and it is worth naming because it looked fixed for longer than it was: the Backups page had been converted while the store still called `apps.ListBackups` on the data disk. On a box that had always been remote its picker was simply empty, which reads as "this app has never been backed up" rather than as a bug. `Set.Locate` — the engine-less form the picker used to justify — is now the compatibility path for a request that names no engine, and nothing in the UI produces one. **One write deliberately stays local**, and is an exception rather than an oversight: the update rollback point (`BackupWith`). It needs a rename it can undo in seconds, and a repository upload is not that. The local engine's own retention prunes it like any other local archive — see [Retention](#retention). Archive-on-uninstall used to be listed here as the second exception. It is not one, and treating it as one was the mistake this document made: an uninstall archive that only exists on the box is a recycle bin, not a backup. It dies with the disk — the failure the offsite engine exists for — and the user who uninstalled an app to reclaim space is exactly the user who then deletes it. See [Uninstalling an app](#uninstalling-an-app). --- ## The engine adapter The seam above is a Go interface, and for the first engine that was enough: a `kopia.Provider` inside Maison translated `Provider` calls into kopia's argv and parsed its JSON back. It worked, and it was also the reason adding restic meant editing Maison, releasing Maison, and shipping a new Maison image to every box — for a change that touched no Maison behaviour at all. So the *interface* stayed where it is and the *implementation* moved out of the binary. An engine is delivered as an **adapter image**: the engine's own binary plus a small `maison-engine` CLI that speaks a fixed protocol. Maison ships one generic provider that speaks that protocol and knows nothing about any engine. The compiled kopia provider survived one release beyond that as a fallback, registered whenever no descriptor had claimed the name, so that a box whose host side had not caught up kept backing up. It has been removed. What replaced it is not another fallback but an alarm: Maison cannot invent an engine — naming the image to run is the deployment's decision, which is the whole reason an adapter is not a remote-execution surface — so a box configured for an engine it does not have raises `backup.missing:` and says what is wrong, while the nightly run lands on the local disk or fails outright. The protocol is specified in its own repository, alongside the first adapter — `maison-kopia-engine`, `docs/protocol.md`. This section is the part that belongs here: why the boundary is shaped the way it is, and what Maison keeps on its side of it. ### The transport is argv and stdio, not HTTP `docker exec maison-engine …`, NDJSON on stdout, diagnostics on stderr, a documented exit code. A glue *service* with an HTTP API is the obvious shape and the wrong one, for four reasons that are all about failure rather than elegance: - **`--network none` stays available.** A repository on a local filesystem needs no network and must not have one. An HTTP daemon needs one by definition. - **Cancellation keeps working.** `engine.Runner` holds one pid file per in-flight exec (the `/run/maison-ops` tmpfs) precisely so a cancelled backup cannot leave the engine holding an app's files open while Maison restarts it. Cancelling an HTTP request does not kill a child process. - **Progress needs no new plumbing.** Engines already report by writing lines; `Runner` already streams them into `apps.Event`. - **There is no new client to get wrong** — no health checking, no retry policy, no connection state, no second definition of "the engine is up". Nothing is gained in exchange. The adapter can be written in any language either way, and a verb that runs for an hour streams progress the same way a request would. ### An unreachable engine is an incident, not a second code path **Maison does not fall back.** If the engine container cannot be reached, the operation fails, the run is recorded as failed, and `backup.engine:` is raised for the owner — the same rule [No fallback to local](#no-fallback-to-local) already states for uninstall, applied to transport instead of destination. An engine that quietly does the work somewhere else is worse than one that stops, because the user is told a backup happened. Two things keep that from being noisy: - **Only an engine that receives writes can be at fault.** An engine holding nothing but history is not broken when it is absent; nothing is trying to write to it. - **Retrying the same destination is not falling back.** The nightly self-check can recreate the engine container, so a reachability check waits a bounded interval before declaring failure. Waiting for the destination the user chose is not substituting another. ### The engine is fully optional The local engine is in-process Go and always present. A box with no adapter image, no engine container and no repository is not a degraded box — it is the default FOSS install, and it must raise nothing. Everything above is conditional on an engine having been provisioned and being ticked to receive a trigger. ### The adapter image is deployment-provisioned, never user-supplied The image is named by the host side (`ensure-kopia-stack.sh` and its pin), not by anything a user can type. This is the line that keeps the adapter from being the remote-execution surface described under [Not yet decided](#not-yet-decided): an image Maison execs as root with app data and storage credentials in scope is first-party infrastructure. *Which repository an engine points at* remains an ordinary user setting. It also collapses a pin that is currently duplicated across two repositories with only a comment holding it together — `KOPIA_IMAGE` in `template-root/scripts/library/kopia.sh` and `kopia.DefaultImage` here. An adapter image built `FROM` a pinned engine carries the engine version inside it, and Maison stops needing to know it. ### What stays on Maison's side The split is the one [What the registry owns](#what-the-registry-owns-and-what-a-provider-owns) already draws, and the adapter does not move it: | Maison | The adapter | |---|---| | Which paths are a source (`app-model.md`) | Where bytes go and come back | | Stop → snapshot → deferred restart, the per-app lock, the two-pass structure | Being incremental between the two passes | | What the user asked to keep (`backupconfig.Mode`) | Whether that expiry is sound on this storage (`RetentionModel`) | | Deriving rate, ETA and percentage from reported events | Reporting whatever it observed | | Resolving an app's declared exclusions once | Spelling them in the engine's own syntax | **Maison passes paths; the adapter never derives them.** An adapter that computed a source path from an app name would be a second definition of the disk layout, and the two would disagree at restore time. **Key escrow reads the file, never the engine.** `repository.password` is read from `AppDataShared/backup//` directly, because the moment the key matters most is the moment the box is broken — an escrow path that needs a working engine container is an escrow path that fails exactly when it is needed. The adapter is not consulted. --- ## Backing up an app ### The shape ``` pass 1 app running snapshot AppData/ → S1 bulk upload; torn, throwaway stop app pass 2 app stopped snapshot AppData/ → S2 delta only; consistent start app deferred — always runs delete S1 ``` This is a **direct port of what `Registry.Backup` already does**, with the engine replacing `mirror()`. The existing code runs a live copy pass, stops the app, runs a second pass, and renames the result into place; its own header comment states the property that carries over unchanged: > downtime is proportional to what the app wrote *during* the first pass, not to how > big it is The source path is identical across both passes, so the engine's size+mtime fast path applies on pass 2 and only what changed during pass 1 is re-read and re-uploaded. S2 is taken with the app down, so it is consistent by exactly the argument the current code already makes. **The restart is deferred and must stay that way.** Any failure after the stop still brings the app back up. Leaving an app down is a worse outcome than a missing backup. ### Why there is no local staging copy Because on a large app there is nowhere to put one, and today that means the backup simply does not happen. `Registry.Backup` mirrors the app folder into a staging directory on the same filesystem — a full second copy — and `EstimateBackup` guards it with `folderHeadroom = 1.1`. A 300 GB app therefore needs 330 GB free, so on a 400 GB disk holding it, `Estimate.Enough` is **false and the backup is refused**. That is current behaviour, not a hypothetical. Snapshotting `AppData/` directly removes the requirement entirely rather than raising the ceiling. It also happens to give the stable source path that real incremental backup needs — the most stable path an app has is its own directory. What it costs: - **Downtime is no longer provider-independent.** A hung repository extends an outage instead of merely failing a backup. Guarded by the deferred restart above and by a **pass-2 timeout**, which turns "the repo is hanging" into "the backup failed, the app is up". - **No local archive for that app**, so its restores are always a `Materialize` download. The uninstall path keeps a rename on the local engine — see [Uninstalling an app](#uninstalling-an-app) — and a rename costs nothing at any size. The local tier remains available for apps that fit it — instant restore is worth having where it is free — but nothing may *require* it. `EstimateBackup` already computes the number that decides which mode an app gets. ### The throwaway snapshot must be pruned S1 is torn. Left in the repository it doubles the snapshot count, pollutes the retention set, and — worst — a user browsing snapshots can restore an inconsistent one. **Delete S1 once S2 succeeds.** Content blobs are shared, so nothing S2 needs is lost with it. If pass-1 churn is high (an Immich mid-import), pass 1 can be repeated until the delta stops shrinking before stopping the app. Not needed for a first version. ### Per-app exclusions `AppData/` is the source, but not all of it is worth storing. Apps keep large **regenerable** caches inside their own directory — thumbnails, transcodes, search indexes, model downloads — and without a way to skip them, every one ships nightly and is billed as if it mattered. On a media app the cache can rival the real data. The declaration lives in **`x-compose-app`**, next to `folders` and `hooks`: the app author knows which of their directories are derived, and it travels with the app instead of living in a list Maison has to maintain per store app. ```yaml x-compose-app: backup: exclude: [cache/, "**/thumbs/"] ``` The grammar is a directory relative to the app folder, `**//` for that name at any depth, or the absolute spelling `folders:` uses — and nothing else. It is documented in full in [x-compose-app.md](x-compose-app.md#backup-exclusions); what matters here is *why* it is that small: **two engines have to mean the same thing by it.** `internal/exclude` parses the declaration once, and the registry hands the result down through `SnapshotOpts.Exclude`, so the local engine (which walks the folder) and kopia (which pushes ignore rules into the repository as policy) can each ask for the form they need but neither can invent a third reading. An app whose backup contents depended on the selected engine would only reveal it at restore time. Two rules keep this from becoming a footgun, and both are implemented rather than aspirational: - **Exclusions are a store-app declaration, not a heuristic.** Maison never infers "this looks like a cache" from a directory name — a wrong guess silently omits real data from a backup, which is the worst failure this system can have. A pattern it cannot honour is dropped and reported, so the backup carries *more* than was asked; the failure direction is always the safe one. - **What was excluded is visible.** `Estimate` carries the canonical patterns, their size on disk, and any refusals, so the Backups tab names them before a backup runs and again beside every restore. A restored app missing its cache is working as declared; a restored app missing something the author wrongly marked derived is a bug report — and only naming the paths tells a user which one they have. Three consequences worth stating plainly: - **A restore does not bring an excluded directory back.** The folder is replaced wholesale (and kopia's in-place restore passes `--delete-extra`), so it returns empty and the app refills it. Declaring the directory in `folders` too is what gives it the right owner before the app starts. - **The estimate sizes the copy, not the folder.** The free-space guard reserves for what will actually be written, so an app whose cache is most of its size stops being refused a backup it comfortably fits. - **The local engine's uninstall archive keeps everything.** That path is a folder rename, where an exclusion would cost work rather than save it; a repository engine's uninstall snapshot honours the declaration like any other backup. The local superset is free, so it is not worth an exception to remove. Kopia has no per-run ignore flag, so the rules are policy on the source path, written by `EnsureIgnore` before the **live** pass — never inside the app's downtime — and written *unconditionally*, including for an app that declares nothing. That last part is what retires a rule an author has removed: policy lives in the repository and outlives a Maison reinstall, so a stale one would go on dropping a directory from every backup with nothing on the box to explain why. A user-level override is still worth having for apps whose authors have declared nothing. The merge already works — an override's `x-compose-app` block wins key by key — so what is missing is only an editing surface. ### Skipping an app outright `backup.skip: true` is the whole-app form of the same declaration: there is nothing in this folder worth keeping. The app is not a target for the nightly run (`Scheduler.skip`), a backup of it by hand is refused (`apps.ErrBackupSkipped`, `403`), and the update rollback point is not taken — the update proceeds with no way back, which is the same state as a box whose local engine has no archive yet, and it says so in the log rather than raising the disk-space incident a genuine failure would. **Restore is deliberately not gated**, because a declaration is not retroactive: an app can be marked skipped while backups taken before it still exist, and those stay listable and restorable. Nor is the uninstall archive, which is the owner's decision at that moment rather than the author's. It is a separate declaration from `view: system` on purpose. That field used to decide skipped-by-backup *and* tile grouping *and* refusal-to-stop from one value, so before `backup.skip` the only way for an ordinary app to opt out of backups was to claim to be a platform piece — and even then only the nightly run honoured it, never the manual path behind `POST /api/apps/{id}/backup`. `view` is now only the grid; `backup.skip` is the one backup decision, on every path, and the scheduler consults nothing else. That is not academic for the engine itself. **Backing up kopia means stopping the container taking the snapshot** — the app path is stop → snapshot → start, and `BackupTo` does its own stop regardless of who called it. The refusal therefore lives in `BackupTo`, before the per-app lock and before anything is stopped, that being the one place every caller funnels through. ### Identity is unchanged `` — `YYYY-MM-DD_HHMMSS` — stays the canonical name of a backup. `stampRe` (`archive.go`) is the traversal guard that makes `DeleteBackup` and `resolveBackup` safe, and **it must not be loosened** to accommodate an engine whose native refs look different. A snapshot ID is a provider-internal detail; `(app, stamp)` is the identity the API routes and the frontend already use, and keeping it removes what would otherwise be the largest refactor in this work. --- ## Uninstalling an app **An uninstall is a backup followed by a removal**, written wherever the default engine writes. It is not a special case with its own storage rule: ``` stop containers stopped, not removed downtime starts backup Snapshot(Consume) through the default engine archive Commit — the commit point; a zip compresses here remove the containers, then the app folder if the engine did not take it ``` **Nothing is destroyed before `Commit` returns.** That ordering is the whole design: until then, any failure restarts the app and leaves it installed with its data intact. It is why the containers are *stopped* rather than removed at the top — removing them first would make "put it back" mean re-creating the stack. ### `Consume`: the verb that keeps it free Routing an uninstall through an engine's ordinary `Snapshot` would have regressed it badly. The local engine's snapshot is a full second copy, guarded by `EstimateBackup`'s `folderHeadroom`, so an uninstall through it would **refuse any app larger than half the free disk** — an app you can install and then cannot remove. That is worse than the problem being fixed. So `SnapshotOpts.Consume` tells the engine the source folder is being destroyed and it may *take* the folder rather than read it. The local engine does, and the result is a single atomic rename from live app to finished archive — exactly what the old uninstall did, now expressed through the seam instead of beside it. An engine streaming to a repository ignores it, reads as usual, and the registry removes the source once `Commit` has succeeded. Either way the caller's guarantee is the same: after a committed `Consume` backup, the app folder is gone. Two rules make it safe, and neither is optional: - **An engine that takes the folder must do it in `Commit`, never in `Snapshot`.** Nothing before the commit point is durable — an interrupted backup is aborted — so moving the folder in `Snapshot` would park the user's only copy of their data in a staging directory that nothing lists, for the whole window until `Commit` ran. It also keeps `Abort` unable to reach the app folder, which is the difference between discarding a staged copy and deleting an app. - **It implies a single pass, and that pass is `Pass: 2`.** The number identifies the *consistent* pass, not the ordinal; there is no live pass for a stopped app to be incremental against, and committing a `Pass: 1` snapshot would discard it as the torn throwaway it normally is. ### No fallback to local If the default engine cannot be reached, the uninstall **fails** and the app stays installed and running, with the error on its tile. Falling back to a local archive is the tempting behaviour and the wrong one: it would put the data somewhere other than where the user was told it goes, at the exact moment the tile disappears and stops being able to say so. The escape hatch is to point the default engine at this server's disk in Settings → Backups and retry. --- ## Scheduling **Daily, default 03:00 local, driven by Maison.** Maison has no scheduler today — backups are uninstall-triggered or manual — so this is new code, not a setting. It cannot be delegated to the engine's own scheduler at any price: a consistent app snapshot requires stopping containers, and no backup tool can do that. Three requirements beyond a ticker: | | Why | |---|---| | **Serialise across apps** | The per-app `enter`/`leave` lock protects one app. Nothing stops a nightly run from taking six apps down at once. | | **Catch up, do not pile up** | A PCS that was off at 03:00 backs up when it returns; a run that overruns its window is skipped, not queued behind itself. | | **Jitter per box** | A fixed 03:00 fleet-wide is a self-inflicted thundering herd against one bucket. The `deviceId` issued to the PCS is a ready-made stable seed. | ### The first run is a different problem from the nightly one Jitter spreads the recurring 03:00 load across minutes. **The first backup a box ever takes needs spreading across days.** Configuration reaches the fleet through `ensure-template-sync`, so every PCS becomes backup-capable at roughly the same moment. Each then seeds its *entire* `/DATA` on its next scheduled run. Two hundred boxes holding even 100 GB each is tens of terabytes converging on one bucket in a single window — and the users are concentrated in a handful of timezones, so "03:00 local" barely spreads it. **Maison owns this**, because Maison owns the trigger. Nothing upstream can stagger it. Derive a per-box offset over a multi-day window from a stable seed (the `deviceId` works), and hold the first run until that offset elapses. Subsequent runs fall back to normal nightly jitter. This matters most on the day the feature ships to existing boxes, which is exactly when it is least convenient to discover. ### Upload throttling A backup that saturates a home uplink is a support ticket, and the user will blame the PCS rather than the backup. Expose a bandwidth cap and set a conservative default; the engines take a flag for this (`--upload-limit-mb` in kopia's case). It belongs in the same settings surface as the schedule. --- ## Retention **Tiered / GFS: keep 7 daily, then one per week, then one per month** — the default of a setting, not a constant. Because each app is one stable source accumulating snapshots over time, this maps directly onto an engine's own retention policy (`--keep-daily 7 --keep-weekly 4 --keep-monthly 12`) instead of having to be reimplemented. That is a direct dividend of the stable source path: with a fresh source per backup, every source would hold exactly one snapshot and per-source policies would be meaningless. ### Retention belongs to the engine, not to the box **One policy per engine, resolved by `Config.Effective(engineID, provisioned)`**: engine override → box-wide setting → provisioned default → compiled default. There is no box-wide number that means the same thing everywhere, because it cannot: | | local engine | repository engine | |---|---|---| | what a backup is | a **full second copy** of the app folder | an incremental snapshot | | cost of 23 generations | 23× the app, on the app's own disk | a little more than one | | who expires them | Maison, through `retention.Plan` | the engine's own policy | So tiers are **unsound on the local engine** — not merely wasteful. `Effective` collapses a tier mode to a count for it (`ModeSmart` → `ModeCount` at `Keep.Latest`, which the smart preset already puts at 2: the copy being replaced and the one before it), and the settings page never offers tier modes for an engine whose `Caps` declare `NeedsLocalSpace`. An explicit count, age or keep-everything is left alone — each is a bound the user chose. This replaced a separate `keep_local` count that sat beside the tiers in one box-wide block. It was the wrong shape twice over: it was a second retention vocabulary for the same question, and on a local-only box — which is most of the fleet — it was the *only* one of the four numbers that did anything, sitting next to three that were inert. **Nothing can empty the local directory.** `retention.Plan` keeps the newest backup whatever the policy says, which is what the old "keep 0 local, but only once another engine has actually listed the backup" path existed to guarantee. That guarantee matters most for the update rollback point, which is always local and has nowhere else to be. Delegating remote retention is clean for kopia. An engine whose policy model differs means Maison expresses the *intent* through the provider interface rather than assuming the flags. --- ## Encryption and the master password > **The master password is generated on the PCS and never reaches Yundera.** It lives at `AppDataShared/backup//repository.password`, mode 0600. It has to stay on the box — an unattended nightly backup cannot prompt for it. The user takes their own copy: the dashboard shows the key on demand, and Maison mails it once on the first boot where it can (see below). Stated plainly, because it must be stated plainly to users too: **the PCS holds it, the user holds a copy, Yundera holds nothing and cannot recover it.** The consequence is accepted rather than mitigated: a user who keeps no copy loses the backups, and no support path exists. Any future softening has to preserve the first line — client-side wrapping under a user credential with only ciphertext stored server-side is the shape that does; server-side derivation is the shape that does not. --- ## Notifications Maison gets an outbound SMTP client. Two jobs, and only two. **1. Alert on backup failure.** What makes backups worthless is silent failure, and an unattended nightly job with no channel out is exactly that. **One mail on transition into failure, one on recovery** — not one per failed run, which becomes noise and then a filter rule. The sticky-error state already tracked per tile is the right trigger source. **2. Handing the user their encryption key.** This is the only reason a copy of the password exists anywhere but the box. > **This one cannot go through the PCS's `smtp` container.** Every PCS runs > `ghcr.io/yundera/mail-gateway`, which parses the message and forwards it over HTTPS to > `mesh-router-backend`, which relays via SendGrid. Mail sent that way traverses **Yundera > infrastructure in plaintext** — which contradicts the claim above that Yundera holds nothing > and cannot recover the key. So the two jobs use different transports: | Job | Transport | |---|---| | Failure alerts | The built-in `smtp` container. Zero config, already working, nothing secret in the body. | | **The key** | Both. **Displayed in the dashboard** on demand — nothing leaves the box. **And mailed once**, automatically, on the first boot where a mail server answers. | Display costs nothing and preserves the security property exactly: the key goes to the browser of someone already authenticated as the owner and stops there. It is offered first for that reason. The automatic mail is a deliberate trade, made against a likelier failure than the one above. A key that only leaves the box when someone presses a button is a key most users never copy, and the day that matters is the day the disk is gone — at which point a plaintext secret sitting in an inbox is the difference between a restore and nothing. Where the relay is Yundera's, the caveat above applies in full and is the accepted cost; a deployment that supplies its own SMTP credentials avoids it. The mail is sent **once per box**, and "once" is anchored on a receipt file rather than on anyone remembering: ``` ${StateDir}/backup-key-sent.json 0600 {"sent_at":…, "to":…, "engine":…, "auto":true, "fingerprint":…} ``` `fingerprint` names *which* key the receipt covers: the first 16 hex digits of the password's SHA-256 — enough to notice the key changed, nothing that gives it back. A space reset from the deployment's dashboard makes the box mint a new key, and without this the receipt would go on saying "already sent" about a key that no longer opens anything. So the automatic send fires again when the key on disk no longer matches the receipt, and the check runs from the five-minute detector loop as well as at boot, once per pass, so a key that changes while Maison is running is mailed without a restart. A receipt **without** a fingerprint — every receipt written before it existed — counts as matching, so an upgrade mails nothing. A key the user typed into the recovery form is recorded as `"held_by_user": true` with its fingerprint and no `sent_at`: they demonstrably hold a copy, and mailing it back would only add a second plaintext one to an inbox. Written *after* a successful send, so a crash mid-send costs a duplicate mail rather than a key that is never handed over — the right way round. A malformed receipt reads as *already sent*, for the same reason: the failure this file exists to prevent is mailing the key on every restart. Absent, it means no copy has ever been mailed, and the settings page says so. The boot-time send retries for a few minutes before giving up until the next boot — on a PCS the mail relay is a sibling container and boot order between the two is not guaranteed, so a single attempt at t=0 would fail on exactly the deployments this is for. It gives up rather than retrying forever: past that window, "no mail server" is an answer and not a race. The button remains, unconditionally: a user asking for the key again has a reason, and refusing because a receipt exists would leave them with no way to reach a secret that is theirs. Both paths are **clearly labelled as unrecoverable**. ### Where the transport comes from Two layers, resolved by `usersettings.Settings.EffectiveSMTP` under the same "store only the override" rule as retention: **the deployment provisions it, the box may override it.** The deployment's layer is the environment, read at boot into `config.Config.SMTP`: | Variable | Meaning | |---|---| | `SMTP_HOST` | The relay. **The switch — no default.** Unset means this deployment ships no mail, which is the standalone local install and must stay silent rather than dial a guess. | | `SMTP_PORT` | Defaults to 587. An unparseable or out-of-range value falls back rather than failing the boot. | | `SMTP_USER` / `SMTP_PASS` | Optional; sent only if the relay advertises `AUTH`. | | `SMTP_FROM` | Defaults to `noreply@${APP_DOMAIN}`. | | `SMTP_TO` | Defaults to `${APP_EMAIL}` — the box's owner. | | `SMTP_SECURITY` | `starttls` (default, opportunistic), `tls`, or `none`. | `From` and `To` fall back to `.env.app`, read **live** like every other value from that file: a box that changes owner or domain must not keep mailing the previous one until someone restarts the dashboard. Those two fallbacks are why a PCS only has to name its relay. Overriding: the box's own `smtp` block in **`settings.json`** wins field by field, except that the **transport travels together** — host, port, credentials and security — for the reason `Mode` carries its own parameter in `Effective`: a host set on the box with the deployment's credentials inherited is a login sent to the wrong server. `From` and `To` resolve independently, because the recipient is the one a user actually changes. It sits in the settings rather than in `backup.json`, where it started, because it is a property of the box and not of the backup schedule — the next thing worth mailing about, a certificate about to expire or a disk filling up, should not have to reach into the backup configuration to find a relay. A box configured before the move keeps its `smtp` key under `backup.json`; `server.adoptLegacySMTP` carries it across on the next boot and clears the old one, so alerting never lapses across the upgrade. On a PCS `SMTP_HOST` is the sibling `smtp` container, which advertises neither STARTTLS nor AUTH; the opportunistic default handles that without configuration, and `SMTP_SECURITY=none` states it. --- ## Restore ### One app Three paths, chosen by **where the backup is** and whether there is room — never by which engine is currently selected, which is the same rule that governs listing. ``` on disk rename the live folder aside, rename the archive in. Instant, atomic, and the displaced state becomes an archive of its own, so the restore is itself undoable. remote, room the engine materialises it into the local tree, then exactly the above. RestoreBackup runs unmodified. remote, no room the engine writes over the live folder. ~1x space. NOT atomic. ``` `EstimateRestore` chooses between the last two — the restore-side sibling of the backup guard, and for the same reason: materialising needs room for a full second copy, which is exactly what an app large enough to need this does not have. **The in-place path is the one that gives something up.** It is not atomic: an interruption leaves the folder holding neither the old state nor the new one, and because there is no local copy the only way back is a remote snapshot — so the restore is reversible only while the repository is reachable. That is the trade for being able to restore an app too large to fit twice on its own disk, and it is why three guards are mandatory rather than best-effort: 1. **An undo snapshot is taken first, and if it fails the restore is refused.** The app is already stopped, so it is consistent; it is incremental, so it costs about the delta. An unrecoverable overwrite is worse than a restore that did not happen. 2. **A `.restoring` marker**, written *outside* the folder being replaced — a restore that deletes files absent from the backup would otherwise delete the marker too. Its name cannot parse as a stamp, so no lister mistakes it for an archive. 3. **The marker gates `EnsureStarted`.** An app whose restore was cut short stays down. Starting it would initialise over the gap — fresh database, default config — and that invented state would become the next night's backup. Two knock-on changes: - **Done.** Listing is every engine's entries, grouped by engine and **not** deduped — `Set.List` for one app, `Set.ListAll` for the global page. The caching the earlier draft called for turned out to be the wrong fix: the cost was never one query, it was one query *per app*. `Provider.ListAll` returns every app in a single call instead, so the page costs one subprocess per engine however many apps it shows, with nothing stale to invalidate. - Install-from-backup still goes through `RestoreBackup`, so it inherits the first two paths unchanged. It does **not** get the in-place path, and does not need it: a fresh install has no live folder to write over. ### The user-data set Two modes, and the safe one is the default in the UI. ``` copy restore into ${DATA_ROOT}/Restored// (or a named directory). Touches nothing that exists, so no undo snapshot and no marker. This is what "get my Documents from three weeks ago back" wants. in place restore over the live tree, entry by entry, with --delete-extra. Exact within each restored entry; destructive by construction. ``` **In place is done entry by entry, never by aiming at the data root, and that is a safety property rather than a style choice.** `--delete-extra` is what makes a restore a restore rather than a merge — it removes what the snapshot does not contain. Aimed at `${DATA_ROOT}` it would therefore delete `AppData/`, which is excluded from this snapshot by policy, taking every app's data on the box with it. Restoring each top-level entry into its own path means `AppData` is never a target and cannot be reached. `TestUserDataRestoreNeverTargetsTheDataRoot` runs the real engine and asserts exactly that, because both spellings read as "restore the user data" and only one of them destroys the box. Three consequences, all of which the UI has to state rather than imply: - Within a restored entry the result is exact: files created since are removed, deletions undone, modifications reverted. - A top-level entry on disk but *not* in the snapshot is **left alone**. A true restore would remove it, but silently deleting a whole tree the user made since is a worse surprise than leaving it. - `AppDataShared/` is skipped in place, though it is in the snapshot on purpose (so that recovering any engine's backup returns the others' credentials). In place it is the one directory whose live copy is newer than the backup by definition, and it holds the configuration the engine performing the restore is reading. The guards mirror the app in-place path, and each refuses rather than proceeding: 1. **An undo snapshot first**, so the restore is reversible; it is incremental, so it costs about the delta. If it fails, nothing is written. 2. **A marker** under `AppData/maison/` — outside the set being restored, so the restore cannot delete its own marker. It outlives a restart, and while it is there the page says the tree is neither state and offers to finish the job. 3. **One at a time, in both directions.** A second restore is refused, not queued: two writers over one tree produce a state that came from neither backup. A restore is also refused while a backup run is going, and the nightly run skips its user-data target while a restore is in flight — a snapshot of a half-restored tree is a state that never existed, and it counts against retention, so it can push out the very snapshot being restored from. Copy mode has a guard of its own, and only it needs one: it writes a full second copy of the set onto the same disk, and the set is the largest thing on the box by design. It is refused when there is not room, with 10% headroom, matching the app path — filling the data disk does not merely fail the restore, it starts failing every app writing to that disk. In place needs no extra room, so the guard deliberately does not apply to it: refusing there would make a full disk unrecoverable by the one operation that could fix it. Apps are **not** stopped for this. The set holds no app state — `AppData/` is excluded — so what is being replaced is Documents, Downloads and Media. An app holding an open handle on a media file being replaced is in the same position as a user overwriting that file over Samba, which is an ordinary thing to do to a NAS. ### Disaster recovery The scenario is a user who has lost the box entirely. 1. A fresh PCS is provisioned. Identity returns the normal way — JWT, domain, mesh-router registration — none of which depends on the backup. 2. The box boots into **recovery mode**: admin stack and domain up, nothing else. 3. The recovery UI asks for the provider and the master password. 4. It connects, lists snapshots, restores. > **The recovery form has one field.** The repository location is derivable from > identity — the fresh box's credentials resolve endpoint, bucket and prefix — so the > only thing a user must supply is the one secret Yundera never had. Treat that as an > invariant; it erodes the moment someone proposes asking the user to pick a bucket. Restoring onto a *different* box needs no new mechanism. Storage is modelled as one space per user with one key per attached device, so a fresh PCS is simply a new device attaching to the same space. **Recovery mode is a state of Maison, not a separate program.** Maison already has the provider layer, `Materialize`, the app registry, per-app progress events, and the requirement to tolerate absent configuration by degrading to "not configured". Recovery mode is that same state plus one question: *there is no config — but is there a repository?* A separate recovery service would duplicate all of it and, being used once per user per lifetime, would be the least-tested code in the product at the exact moment it matters most. Sharing the path with ordinary single-app restore means it is exercised continuously. #### Four things that go wrong if they are not designed in | | | |---|---| | **Verify before destroying** | Connect and list snapshots *first*, so a mistyped password is caught before the box has been reprovisioned or wiped. | | **The scheduler stays off until recovery is explicitly confirmed complete** | A nightly run firing mid-restore snapshots a half-restored PCS, and under keep-7-daily a week of that evicts every good daily. Weeklies and monthlies survive, so it is recoverable — but it must be impossible. Require an explicit action; **inferring completion is how this happens.** | | **No app starts before its data has landed** | An app brought up against an empty `AppData/` initialises fresh — new database, default config — and then either the restore writes over a running app, or that fresh state becomes the next night's backup. Gate per app: restore → verify → start. | | **Restore is resumable** | A few hundred gigabytes over a home uplink is days, and the box will reboot inside that window. Persist per-source progress; resume, never restart. | #### Restore order is a feature ``` 1. engine config near-instant; on a multi-engine box this returns engine #2 2. small apps, one by one dashboard, notes, passwords — a working PCS in minutes 3. user data / media streams in the background for hours or days ``` Restores are per-source anyway, so ordering costs nothing to implement and changes the experience from "unusable for three days" to "usable in ten minutes, complete in three days". #### Recovering on a rebuilt box (what ships today) The first slice of the above, for the commonest case: a box reinstalled onto the same backup space. The host side finds a repository it has no password for, refuses to create a second one, and leaves a `needs-recovery` marker; from then on the engine's `status` answers `needsRecovery: true`. - **The page asks for the key.** The engine's row turns red ("needs your backup key"), the page leads with it, and the engine's tab shows a key field in place of the key block. The field is offered only when the adapter declares `Caps.Recover`; an older adapter hides it. A link to start fresh appears when the host wrote a `recoveryHelpUrl` into `state.json`. - **`POST /api/backup/engines/{id}/recover`** with `{key}`. Maison writes the key to `repository.password.candidate` (0600) in the engine directory — the only file it ever writes there — and runs the adapter's `recover` verb, which connects (never creates), promotes the candidate to `repository.password` and pins every existing snapshot. The candidate is removed on every path. 400 is a wrong key (exit 14), 409 an engine not waiting for one, 501 an engine that cannot take one. The key never reaches a log, and the receipt records it as held by the user. - **The existing snapshots are pinned**, so retention cannot expire them, whatever happens next. - **Then the schedule question**, as a dialog: the box was just reinstalled, its apps may be empty, and nightly runs would save that emptiness and push the real history out of retention. The default is **pause** (`paused` / `paused_at` in `backup.json`): the timer stops, a manual run still works, and a banner with *Resume* stays on the page until it ends. This is the "scheduler stays off until recovery is explicitly confirmed" rule above, in the form a user can act on. - **Two incidents** keep both states visible: `backup.recovery:` (critical, raised whatever the schedule says, closed once the key is in) and `backup.paused` (warning, open while the schedule is on and paused). With only backup incidents open, the bell opens Settings → Backups. What it does not do yet: restore order, gating app start on restored data, or resumable restores. Those remain the design above. Recovery is a **one-way restore, not a merge.** Restoring an older snapshot and then letting the nightly run creates a rollback point in the chain — correct under GFS retention, but a deliberate choice rather than a surprise. --- ## A backup belongs to one engine **The identity of a backup is `(engine, app, stamp)`.** Engines run in parallel and never coordinate: a stamp present both in the local archive tree and in a repository is two backups that happen to share a name. They were written by different runs, they can be deleted independently, and restoring one is not restoring the other. Listing used to fold them into a single row carrying a three-valued tier — `local`, `remote`, `both`. That worked only while exactly one engine was offsite. Add a second remote engine and two unrelated snapshots merge into one row reported as "on this disk + offsite", which is true of neither; `Locate` then picks between them arbitrarily and a delete takes both. The tier is now a property of the engine holding the backup — `local` or `remote`, never a summary — and `both` is gone. What follows from it: - **Delete is per engine, and the engine is required** — no "delete everywhere" default. Clearing space on the data disk must not take the offsite copy with it: that copy is the entire disaster-recovery story, and a delete that silently spanned engines would destroy it while appearing to tidy a local folder. - **Restore names its engine.** `Set.LocateIn` is the shape the UI uses, because the user clicked a row and a row belongs to an engine — including in the store's install-from-backup picker, which offers a group per engine for the same reason the page tabs by one. `Set.Locate` — "whichever engine has it", preferring an instant restore — is kept only for a request that arrives without an engine. - **The selected engine still governs writes only.** It decides where the nightly run, an uninstall and "Back up now" put the next backup. It has no bearing on what is listed, restorable or deletable, which is what keeps a user's history reachable after they switch. - **The Backups page is a tab per engine.** Everything inside a tab — the app list, the totals, the "your files" card, the same-disk warning — belongs to that engine, and so does every button in it. The app's own Backups tab keeps a single list with a group header per engine, because it is one app across all of them. - **The user-data set is read per engine too.** `UserData.Engine()` is still the writer, because writing the set is the schedule's job and the schedule has one destination; but `ListIn`/`AvailableIn`/`Restore` take an engine, so a box that wrote its files to a repository and then switched its default still lists and restores them. Previously that card showed only the writer's snapshots, so switching to the local engine — which can never hold the set — made them look deleted. - **The restore state is box-wide.** One restore at a time, whichever engine it came from, so every tab reports the same one rather than each pretending to have its own. Engines are named on screen by the deployment, not by this package: `engineInfo.Name` comes from the host-written state file, so a provisioned PCS can call its space "Yundera Backup Storage" while the identical engine pointed at a self-hoster's own bucket claims nothing of the sort. The engine **ID** stays a bare engine name — it is recorded on every backup it writes and is machine identity, so branding has no business in it. --- ## What Maison consumes from the PCS Maison **never fetches credentials.** Two self-check scripts on the host do it: `ensure-backup-credentials.sh` exchanges the box's JWT for a scoped, expiring storage key and parks it in the PCS secret env, and `ensure-backup-config.sh` turns that into a connected repository under `AppDataShared/backup//`. Maison reads what it finds there and shells out to the engine. This keeps credential fetch, key rotation and suspended-space handling in the host path that already does that work, and keeps Maison's blast radius at *reads files it does not own*. Four of them, and the split between them is the whole design: | File | Written | Read by Maison | |---|---|---| | `repository.config` | **once**, at first connect | identity and storage type | | `repository.password` | once, with the repository | every invocation | | `credentials.env` | on **every** rotation | every invocation, fresh | | `state.json` | every host run | writability and credential expiry | **The storage credentials are deliberately not in `repository.config`.** kopia persists whatever it was given at connect time and offers no flag to stop it (`--no-persist-credentials` governs the repository *password*). The host script therefore blanks the two persisted fields and rotates `credentials.env` alone; kopia accepts `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` from the environment for every ordinary operation and does not write them back. The reason is identity. Installing a rotated key by re-running `repository connect` rewrites the whole configuration file, and a reconnect that omits `--override-hostname` / `--override-username` silently refiles the box under the engine container's random hostname and `root`. kopia keys snapshots `user@host:path`, so that costs a full re-hash of every file and leaves per-source retention pointed at a lineage nothing writes to any more, while the new one is covered by no policy at all. `snapshot list --all` still shows everything, so none of it is visible from here until the storage bill grows. Writing the configuration exactly once removes the failure mode rather than documenting it. Three requirements follow: - **Tolerate the files being absent.** On a box where the host side has not run, the backup UI degrades to "not configured" — it does not error. This is the same state recovery mode builds on. - **Tolerate them being stale.** A rotated credential is picked up on the next invocation because `credentials.env` is read per run, not at construction. - **Never branch on who wrote them.** A hand-written `repository.config` pointing at a local MinIO or a filesystem repository must exercise every code path — that is how this is developed and tested. A filesystem repository has no `credentials.env` at all, and that is not an error. **Two more things the engine image dictates, both found by running it rather than reading it.** kopia's container bakes `KOPIA_CACHE_DIRECTORY=/app/cache` and `KOPIA_LOG_DIR=/app/logs` into its own environment, and those **outrank** `--cache-directory` and `--log-dir` on the command line. `/app` belongs to root, so an engine running as `PUID` cannot create either and every command fails with `unable to create cache directory: mkdir /app/cache: permission denied` before it ever reaches the repository — which is why `engine.Spec` carries `Env` (by value) alongside `Secrets` (by name). And kopia's `--endpoint` is a `host[:port]`, not a URL: the credential API returns a URL because that contract is engine-independent, and the host script strips the scheme because only it knows what kopia parses. The host side also leaves two markers in the same directory, both currently written and not yet read: - `needs-credentials` — the storage refused the key. The credentials script consumes it on the next cycle and mints a fresh one. - `needs-recovery` — the backup space already holds a repository and this box has no password, i.e. the box was rebuilt. The host script stops rather than initialising a second repository under the same prefix. This is the seam recovery mode plugs into. A credential may also arrive marked **not writable** (a storage space suspended for quota abuse). That path must degrade correctly rather than error: **restores and prunes work, writes fail.** It is worth building from the start, because the graceful handling lives on this side and retrofitting it during an incident is expensive. --- ## Not yet decided - **What "custom" means** in the engine picker. Half-answered by [The engine adapter](#the-engine-adapter): *user supplies a repository target for an engine the deployment ships* is ordinary and is what the picker offers, while *user supplies an image or a command* remains a deliberate remote-execution surface and is not offered. What is still undecided is whether a user-installed adapter ever becomes a supported thing for a self-hoster who wants a destination Yundera does not ship, and if so what distinguishes it from the provisioned one in the UI. - **Whether the engine container drops to `--network none` for a filesystem repository.** It gets that isolation today only as a side effect of running one-shot, which [The engine adapter](#the-engine-adapter) removes. The stack would have to choose its network at deploy time from what `repository.config` says. - **Scheduling surface** — how much of the schedule a user can change (time, tiers, per-app opt-out). - **Where recovery mode sits in the boot sequence**, and how a fresh box knows to enter it rather than come up empty. --- ## Provenance Derived from `BACKUP-STORAGE-PLAN.md` at the repo root, which remains the working document and additionally carries the server-side storage design (per-user space provisioning, scoped credential issuance, usage metering and reconciliation) that this document treats as an external contract.