# Lessons Learned Non-obvious knowledge this project paid for, kept because it prevents a costly mistake from being repeated. Current-state facts belong in the design documents and `lab/README.md`; how the project got here belongs in Git. ## Windows writes an unflushed file out on its own within seconds **Area:** lab recovery rows, fault injection **Status:** Active **Symptom:** The first power-off rows for a deliberately unflushed rename (`skip_flush`) found the file intact after the cut, and recovery correctly reported nothing to repair — so the rows proved nothing about the zero-fill detector they exist for. **Cause:** Windows' lazy writer flushes the dirty pages by itself within a few seconds. The damaged-file symptom exists only when the cut lands about a second after the unflushed write; with a 20 s gap the data had already reached the disk. It reproduced for a 4 MB and a 1 KB file alike, so the file's size is not what decides it. **Do not:** read an intact file after a `skip_flush` power-off as proof that unflushed writes are safe, and do not lengthen the hold at the fault point to "make sure the cut lands inside the window" — the hold is what gives the data time to reach the disk. **Use instead:** interrupt on the boundary rather than on a clock. The engine's `--fault-signal` creates a file when a fault reaches its boundary and then holds; the lab's `signal` interruption trigger polls for that file every 100 ms and cuts the power then, so the cut lands inside the window by construction rather than by aim. A row whose file is nonetheless intact still reports WARN ("proved nothing") rather than FAIL, because an interruption that damaged nothing is absence of evidence. **Prevented by:** `durability.rs::a_fault_signals_the_boundary_before_it_acts` asserts the announcement precedes the action; `lab/Invoke-RecoveryRows.ps1` builds every recovery row's interruption from that signal; the local zero-fill test in `crates/tigersetup-setup/tests/durability.rs` covers the detector itself. **Generalization candidate:** none any more — the generic capability this lesson used to ask for, interrupting at a marker the product writes, exists in TigerWinLab and is what the rows now use. ## Halves of a privilege transition do not prove the handoff **Area:** elevation acceptance, lab rows, verification design **Status:** Active **Symptom:** the 0.5.1 self-installer hung after a real UAC consent — wizard responsive, "elevating", never finishing — while every elevation check was green: the unelevated wizard raised the prompt and stayed responsive, a refusal recovered, the already-elevated wizard installed machine scope, and a process test handed an outcome document back across the boundary. **Cause:** `run_elevated` had been written for hidden `/quiet` dependency installers and passed `SW_HIDE` to `ShellExecuteEx`; a process's first window follows the show state it was started with, so the elevated wizard child was created invisible and waited for a click nobody could give, and the parent waited for the child. Every half of the acceptance ran on one side of the consent or the other — the lab could not answer a prompt, and the substitutes were chosen to be the same *authority* — so no check ever started the child the way the product starts it. **Do not:** accept evidence for a privilege transition assembled from a pre-consent half and a post-elevation half, however faithful each is. A substitute that shares the authority does not share the launch. And do not weaken Windows to make the transition testable — no `PromptOnSecureDesktop=0`, no UAC policy change, no auto-approval. **Use instead:** one wizard end to end on the artifact under release, with the genuine prompt answered on the real secure desktop: the lab's host console types on the VM's console as a person would (`Approve-ElevationPrompt`), and the row asserts the transition rather than assuming it — the prompt this press raised, consent.exe gone and the input desktop back, exactly one elevated process and it is this installer with `--elevated-result`, a visible child wizard, the parent stepped aside and responsive, and the child's outcome read from the parent's exit code and printed document. A high-integrity window is driven through the lab's elevated agent; the standard session cannot click into it (UIPI). The prompt kind follows the desktop — credential prompt for a standard user, Yes/No consent prompt with *No* as its default for an administrator — so both are rows. **Prevented by:** `lab/Invoke-ElevationRows.ps1` rows `complete-uac` and `complete-uac-admin` (`TigerSetup-Validation.md` §5.2); the unit test `elevation.rs::a_wizard_child_is_started_shown_and_a_quiet_one_hidden` pins the show state to the child's command line. **Generalization candidate:** TigerWinLab has the mechanism (R17) and its README carries the Windows half of the lesson; the method — a transition is proven on one operation, never from its halves — is TigerAiCore material and must be raised there as its own change. ## A capability recorded as missing may never have been missing **Area:** consuming a Lab, external blockers **Status:** Active **Symptom:** Every machine-scope interactive row failed identically on all three platforms and in both languages: the run advanced past the wizard's scope page, installed per-user, and then failed on the install root a machine-scope run must create. It was recorded as a TigerWinLab requirement — "a wizard driver must be able to choose a radio button" — and five rows were left failing as blocked on another project. **Cause:** the lab had clicked the radio button, and the wizard had taken the choice. The lab's own `interactive-install.json` for the failing row records `"prepared": "Install for all users"` on the scope page and, four pages later, a summary reading `Install mode: For all users` beside a destination still inside one user's profile. The defect was TigerSetup's: the destination page was filled once from the scope the run started in and never followed the scope the person then chose. Two pages of one wizard disagreed, and the row reported that disagreement where it became visible rather than where it was caused. **Do not:** conclude that a provider lacks a capability from the shape of a failure. A whole class of failures looking identical is evidence that one thing is wrong, not evidence of where it is. **Use instead:** before recording a requirement against another repository, open the provider's own evidence artifact for the exact step claimed to have failed, and read what it says happened. The lab writes one per phase; reading it costs a minute and it is the difference between a defect fixed here and a blocker filed there. Where the artifact shows the provider did its part, the defect is the consumer's. **Prevented by:** `wizard.rs::the_destination_follows_the_scope_until_it_is_edited` drives the real window and asserts the destination moves with the scope, and stops moving once the person has typed their own. **Generalization candidate:** TigerAiCore — "read the provider's evidence before filing a provider requirement" is method knowledge about consuming a Lab, not anything specific to installers. ## A worker pinned to a worktree by instruction still drifts **Area:** delegation, worktrees **Status:** Active **Symptom:** With the Agent tool's worktree isolation unusable on this machine, workers were told to `cd` into their worktree in every shell command. One worker resumed after an interruption ran `cargo fmt`, `cargo build` and `cargo test` three times in the main checkout before noticing; only build output landed there, but that was luck, not design. **Cause:** the shell's working directory resets between a worker's tool calls, and an instruction is not a boundary. **Do not:** treat an instruction to work in a worktree as isolation, or skip checking the main checkout after a worker finishes. **Use instead:** `git status --porcelain` of the main checkout after every worker report, before integrating; file edits by absolute path (which did hold); and real isolation when the tooling provides it. **Prevented by:** the Lead's integration step reads `git status` first; nothing mechanical stops the worker. **Generalization candidate:** TigerAiCore — the check belongs to the autonomous development method, not to this project. ## Read lab evidence from the job that produced it **Area:** lab driver, evidence handling **Status:** Active **Symptom:** A wizard capture read as "English under the Polish account" and became a finding about language detection, complete with a planned change to the detection order. The lab's own artifact for that job showed Polish. **Cause:** the capture runner copied a row's screenshots by searching the whole lab-artifact tree for a file name that every language × scale combination shares, and took the first match — another combination's file. **Do not:** derive a finding from a file copied or renamed by a driver without checking the job artifact it claims to come from, and do not search an artifact tree by a name that more than one job produces. **Use instead:** copy and read from `\\` — the lab's result names the job — and record the job id beside every copied artifact. **Prevented by:** `lab/Invoke-UiCaptureRows.ps1` copies only from the row's job directory and reports the job id in a check. **Generalization candidate:** none — the TigerWinLab result already names the job; the mistake was the consumer's. ## A finished process is not a finished task, and "timed out" is not "stopped" **Area:** delegation, monitors, completion reporting **Status:** Active **Symptom:** A run reported "all subagents finished, all lab and build processes terminated ... no lab lease held" while the session still showed two shells running and an active researcher whose status read *Confirming no lingering probe processes*. Investigated afterwards, the machine held **seven** orphans and the session held **three** live task records — the youngest already an hour old, the oldest an hour and three quarters. **Cause:** one bad verification, and three ways for background work to survive a completion claim. - **The check measured the wrong things, twice over.** It was a `Get-CimInstance Win32_Process` sweep filtered to `pwsh`, `cargo`, `rustc`, `dotnet`, `zopfli-probe`, `tiger-setup`, `tigersetup-setup`. Every orphan was a `tail`, a `grep` or a `find`, so the filter excluded all of them — and a process sweep of any width cannot see a task whose process has died but whose registry entry never reached a terminal state, which is what the two shells were. - **A monitor's timeout does not stop its pipeline.** Three monitors reported *"[Monitor timed out — re-arm if needed]"*, and all three left their `tail -f … | grep --line-buffered …` running. `tail -f` has no reason to exit; the notification is about the watch, not about the processes. - **A delegated worker's background work is the Lead's.** A researcher had started `find / -iname "zopfli-probe"` and `until grep -q "^264," ; do sleep 15; done`. Both were registered against the Lead's session and outlived the worker's own final report. Neither was bounded: a `find /` walks a filesystem, and the marker the waiter polled for went to a different file once the run it watched was restarted, so it could never appear. - **The worker's self-report was believed.** It stated "no orphaned process remains" while holding both. **Do not:** answer "is my background work finished?" with a process list, and never with a list filtered by the executable names you happened to think of. Do not read a monitor's timeout as a termination. Do not let a delegated worker's account of its own cleanliness stand in for checking it. **Use instead:** the task registry — `/tasks`, or the task tools — as the authoritative account, and stop each live entry explicitly. Supplement it with a sweep keyed on **what the work touches** (the repository path, the session's scratchpad directory, the Lab, the probe name) rather than on process names, so that a `tail` or a `find` cannot hide behind a filter. Stop a monitor's pipeline yourself; give every waiter a bound it cannot miss; and require the same of a delegated task in its prompt. **Prevented by:** nothing mechanical — the registry belongs to the harness, not to this repository, so no check inside TigerSetup can enforce it. The compensating controls are procedural and deliberately high-salience: `AGENTS.md` makes reading the registry and sweeping by work-touched the last step before reporting completion, and `TigerSetup-AI-Approach.md` §3 puts the constraint on delegated tasks where this project's delegation policy lives. **Generalization candidate:** TigerAiCore. The rule this violated is already there and is explicit — *"Polling indefinitely for a string … is not a termination condition"* describes one of these orphans exactly. What is missing is the sentence naming **what counts as evidence** that the rule was honoured. That belongs in TigerAiCore and must be made as its own change, not from here. ## A lab row measures the engine in the installer, not the one in the workspace **Area:** lab rows, build provenance **Status:** Active **Symptom:** Recovery rows that had just been rewritten to interrupt on a boundary the engine announces reported that the interruption trigger never fired, twice, at about six minutes a time. The engine code was right, its local test passed, and the installers had been rebuilt after the change. **Cause:** the installers had been rebuilt, but the **engine had not**. `tiger-setup build` embeds `tigersetup-setup.exe` from beside itself, so the installers carried a release engine from before the change and rejected the new `--fault-signal` argument outright. Nothing in the result said so: the run simply never reached the fault, so no signal was written and the trigger timed out looking for it. `cargo build --workspace` had been run many times; it does not touch the release profile the builder reads. **Do not:** rebuild an installer to pick up an engine change. That is not what it does. And do not read a lab row as evidence about the code in the working tree — it is evidence about the bytes inside the installer, which is exactly why the exact-artifact principle exists. **Use instead:** `cargo build --release` first, then the installers, then the rows. Where a row is about to run, the two can be compared: the installer declares the SHA-256 of the engine it carries, and the builder's engine is the file beside the builder. **Prevented by:** `Assert-TigerSetupEngineIsCurrent` in `lab/TigerSetupLab.psm1`, called by both row drivers before any guest work, refuses a mismatch with both hashes and the command that fixes it. It costs a second; discovering the same thing from a matrix costs the matrix. **Generalization candidate:** a shared Lab or TigerAiCore — "assert the artifact under test was built from the tree you think it was" is general to any harness that validates a built artifact rather than a source tree. ## A lab row's own defects must be found before the guest time is spent **Area:** lab driver, PowerShell **Status:** Active **Symptom:** Three separate rows died *after* their guest work succeeded, each on a property that was not there, each costing the minutes the row had already spent. `The property 'status' cannot be found on this object`; `The variable '$Theme' cannot be retrieved because it has not been set`; `The property 'Count' cannot be found on this object`. **Cause:** three different ways a PowerShell value silently is not what the code assumes, all fatal only under `Set-StrictMode -Version Latest`, and all reached only after the expensive part of the row: - **A local overwrote a parameter differing only in case.** `$prepare` and `[string[]] $Prepare` are one variable; the type constraint survives the binding, so a run object was coerced to a single-element string array. - **A line was pasted into a function that has no such parameter.** An added `$Theme` read compiled fine and threw at run time. - **A function returned an empty array.** PowerShell unrolls a returned collection, so `return @()` returns *nothing*, the caller's variable becomes `$null`, and the next `.Count` ends the row. It only ever happens on the path where the guest produced no logs — a `BUSY` lease, a failed job — which is exactly the path a row is least often exercised on. **Do not:** rely on running a row to find out whether the row is correct. A row is the most expensive test in the project and the worst debugger. **Use instead:** `pwsh -File lab\Test-LabScripts.ps1` before any lab run — it parses everything and reports any function reading a variable nothing declares. Return `, @()` rather than `@()` from a helper whose empty result is a legitimate outcome. Treat a lab run that is not `OK` as an outcome to report and stop on, never as a base to keep reading evidence off: `BUSY` means nothing ran, so every reader after it is interpreting an absence. **Prevented by:** `lab/Test-LabScripts.ps1` — the undeclared-variable class, and the case-collision class beside it, which cost a second run before it was guarded: an assignment whose name matches a parameter of its own scope in a different spelling is that parameter, and the value the caller supplied is gone; `Test-LabRunUsable` ends a recovery row on a lab run that produced no evidence; `ConvertTo-TigerSetupFlattenedChecks` rejects an object that is not an entry-point run; and both row drivers record the failing statement and the script stack in an `ERROR` summary, so the next attempt starts from the line rather than from the message. **Generalization candidate:** TigerAiCore or a shared Lab — "a driver whose each attempt costs minutes needs a static gate and an error record that keeps its position" is general to any expensive-loop harness. The unrolling trap is not merely ordinary language knowledge either: the same mistake, in the same shape, is what TigerWinLab's R12 turns out to be (`TigerWinLab-Requirements.md`), and reading the helper is what hides it — `$list.ToArray()` is never `$null`, but a function returning it *when the list is empty* hands the caller `$null` all the same. **The other half of that trap, which cost a second matrix pass:** `, $list` fixes the helper and moves the defect to its callers. A helper that emits its collection as one object is only correct for callers that *assign* it; a caller that enumerates — `@(& $body)`, or piping into the property form `Where-Object status -eq 'FAIL'` — now receives one `Object[]` instead of the checks, and reports a count of 1 however many there were. Under `$ErrorActionPreference = 'Stop'` the property form throws `The input name "status" cannot be resolved to a property`; under the default preference it silently matches nothing, which is worse. That is R15, and it is why TigerSetup's own `Get-JobLog` was checked rather than assumed: every one of its callers assigns, so the convention is safe *there*. **Both ends of a collection boundary have to agree, so a return convention is only verifiable at its call sites** — and a regression test that exercises the helper alone, as the lab's did for 0, 1 and 3 items, passes either way. TigerWinLab now normalizes at the boundary that had the mismatch — every scenario collects its phase body through `ConvertTo-CheckArray` — and tests the public `Invoke-Phase` with each shape a body can produce rather than the helper alone. **The third shape of the same trap, which cost a feature-rows pass on 2026-09-19:** a conditional unrolls too. `$x = $(if ($Short) { @() } else { @('a') })` binds a *string* on the else branch, and `$x + @('b')` is then string concatenation (`ab`), not a two-element array — the check compared `a|b` with `ab` and failed on every step. Likewise `@(Get-Member2 $doc 'actions')` for a member the document omits is `@($null)`, one element, not none. Wrap a conditional that yields a collection in `@(...)`, and filter `$null` out of anything built from an optional member. Neither is caught by `Test-LabScripts.ps1`, which parses; a ten-line `pwsh` probe of the expression with each branch is, and costs seconds where the row costs minutes. ## A graceful Restart Manager shutdown leaves a process without a message loop running **Area:** engine, quiescence around a transaction **Status:** Active **Symptom:** `RmShutdown` returned `ERROR_SUCCESS` for a console process holding a payload file open, yet the file was still locked and the process was still running. Trusting that return value would have let the transaction open and then fail halfway through replacing the file. **Cause:** without `RmForceShutdown`, the Restart Manager asks and does not insist. An application with no message loop never answers, so it is neither closed nor terminated, and the call still reports success. **Do not:** treat a successful `RmShutdown` as evidence that the files are free, and do not reach for `RmForceShutdown` to make it so — forcing a running application to die is exactly what the design refuses to do. **Use instead:** decide from who still holds the files, over a bounded grace period rather than in one look — the complement of "success does not mean they are gone" is "still there does not mean they refused": an application with a message loop needs a moment to answer the request, save and exit, and concluding at once reports it as a holder that would not close. Ask that question of the machine, not of `RmGetList`; see the entry below for why. **Prevented by:** `restart::Quiescence::acquire` waits through `wait_for_holders_to_go`, and `quiescence.rs` asserts that an upgrade blocked by a console holder either succeeded outright or changed nothing. **Generalization candidate:** none — this is Restart Manager behaviour, and TigerSetup is where it matters. ## `RmGetList` answers with what the session was told, not with what is running **Area:** engine, quiescence around a transaction **Status:** Active **Symptom:** M6 exited 6 `package_in_use` after the full grace period, twice, against an application the Restart Manager had listed as a *main window* application and had asked to close. Lengthening the grace from fifteen seconds to sixty changed only how long it took to say so. Every other signal said the arrangement was right: the application had a window, it ran on the signed-in user's desktop, and the elevated upgrade ran beside it. **Cause:** the wait re-listed the holders with `RmGetList`, and `RmGetList` answers with the applications the Restart Manager session knows about — the list taken when the resources were registered — rather than with a fresh look at the machine. It keeps naming a process that has already exited. A wait that ends when that list empties therefore never ends, and **every** upgrade over a running application was going to report `package_in_use` however promptly the application closed. The application had been closing all along; nothing in the run could see it — asked the other way, the same application in the same row closes 379 ms after the request, and M6's upgrade went from a sixty-three second refusal to a three-and-a-half second success. **Do not:** treat `RmGetList` as an observation of the machine, and do not read a holder still listed after a shutdown as an application that refused. **Use instead:** ask Windows about the holder's process. The Restart Manager gives a `RM_UNIQUE_PROCESS` — a process id together with the moment it started — which is exactly the identity needed, because Windows reuses a process id as soon as the process that had it is gone. A handle that will not open because this run may not look at that process is not evidence that it ended: an unelevated run keeps waiting rather than concluding that another account's application has closed. **Prevented by:** `restart.rs::a_holder_that_goes_away_by_itself_ends_the_wait` — a holder that closes the file and exits on its own, with no shutdown to attribute it to, must end the wait. It runs in three seconds and fails in thirty against the old code; the whole class was invisible to the tests that existed, because the only holder they used was one that never closes. **Generalization candidate:** none — this is Restart Manager behaviour, and TigerSetup is where it matters. ## A quiescence row must run the application where a person would run it **Area:** lab rows, M6; any row that asks the Restart Manager to close a GUI application **Status:** Active **Symptom:** M6 failed identically every time it was run — the upgrade exits 6 `package_in_use`, the row reports "the upgrade log records no shutdown", and the installed version stays at the old one. Two distinct causes produced that one verdict, and fixing the first changed the counts by two and nothing else, which reads as "the fix did nothing" and is not what happened. **Cause:** the row started TigerMarkView with a plain guest-job command, and those run as the job account in the lab's non-interactive session. The Restart Manager lists a holder from its open file handles, which crosses Windows sessions, but it closes a GUI application by messaging its windows, which does not. An application started that way has no interactive desktop to be messaged on, so it was asked and never answered — the same shape as the console holder above, reached by a different route. Underneath it, and invisible until the arrangement was right, was the engine's own defect: it asked `RmGetList` whether the holder had gone, and `RmGetList` never says so (see above). **What makes this hard to read:** the engine emits `[restart_manager_shutdown]` only once the holders have actually gone, so the check that looks for it reports "no shutdown" for a shutdown that was requested and honoured. And "still running after the grace period" reads identically for an application that refused, one that was never asked, and one that had already closed. **Do not:** read `package_in_use` in a quiescence row as evidence that the product's Restart Manager handling is broken, and do not stop at the first plausible cause because the counts barely moved when it was fixed. **Use instead:** make the row say where each command ran and what Windows could ask of the holder, then read those rather than the verdict. The Restart Manager's own `RM_PROCESS_INFO.ApplicationType` is what separates "the request that works was made" from "no window was ever found to ask", and it costs nothing to record. **Prevented by:** the engine records the classification beside every holder (`restart.rs::classify`, in the log line and in the in-use message) and derives the grace period from it — `restart.rs::a_windowed_holder_is_given_longer_than_one_windows_cannot_ask`; M6 asserts that the application ran on the signed-in user's desktop, that the elevated upgrade ran beside it, that the application presented a window, and that the Restart Manager listed it as one it could ask. **Generalization candidate:** none — the Windows behaviour is general, but what it cost was a TigerSetup row's ability to evidence its own requirement. ## A Restart Manager shutdown kills the console it is asked from, not just the holder **Area:** tests, quiescence; any harness that spawns a console process the Restart Manager will be asked to close **Status:** Active **Symptom:** `cargo test --workspace` killed the terminal it was started from. Running it from an agent session closed that session outright — twice — with no error, no exception, no crash dialog and no entry in any event log. The test run simply stopped existing partway through `quiescence`. **Cause:** the Restart Manager closes a **console** application by delivering a console control event to it, and a console control event goes to every process attached to that console rather than to the one process being closed. `quiescence`'s `Holder` spawns `powershell.exe` to hold a payload file open, and it was spawned with no creation flags — so it inherited the console of the test runner, which is the console of whatever started the suite. When the engine asked for the graceful shutdown the quiescence path exists to perform, Windows delivered the event to that whole console: the holder, the test binary, `cargo`, the shell, and the agent session hosting them all died together with `STATUS_CONTROL_C_EXIT` (`0xC000013A`, seen as exit code `-1073741510`). The silence is the trap. A console control event is not a fault, so there is nothing to report: no `Application Error`, no Restart Manager event, no Defender detection, no resource exhaustion. Every log is clean and every plausible cause — a crash, memory, the disk, antivirus, the engine's own `AttachConsole` — is innocent. What identifies it is the **exit code of the run**, and the run has to be observed from outside its own console to have one. **Do not:** spawn a process the Restart Manager will be asked to close without giving it a console of its own, and do not diagnose a silently vanishing test run by reading event logs — this failure writes to none of them. **Use instead:** `CREATE_NO_WINDOW | CREATE_NEW_PROCESS_GROUP` on any such child, which keeps it a console application with no message loop — the thing the quiescence path is actually being tested against — while containing the event to a console nothing else shares. More generally, when a test run disappears without a diagnostic, re-run it detached from the session's own console (`Start-Process` with redirected output) and read its exit code: an observer inside the blast radius cannot report what killed it. **Prevented by:** `quiescence.rs::HOLDER_ISOLATION`, applied where the holder is spawned. **Generalization candidate:** TigerAiCore — "an agent session must not share a console with a child that will be signalled, and a run that vanishes without a diagnostic must be re-observed from outside its own console" is method knowledge about running expensive suites from an agent session, not a fact about installers. ## Windows rewrites an access control list in its own spelling **Area:** engine, machine-scope state directory **Status:** Active **Symptom:** A DACL written as `D:(A;OICI;FA;;;SY)(A;OICI;FA;;;BA)(A;OICI;0x1200a9;;;BU)` reads back as `D:P(A;OICI;FA;;;SY)(A;OICI;FA;;;BA)(A;OICI;FRFX;;;BU)`. Comparing the two as text says the list drifted, so every run would rewrite it and no run could ever detect real tampering. **Cause:** the SDDL writer renders an access mask with the abbreviations that add up to it (`FR|FX` is `0x1200a9`), reorders nothing but reformats everything, and adds the header flags the object actually carries. **Do not:** compare access control lists, or any other Windows security descriptor, by string equality or substring search. **Use instead:** parse both sides into entries and compare rights as numbers (`win::acl::Dacl`). **Generalization candidate:** none — it belongs with the code that reads DACLs. ## `%ProgramData%` is writable by a standard user, so it cannot decide elevation **Area:** engine, elevation **Status:** Active **Symptom:** A machine-scope run decided it needed no administrator because it could create files under `%ProgramData%`, and then failed creating its install root under `%ProgramFiles%` — after the elevation prompt was no longer on offer. **Cause:** the default access control list of `%ProgramData%` grants `CREATOR OWNER` and lets any user create files there. Only `%ProgramFiles%` is closed to a standard user from the start; the state directory becomes closed later, once TigerSetup writes its own list on it. **Do not:** decide that a run needs elevation from the scope alone, or from a single root. **Use instead:** probe every root the run will write — the state directory and the install root — and require elevation if any of them refuses (`elevation::required`). **Generalization candidate:** none — the roots are TigerSetup's. ## A known folder can move under a live installation **Area:** engine, ownership and scope confinement **Status:** Active **Symptom:** A user-scope installation could not be uninstalled at all. The run refused with `path_outside_root: stored shortcut ...\Desktop\App.lnk is outside ...\Start Menu\Programs and C:\Users\\OneDrive\Desktop`, before opening a transaction, so the product stayed installed with no way to remove it. **Cause:** confining stored resource paths to the scope's own roots is right, but it was applied to shortcuts as a reason to refuse the whole run. A hive cannot move, so a registry key outside the scope really does mean the database is wrong about the machine. A shortcut folder is different: OneDrive's Known Folder Move relocates the desktop, and policy can redirect the Start Menu, so a link recorded before the move is legitimately outside the folders the scope resolves afterwards. The check could not tell the two apart and treated both as tampering. **Do not:** treat "this stored path is outside the roots I resolve now" as proof of tampering for any resource whose location Windows may move, and do not let a conservative refusal end in a state the user cannot get out of. An installation that cannot be uninstalled is a worse outcome than a resource left behind. **Use instead:** decide per resource kind. Where the location cannot move (registry hives, the install root), refuse the run. Where it can, do not touch the resource, report it with its own code (`shortcut_outside_scope_preserved`) and finish removing everything else. Both answers leave the out-of-scope resource untouched, which is what the confinement exists to guarantee. **Prevented by:** `plan::reconcile` takes the scope's shortcut folders as an input and preserves a link outside them; `resources.rs::a_shortcut_whose_folder_moved_is_preserved_and_the_rest_is_removed` moves the folder and asserts the uninstall completes; `machine_scope.rs::a_shortcut_row_outside_the_scope_is_preserved_and_never_deleted` asserts a tampered row is still never acted on. **Generalization candidate:** the principle — separate "cannot legitimately differ" from "may legitimately differ" before treating a mismatch as an attack — is general, but the resource kinds are TigerSetup's. ## Locking a directory down is not the same as owning it **Area:** engine, machine-scope privilege boundary **Status:** Active **Symptom:** An independent review traced a working confused-deputy path through code that looked correct: the machine-scope state directory was given an explicit access control list granting only SYSTEM and Administrators, yet a standard user could still replace the `uninstall.exe` inside it that Add/Remove Programs later runs elevated. **Cause:** two facts that are individually unremarkable and together decisive. `%ProgramData%` grants `BUILTIN\Users` the right to create subdirectories, so a standard user can create the product's state directory before any install ever runs — and the creator owns what it creates. An object's **owner** keeps `WRITE_DAC` no matter what the list says, so it can hand the rights back to itself at any later moment. Writing a DACL onto a directory somebody else owns protects it only until that owner objects. `create_dir_all` also adopted a pre-existing directory without ever asking who owned it. **Do not:** treat "I set the access control list" as "this object is mine", and do not stage an executable you are about to run elevated in a directory the invoking user can write — under same-account elevation `%TEMP%` is still that user's own folder, with inheritable full control over anything created in it. **Use instead:** set the owner with the list (`OWNER_SECURITY_INFORMATION`, an `O:BA` prefix in the SDDL), read the owner back before trusting a directory that already exists, and take ownership when it is someone else's. For a binary about to be executed elevated, create a fresh directory under a system-owned root with an unpredictable name and protect it before writing into it. **Prevented by:** `scope::protect_state_directory` checks owner and list and reports `state_directory_ownership_claimed` when it takes a directory over; `lib::staging_directory` refuses to stage an elevated relaunch in the user's `%TEMP%`. **Generalization candidate:** the principle is Windows-wide and belongs in whatever shared guidance covers privileged Windows code — an installer is simply where it bites first. ## Coordination is not a precondition: an unavailable service must not fail the run **Area:** engine, quiescence **Status:** Active **Symptom:** A lab row installed nothing and reported `restart_manager_failed: the Restart Manager could not start a session: The system cannot write to the specified device`. The run had not reached its transaction; the machine was left untouched but the product was not installed. It happened on a guest that had just been power-cut and restarted — that is, during the recovery scenario the row exists to test. **Cause:** `RmStartSession` was treated as a step that must succeed. Restart Manager is a service, and a service can be unavailable, most plausibly right after an unclean boot, which is exactly when a recovery run happens. The engine turned "I could not ask who holds these files" into "I refuse to install". **Do not:** let an advisory mechanism become a precondition. Quiescence exists to make replacing a file in use *pleasant*; it is not what makes the transaction safe. The journal is. **Use instead:** report the mechanism as unavailable (`restart_manager_unavailable`) and continue. A file genuinely held open then fails its own operation, and the transaction rolls back — a recoverable outcome the engine already handles — instead of a run that will not start at all. **Prevented by:** `restart::Quiescence::acquire` degrades to `none()` when the session cannot be started, files cannot be registered, or holders cannot be listed. **Generalization candidate:** the principle is general — decide for every external service whether it is load-bearing or advisory, and make an advisory one fail open — but the mechanism is Windows-specific. ## Windows writes the registry back lazily, so a journal can outlive its own mutation **Area:** engine, durability of registry operations **Status:** Active **Symptom:** After a power cut during an upgrade, recovery reported nothing to do and exited in 0.3 s, but `verify` failed with `registration_modified` on `DisplayVersion`: the state database said 0.8.2 and the machine's Add/Remove Programs entry still said 0.8.1. The transaction had committed, so nothing would ever reconcile the two — a same-version re-run short-circuits to `already_installed`. **Cause:** files are flushed as they are written (`win::fs`), but registry values were not. `RegSetValueEx` returns once the change is in memory; Windows writes the hive back when it chooses. So a power cut could take a value the journal had already marked applied, and the transaction that committed afterwards recorded state the machine did not have. The invariant the design states — durable undo before the mutation — was upheld; its unstated other half, *the mutation durable before the record of it*, was not. **Do not:** assume a Windows API that returned success has written anything to disk. `RegSetValueEx`, unlike a flushed file handle, has not. **Use instead:** flush the hive once before the commit, not once per value — `RegFlushKey` flushes the hive file its key belongs to. A cut before the flush leaves the transaction open and recovery reconciles it; a cut after it finds the values already on disk. One call per transaction is affordable; one per value is not. **The same lesson twice more, on 2026-09-20, once the recovery rows ran on a transaction fast enough to expose it:** `install-reboot-before-commit` and `upgrade-poweroff-prepared` came back with `firewall_rule_missing` — the rule the Windows Firewall policy object had accepted lived in the `SYSTEM` hive's memory only, and the reset took it; in the second row it took a rule the *prepare* install had committed more than a minute earlier, so nothing about elapsed time can be relied on. And `HKLM` is not a hive file but a master key over `SOFTWARE` and `SYSTEM`: flushing the predefined `HKLM` handle flushes neither, so the machine scope's own values were never flushed either. Every mutation that lands in a hive file the process did not write itself — a firewall rule — is flushed at the mutation, and the commit flushes each hive file the scope writes by a key inside it. **Prevented by:** `txn::Executor::commit` calls `win::registry::flush` before `journal::commit_*` — `HKCU` for the user scope, `HKLM\SOFTWARE` and `HKLM\SYSTEM` for the machine scope — and `win::firewall::Store::put` and `remove` flush the firewall policy's hive before they return; lab row M5b (power cut during an upgrade) verifies the registration afterwards, and the `install-reboot-before-commit` and `upgrade-poweroff-prepared` rows verify the firewall rule. **Generalization candidate:** the question "does success mean durable?" is worth asking of every mutation an installer makes, on any platform. ## A vendor's installer may report failure and still have done the job **Area:** engine, dependency acquisition **Status:** Active **Symptom:** Two Windows 10 rows failed to install the product at all. The log showed TigerSetup doing everything right — 258 MB downloaded from Microsoft's own URL, SHA-256 verified, the vendor's installer run with `/silent /install` — and then `dependency_install_failed: the installer exited with code -2147219970`, a clean rollback and exit 3. The very next run's log records `dependency_detected Microsoft.EdgeWebView2Runtime: version 152.0.4191.66`. The runtime was there. The installer had succeeded and said otherwise. **Cause:** the exit code was treated as the verdict. For a *resource* that would be right; for a *requirement* it is not. `TigerSetup-Design.md` §7.2 separates identity, detection, acquisition, installation and verification precisely because the question is "is the requirement met on this machine", and only detection answers that. Microsoft's WebView2 evergreen bootstrapper returns vendor-specific codes for conditions that are not failures. **Do not:** let a third-party installer's exit code be the last word on whether a dependency is present, in either direction. Success already had to be confirmed by detection (`dependency_unverified` exists for an installer that claims success and leaves nothing); failure deserves the same check. **Use instead:** on a non-zero exit, detect again. If the dependency is now present the requirement is met — report `dependency_installed_despite_exit_code` and continue. If it is absent, the failure stands. **Prevented by:** `dependency::install` re-detects before believing a failing exit code; lab rows S3, S5 and S6 are the ones that exercise acquiring a genuinely absent WebView2, on Server 2019, the one baseline without it. **Generalization candidate:** the shape is general — when something else performs the work, verify the world rather than trusting the report — but the detectors are TigerSetup's. ## A wizard row's keys are aimed at a page number, and pages move **Area:** lab rows, wizard capture; any row that drives a wizard by keystroke **Status:** Active **Symptom:** the interactive dependency row (W6 then; S6 since the WebView2 premise moved to Server 2019) captured one page and then died with TigerWinLab's Desktop Capture reporting `The handle is invalid`. Every other signal said the wizard was fine: the lab's `wait-window` found the window, `hit-test` answered from it, and the window's bounds were unchanged. Refreshing the guest support scripts changed nothing, and a secure-desktop transition stayed a hypothesis for two sessions. **Cause:** the row sent Alt+A to page 1 to accept a licence. Page 1 is the **scope** page, whose second choice is `Install for &all users` — so Alt+A selected the machine-scope install, Enter asked for elevation, and Windows put the UAC prompt on the secure desktop. Nothing on the secure desktop can be photographed by a session that may not open it, so the very next capture failed with a bare Win32 message, six steps away from the keystroke that caused it. The row is the one row of the matrix whose whole point is acquiring a dependency *without* elevation. **Do not:** number a wizard's pages in a row and aim keys at the numbers without evidence of what each page offers, and do not read a capture failure as a capture defect — the window is still there and still enumerable while the desktop in front of it is one the session cannot read. **Use instead:** read the page's own UI Automation tree, which every capture already writes beside the screenshot, and derive the keys from what the page offers. The wizard's control identifiers are grouped by page and stable (`ui::window`'s `ID_SCOPE_*`, `ID_LICENSE_*`, …), so `controls` on each captured page names which page it was, in every language. TigerWinLab now names the input desktop when a capture fails there, so the same mistake reads as an elevation prompt rather than as a broken handle. **Prevented by:** the page record carries its control identifiers, so a row's key plan can be checked against what was actually in front of it; S6's plan answers the scope page with Alt+M and the licence page that follows with Alt+A. **Generalization candidate:** none — the Windows fact (a secure desktop cannot be captured) is TigerWinLab's and is now documented there; aiming keys at page numbers is a TigerSetup row's own discipline. ## A page that is working looks exactly like a page that was not answered **Area:** lab rows, wizard capture **Status:** Active **Symptom:** the interactive dependency row (W6 then, S6 now) could not have completed even with the right keys. Its wizard budget was ten pages and its loop advanced every 800 ms, so the dependency download — minutes on one page — would have exhausted the budget and killed the wizard part-way through installing WebView2. **Cause:** the wizard deliberately keeps its forward control visible and disabled while it is busy (`ui::window::update_buttons`), so a key pressed then is lost rather than queued. A capture loop that advances on a timer therefore photographs the same page repeatedly and calls each one a page. **Do not:** give such a row a bigger page budget or a longer settle delay. Both are guesses about how long a download takes, and the row still cannot tell "still working" from "not answered". **Use instead:** wait for the page itself to change. A wizard keeps one window with one title and rewrites a progress page's label as it works, so the stable identity is the set of control identifiers the page shows; `lab/guest/Invoke-WizardCapture.ps1` waits for that set to change or for the process to exit, bounded by the wizard's `pageTimeoutSeconds`. A page that never changes then ends the wizard naming what it was still showing, which is a wrong key rather than a page budget quietly running out. **Prevented by:** the wait itself; S6 captures seven distinct pages — scope, licence, destination, options, ready, progress, finish — and its process ends `exited` rather than `killed`. **Generalization candidate:** possibly TigerWinLab — its own `Invoke-WizardRun` already waits for a *named* advance control, which is the label-driven form of the same rule. Promote only if a second consumer needs the language-independent form. ## Ending a lab session and reclaiming its VM are two transitions **Area:** lab drivers, Lab sessions, background-work accounting **Status:** Active **Symptom:** A full matrix run was killed by the operating system for want of memory, part-way through its first row. Nothing about that run was wrong: a VM from the *previous* run had been left powered on, and two guests plus the host's own work did not fit. **Cause:** the earlier run ended its Server session correctly and the lab reported, per resource, that ending it had **not** reclaimed the VM — `cleanup: Failed`, `Shutting down VM … failed: … Access is denied. (0x80070005)`. The driver discarded that report. Ending a session is authoritative and stays ended; stopping the VM it released is a separate transition, and a stop the guest is allowed to refuse is one it can fail. **Do not:** treat "the session ended" as "the VM was released", and do not discard the per-resource report because the rows all passed. A VM left running costs host memory and one of the host's few running slots, so the run that pays is the next one — killed, or refused a baseline it cannot start. **Use instead:** read the report the close writes and name anything the lab did not reclaim (`LEFT RUNNING:`). Completion accounting reads `Get-TigerWinLabSession.ps1`, which answers from the lab's own records: no process list shows a VM. **Prevented by:** `Exit-TigerSetupLabSession` reads the close report and prints every resource whose cleanup is not `Completed`, `NotRequired` or `Protected`. The refusal that produced this entry no longer happens: TigerHyperLab's stop is a shutdown the guest cannot veto, which is what a Windows Server guest was doing. **Generalization candidate:** none — the provider now owns the refusal, and this entry keeps only the accounting rule that outlived it. ## A shortcut's target reads back in Windows' own spelling: drive letter and long names **Area:** process-level tests; any assertion comparing a path the engine wrote against a path a Windows API read back; the engine's shortcut ownership checks **Status:** Active **Symptom:** the whole `durability` suite (and every other suite that calls `assert_verified`) failed on one build and passed on the next with no source change, on `assertion left == right` where the two paths differed only in the drive letter: the link's target was `C:\…` and the expected install root was `c:\…`. **Cause:** `assert_verified` compared the Start Menu link's target byte-for-byte against `install_root().join(...)`. Windows canonicalises the drive letter of a link's target when it saves the `.lnk`, so the target always reads back upper-case; the expected side is the test's temporary directory, whose drive-letter case is whatever `CARGO_TARGET_TMPDIR` had **baked in at compile time** — which flips with how cargo was invoked (a `c:\` path from Git Bash, a `C:\` path from PowerShell). So the assertion passed only when the compile happened to bake an upper-case drive. The engine's own `verify` never had the problem: it compares shortcut targets with `eq_ignore_ascii_case`. **Do not:** compare a Windows path the OS round-tripped against one the test assembled with `==`, and do not chase a pass/fail that flips between runs as a real regression before checking whether only the drive-letter case differs. **Use instead:** compare Windows paths case-insensitively, at the same contract the engine holds (`eq_ignore_ascii_case`). Windows paths are case-insensitive; a case-sensitive comparison is the defect. **Prevented by:** the helper now compares the shortcut target case-insensitively, so the assertion no longer depends on the drive-letter case the build directory happened to carry. **The same, for short names:** the first hosted CI run failed three engine tests because a GitHub runner's `%TEMP%` is `C:\Users\RUNNER~1\…`. `IShellLink` stores the target as the item it parsed, so a target written through a short spelling reads back long (`C:\Users\runneradmin\…`) while the icon and working directory keep the spelling they were written with. A case-insensitive comparison still called the link modified, and the uninstall planner's text-prefix check preserved a link it owned. Case folding is not enough: the engine compares a link's path fields with `resource::shortcut::same_path`, which also expands 8.3 names (`win::fs::long_form`, a spelling and never a resolution — junctions stay junctions), and `targets_install_root` compares long forms. Reproduce it locally by pointing `TMP`/`TEMP` at a short spelling of a real directory; the tests `a_short_spelled_target_reads_back_as_the_same_path` and `a_short_and_a_long_spelling_of_one_path_are_the_same_path` do it on any volume that generates short names. **Generalization candidate:** none — it is specific to paths that cross a Windows API, and the engine applies it in one place. ## An interrupted wizard test leaks a GUI installer that poisons a later build **Area:** process-level wizard tests; background-work ownership **Status:** Active **Symptom:** every `wizard` test failed at the very first step — building the shared fixture — with a bare `io_error` "Access is denied. (os error 5)", six steps removed from anything the failing test does. The same fixture code built fine for every other suite. **Cause:** a `Setup.exe` a wizard test had started (an interactive run left mid-flight when a `cargo test` was moved to the background and not awaited to a terminal state) stayed running, holding a handle to its installer file under `target\…\tmp\fixture-wizard\out\`. The next `wizard` run's fixture build does `remove_dir_all` on that directory and could not delete the open file, so the build failed — in a suite, and at a step, that had nothing to do with the leak. **Do not:** move a wizard `cargo test` to the background and start other work without driving it to completion or killing it; a wizard test spawns real GUI `Setup.exe` processes (and, for a scope choice, a child process), and an abandoned one keeps a window and a file handle alive. Do not read the resulting access-denied as a build bug. **Use instead:** account for wizard-spawned processes as task-owned background work (`AGENTS.md`, and *A finished process is not a finished task*): before reporting or re-running, sweep for `Setup.exe` / `-Setup` / the engine by the directory they run from, not by a fixed name list, and stop any that a prior run left behind. `Get-Process … | Where Path -like '…\fixture-*\out\*'` finds them. **Prevented by:** nothing mechanical yet — a wizard run interrupted from outside the test cannot clean up after itself. The discipline is to await or kill a backgrounded wizard test, and to sweep before trusting a fresh run. **Generalization candidate:** none — it is the general background-ownership rule applied to this project's GUI tests. ## A wizard page is seen before its buttons are, and a synthetic run outruns the reader **Area:** wizard tests, wizard automation **Status:** Active **Symptom:** `wizard.rs`'s install test, green for weeks, failed twice in a row on "a busy page keeps its forward button visible but disabled": it had waited for the progress page and then read the Next button once, and found it enabled. The wizard, when looked at, was sitting on its finish page with the install complete. **Cause:** two things that are each fine alone. `enter_page` shows the new page's controls first and updates the buttons after — a few statements later on the same thread, so a click can never land in between, but a reader in another process, which reads window state directly, can. And a synthetic install finishes in well under a second, so the reader that missed the busy state may already be looking at the finish page, where the forward button is enabled by contract. A single sample taken the instant a page marker appears answers neither the page's contract nor the run's state. **Do not:** assert a page's button state from one read taken right after its marker became visible, and do not read a fast run's transient page and expect it to still be there. **Use instead:** what the cancel test already does — wait for the settled contract (`wait_until(… !enabled(ID_NEXT))`, as it waits for Cancel), and hold the run at a fault point (`--fault after_prepare@20:hold:`) when the page would otherwise be gone before the reader arrives. A held run makes "the button stays enabled" a real violation rather than the finish page arriving. The same holds for any automation that drives the wizard from outside: wait for the page *and* the button state it wants, never for the page alone. The same reader could also see a page with *some* of its controls: `enter_page` used to show a page's controls one `ShowWindow` at a time, and a test that acted only on the controls it found visible right after the page's marker skipped a radio button that appeared a moment later — once in three runs. Showing and hiding a whole page is one `BeginDeferWindowPos` … `EndDeferWindowPos` operation now, so a page is either not there or complete. Complete is not settled, though: the finish page appears with all its controls, and only then does `show_outcome` fill the body and hide the launch offer for a run that did not install. A test that read "a run that rolled back offers nothing to start" the instant the body became visible saw the offer, once in several runs. Every test now waits for the finish page with `wait_for_finish`: the body visible, Cancel gone and Finish enabled, which `update_buttons` sets only after the outcome has been applied. **Prevented by:** the install test in `crates/tigersetup-setup/tests/wizard.rs` holds its run and waits for the contract; `Wizard::enter_page` switches pages atomically with `DeferWindowPos`; `wait_for_finish` is the one way the wizard tests wait for the finish page. **Generalization candidate:** the observation half — a cross-process GUI reader sees intermediate states the UI thread never exposes to input — may belong with TigerWinLab's wizard driver; it is recorded here until a lab row shows it. ## A COM apartment must outlive the interfaces it owns, and only the release build says so **Area:** engine, hand-declared COM (`win::firewall`, `win::shortcut`); anything with `CoInitializeEx` and raw interface pointers **Status:** Active **Symptom:** the first lab row of the 0.6 batch died inside the guest with exit code `-1073741819` (`STATUS_ACCESS_VIOLATION`) between `uninstaller_written` and `transaction_started`, in a release engine whose every unit test was green. The same test binary rebuilt with `--release` crashed on the host too, at the end of the test, after every call had already succeeded. **Cause:** Rust drops a struct's fields in declaration order. `Policy` held an `Apartment` (whose `Drop` calls `CoUninitialize`) *before* its two interface pointers, so the apartment was torn down first and the two `Release` calls that followed ran against a COM runtime that was no longer there. The debug build survived that by accident; the optimised one did not. The shell-link wrapper has the same field order and never hit it, because its explicit `Drop` releases the interfaces before any field is dropped. **Do not:** trust a debug-profile run of hand-written COM code, and do not let field order carry a lifetime rule silently. **Use instead:** declare the apartment last (or release explicitly in `Drop`), say why in a comment beside the struct, and run the Win32 layer's unit tests in the release profile too — `cargo test -p tigersetup-engine --release --lib win::` takes about a minute and is part of the gate. **Prevented by:** `Policy`'s field order and comment in `crates/tigersetup-engine/src/win/firewall.rs`; the release-profile test run in the verification gate (`AGENTS.md`). **Generalization candidate:** the rule is Rust-general — a guard that owns a runtime must be declared after everything that needs the runtime — and worth TigerAiCore's attention only if a second project meets it. ## An elevated process does not see a per-user `App Paths` entry **Area:** lab acceptance of `[[app_paths]]`; any check that resolves a bare executable name through the shell **Status:** Active **Symptom:** the feature rows registered `App Paths\TigerSetupTestApp.exe` under `HKCU` exactly as declared, and the same job's `ShellExecute` of the bare name answered `ERROR_FILE_NOT_FOUND` (2) after every step, in a hive the job itself had just read the key from. **Cause:** the job account is an administrator in session 0 and the collector ran in its elevated process. Windows does not consult `HKCU\...\App Paths` from an elevated process — a per-user hive must not be able to redirect an administrator's command — while an `HKLM` entry resolves from any session. One guest experiment (the probe from the job, from the signed-in standard user and from the elevated agent, against `HKCU` and `HKLM`) showed all three at once; the host had hidden it because a developer's console is not elevated. **Do not:** read a per-user shell registration from an elevated probe, nor conclude from such a probe that the registration is broken. **Use instead:** run the probe as its own guest command in the session whose registry the installer wrote (`New-ShellProbeCommand` in `lab/Invoke-FeatureRows.ps1`), only where Windows consults that registry — unelevated for `HKCU`, any session for `HKLM` — and let the elevated per-user rows prove the key alone. **Prevented by:** `New-ReadCommands -ProbeRunAs` and the row's `$probeRunAs` rule in `lab/Invoke-FeatureRows.ps1`; `TigerSetup-Validation.md` §5.3. **Generalization candidate:** the same rule applies to per-user file and protocol associations, which an elevated process also ignores; Tiger projects that check a shell registration from a lab should know which session they ask. ## The dark visual style dims a radio button's label **Area:** wizard, dark theme (`ui/theme.rs`); any Win32 UI moved onto `DarkMode_Explorer` **Status:** Active **Symptom:** the Windows 11 dark capture of the options page showed the choice option's radio labels in a dim grey next to check-box labels in the palette's white — legible only just, and unmistakably not one control family. **Cause:** with `SetWindowTheme(hwnd, "DarkMode_Explorer")` a check box draws its label in the colour the parent answers `WM_CTLCOLORSTATIC` with, but a radio button draws its label in the style's own text colour, which is the dim one. The same window, the same message handler and the same font produced two label colours, and nothing on the light side ever showed it. **Do not:** expect the dark style to colour every button class alike, or accept the dim label as "what Windows does" — the check boxes beside it prove otherwise. **Use instead:** let the button keep its behaviour and its glyph and paint the label yourself in dark mode: subclass the radio, fill with the parent's `WM_CTLCOLORSTATIC` brush, draw the glyph with the style's `BP_RADIOBUTTON` part in the control's state, draw the text in the DC's colour, and the focus rectangle where `WM_QUERYUISTATE` allows one (`theme::paint_radio`). **Prevented by:** `theme::apply_to_control` subclassing every radio button; the dark options-page capture in the Windows 11 UI matrix (`TigerSetup-Validation.md` §8), which is where this was seen. **Generalization candidate:** yes — any Tiger Win32 UI that uses radio buttons in dark mode meets the same colour; worth a note beside the shared desktop-experience rule if a second project does. ## A third-party installer's process is not its lifecycle, and its silent switch is not optional **Area:** the lab's guest command runner (`lab/guest/Invoke-SetupCommands.ps1`) and any row that runs an installer TigerSetup did not build — the installer-technology benchmark (`benchmark/`) **Status:** Active **Symptom:** the benchmark's prototype run "completed" a WinMerge row, then its NSIS row sat for the whole command timeout while a shell stayed alive; the prototype's NSIS uninstall times were about one second and the harness had grown a fixed 15-second sleep before reading the machine. **Cause:** two unrelated facts about NSIS, each invisible from the harness's own vantage. The NSIS rows for the per-user applications were started with `/CurrentUser /D=…` and no `/S`, so the wizard opened on the desktop-less job account and waited for a click nobody could give; the job's own output is printed only when it ends, so "Waiting for PowerShell Direct" was the last thing seen. And an NSIS uninstaller copies itself to `%TEMP%\~nsuX.tmp\Au_.exe` and exits at once while the copy does the work, so the lifetime of `Uninstall.exe` measures nothing and the machine read right after it is read mid-uninstall — the sleep was papering over that. Inno Setup's `unins000.exe` first phase waits for its second phase only until the second phase releases it with `WM_KillFirstPhase`, which comes before the install root is removed, so its lifetime is not the uninstall's either (see *An Inno Setup uninstaller's process and its registration both end before its uninstall does*). **Do not:** time or read after a foreign installer's process without knowing what it hands off to; add a sleep where a completion condition is missing; assume a technology's silent behaviour from another technology's switches (`/VERYSILENT` is Inno's, `/S` is NSIS's, `--quiet` is TigerSetup's); read "the job printed nothing for minutes" as a lab problem before checking what the guest command is waiting for. **Use instead:** name the completion the command means — the guest runner's per-command `waitForProcesses` (names started after the command) and `waitForAbsentPaths` (the install root) keep the command's clock running until they hold, bounded by the command's timeout, and record the process lifetime and the wait separately; MultiUser NSIS packages need `/CurrentUser`/`/AllUsers` on the uninstaller too (the benchmark writes the mode into `UninstallString`), and `MULTIUSER_INIT` overwrites `$INSTDIR`, so a `/D=` is kept only by saving `$INSTDIR` before it and restoring it after. **Prevented by:** the harness's row verdict requires the completion wait to be satisfied and the install root gone; a one-row smoke test before a campaign, into a results root of its own (`benchmark/README.md`). **Generalization candidate:** yes for the guest runner's completion options (any Tiger consumer driving a foreign installer through TigerWinLab meets the same hand-off); the NSIS facts stay with the benchmark. ## An Inno Setup 7 installer cannot be unpacked without running it **Area:** any corpus or benchmark that needs an application's real payload — the installer-technology benchmark (`benchmark/`) and the compression spike (`benchmark/compression-spike/`) **Status:** Active **Symptom:** the compression spike planned to take four payloads out of Inno Setup installers (GIMP 3.2.6, and the IT Tiger applications TigerMarkView, TigerWrap and TigerSqlCmd). `innoextract` 1.9, installed for the purpose, reported "Unexpected setup loader revision: 2 … Could not determine setup data version" on every one of them. **Cause:** all four are Inno Setup 7 installers (GIMP's build script pulls the latest Inno release; the Tiger installers are compiled with Inno Setup 7.1). No open unpacker reads the Inno 7 format: `innoextract` 1.9 stops at 6.2.2 and its unreleased master at 6.7.0, and 7-Zip has no Inno reader at all. An Inno installer is opaque until it runs, and running it on a developer machine or a lab VM to harvest `{app}` is a different, far more expensive kind of acquisition. **Do not:** plan a corpus on "extract the installer" for an Inno Setup application without checking the setup-data version string inside the executable first (`Inno Setup Setup Data (x.y.z)`, findable with a plain string search); assume a Windows unpacker exists for a format because it existed for the previous major version. **Use instead:** for IT Tiger applications, the staging tree the installer was compiled from (`WorkingDir\` beside the `.iss`, or the publish directory the script names), verified against the installer by its timestamp and version stamp, with the `[Files]` excludes applied by hand — that is the exact payload and needs no unpacker. For a TigerSetup-built application, the installer's own payload block (`tiger-setup inspect --output-zip`). For a third-party Inno 7 application with no archive artifact, either a lab install that copies `{app}` out, or leave it out with the reason recorded, as the spike did with GIMP. **Prevented by:** `results/corpus.json` in the spike records each source's kind and, for an installer, the reason it can or cannot be unpacked; the manifest is read before anything is downloaded. **Generalization candidate:** no — the fact is about Inno Setup 7 and belongs with the benchmarks that meet it. ## The first read of a file an installation just wrote is the scanner's **Area:** engine performance; any plan or check that reads files the same run or the previous run wrote **Status:** Active **Symptom:** the 0.7.1 uninstall of a 1,200-file, 553 MB installation spent 7.5 s before its transaction even opened, hashing every owned file at about 74 MB/s on an NVMe drive whose SHA-256 throughput is 300–700 MB/s. Reading the same files a second time took 0.4 s. Removing the hashing moved the same 6.6 s into the Restart Manager check, which had been 155 ms. **Cause:** Windows Defender's real-time protection scans a file on its first open for data after it was written, and the install had just written all of them. Everything that opens the files pays once: TigerSetup's own hashing, and the Restart Manager, which opens every registered file to find its holders. The cost is per file opened, not per byte hashed, so a micro-benchmark of the hash (warm files) never shows it, and it is largest exactly where the benchmark measures — an uninstall right after an install. **Do not:** open a freshly written file for data unless the bytes are needed, and do not hand the Restart Manager a file nothing holds; do not read a slow first pass as slow hashing. **Use instead:** decide from the directory entry where the decision allows it — an owned file whose size and last-write time are still what TigerSetup recorded is the file TigerSetup wrote, so its recorded hash stands without a read (`win::fs::inspect_unless_unchanged`) — and probe a file's holders with an open for `DELETE` and write access that grants every sharing mode and asks for no data (`win::fs::is_held`; see *A running program's image can be renamed, so a delete-access probe does not see it* for why delete alone is not enough), registering with the Restart Manager only what the probe says is held. **Prevented by:** `benchmark/scripts/Measure-LocalUninstall.ps1` reports the engine's own span and the log carries `plan_completed` and `restart_manager_checked` with milliseconds, so the phase that regressed is named; `win::fs::probe_tests` pins the probe's semantics. **Generalization candidate:** the Windows fact belongs to any Tiger tool that reads what it just wrote; the method — attribute a phase before optimizing a loop — is general. ## A fixture binary the tests copy is not rebuilt by the tests **Area:** process-level tests (`crates/tigersetup-setup/tests`), the controlled programs (`TigerSetupTestAction.exe`, `TigerSetupTestPrereq.exe`) **Status:** Active **Symptom:** a new quiescence test kept failing on the assertion that a held file cannot be renamed, after the test program's `--hold` had been changed to hold the file without delete sharing, and a manual run of the same binary confirmed the rename succeeding. The source was right; the binary was old. **Cause:** `common::workspace_fixture` copies the program from beside the engine if it is there and builds it only when it is not, and `cargo test -p tigersetup-setup` builds nothing of another package. A test program edited in the same session as the tests that use it runs as it was last built by a workspace build. **Do not:** read a failing process test as a defect in the engine or the test before checking that the fixture binary is as new as its source. **Use instead:** `cargo build -p tigersetup-test-action` (or a workspace build) before running the suite that uses it; `cargo test --workspace`, the gate, always builds it. **Prevented by:** nothing mechanical — the fixture deliberately avoids a rebuild per test binary. The gate builds the workspace. **Generalization candidate:** none — the shape is this repository's fixture convention. ## A chained lab step without a lease policy starts from the baseline **Area:** the lab drivers (`lab/Invoke-*Rows.ps1`), `TigerSetupLab.psm1` **Status:** Active **Symptom:** the recovery driver's `uninstall` row reported every command as `The executable 'C:\TigerSetupLab\...' does not exist.` after its prepare job had installed the product and preserved the VM; the read step after a recovery scenario would have failed the same way. **Cause:** `Invoke-TigerSetupGuestCommands` takes the lab's lease policies as parameters and, given none, lets the lab start the job from the baseline — the right default for a lone job, and silently wrong for the second step of a row. Two steps of the recovery driver had never been given `Get-TigerSetupRowStepPolicy` when the drivers moved to lease policies, and the recovery rows had not been run since. **Do not:** read a chained step that finds nothing on the VM as an engine defect, and do not add a step to a row without deciding its policy. **Use instead:** every step of a chained row splats `Get-TigerSetupRowStepPolicy` (`-FromBaseline` for the first one) into the job helper; a row that ends with an omitted policy is a row that measured a clean VM. **The inverse omission, found on 2026-09-20 once the rows ran at all:** the install-interruption rows, each a single step, never passed `-FromBaseline` either, so every one of them ran on the VM the previous row left — where the product was already installed and the run ended in 0.3 s with `already_installed`, reaching no boundary and announcing nothing. The symptom read as "the interruption trigger did not fire", which is also what two genuine lab defects produced the same day; the launcher's record now says whether the run exited before the job returned, and with what code, so the three are told apart from the first row rather than after a second run. A row's first step, chained or not, is `-FromBaseline`. **A third instance, 2026-09-21:** row W6's wizard-capture step, the second of three, had never been given a policy either; the wizard ran from the baseline, the lab took the VM back, and the read step found no installer and no installation. Found by the full matrix, which is the first run of that row since the drivers moved to lease policies. **Prevented by:** `Test-LabScripts.ps1` now reports any function under `lab/` that starts a lab job (`Invoke-TigerSetupGuestCommands`, `Invoke-TigerSetupWizardCapture`, `Invoke-TigerWinLabEntryPoint`) without naming what the job starts from anywhere in its body — `Get-TigerSetupRowStepPolicy`, or the lab's `EntryPolicy`/`ExitPolicy`. The check fires on the W6 driver as it was and passes it as it is. **Generalization candidate:** none — the policy default is the lab's documented contract; the omission was this consumer's. ## The engine carries no compressor only while no engine code path names one **Area:** engine size, the installer format crate **Status:** Active **Symptom:** The first release build after the metadata block became compressed grew the engine by 349 KB of code — the whole Zstandard compressor, every block strategy from `fast` to `btultra2` — although the engine composes only the payload-less uninstaller copy and had never carried the compressor before. Compressed, that is about 170 KB more in every generated installer. **Cause:** the engine had always instantiated the same `compose` as the builder, encoder path included; the compressor stayed out only because fat LTO folded `files.is_empty()` for the engine's constant empty file list and dropped the payload writer with it. Compressing the metadata block added a second encoder call on the engine's path — through a function the optimizer no longer inlined — and the whole compressor came with it. Nothing in the build or the gate says which halves of libzstd an executable links. **Do not:** rely on the optimizer to keep a library's unused half out of a binary, and do not read "the engine was small last time" as a property of the code rather than of one build's inlining decisions. **Use instead:** make it structural. The engine calls `compose::compose_without_payload`, whose metadata block is a stored Zstandard frame written by hand (`payload::stored_frame`) and whose payload closure writes nothing, so no engine code path names an encoder at all. When composition or the format changes, link the engine with `/MAP` and attribute `.text` by object (`cargo rustc -p tigersetup-setup --release --bin tigersetup-setup -- -C link-arg=/MAP:`): libzstd's share of the engine is about 64 KB with the decoder alone and about 390 KB with the compressor. **Prevented by:** the engine's only composition has no encoder path by construction; the release turn reports the engine's raw and compressed sizes against `benchmark/README.md`, where a jump of this size is visible. **Generalization candidate:** any Tiger tool that links a library with a large optional half — the shape is "a size property that was really an optimizer property". ## A running program's image can be renamed, so a delete-access probe does not see it **Area:** Restart Manager coordination, `win::fs::is_held`, acceptance row M6 **Status:** Active **Symptom:** The first full acceptance matrix after the 0.8.0 quiescence optimization failed M6, the running-application upgrade: the application was running on the signed-in user's desktop, the upgrade log recorded no Restart Manager holder at all, the upgrade replaced eight files under the live process and ended `installed`, and its staging area could not be removed (`staging_cleanup_failed`, access denied) because the previous files it had moved there were still the running process's image. **Cause:** 0.8.0 stopped handing every file to the Restart Manager and probed each one first with an open for `DELETE` access that granted every sharing mode, on the reasoning that a file which can be renamed this moment has no holder the mutation would trip over. That is true of the mutation and false of the policy: Windows maps a running program's executable and DLLs with `FILE_SHARE_READ | FILE_SHARE_DELETE`, which is exactly why a running program can be renamed but not deleted or written. The probe opened a running application's own files without error and called them free, so the one kind of holder the Restart Manager exists to close never reached it. Nothing local caught it: the process-level holder model was a data file opened without delete sharing, which the probe does see, and the matrix row that would have failed was not rerun after the optimization — only the recovery, feature, elevation and UI rows were. **Do not:** reason about "held" from what the mutation needs; a rename going through is not evidence that nothing is using the file. Do not model a running application in a test as a file opened without delete sharing. And do not change the quiescence path without rerunning M6, whatever the cheaper rows say. **Use instead:** probe for `DELETE | FILE_WRITE_DATA` with every sharing mode granted: a data file held without delete sharing refuses the delete, a running program's image refuses the write, and a sharing violation on either is a holder. Where the write is refused for a reason that is not a holder — a read-only attribute, an ACL — fall back to the delete-access open alone. **Prevented by:** `win::fs::probe_tests::the_image_of_a_running_program_is_held_until_it_exits`, which runs a copied program and asserts its image is held while it runs and free once it exits; row M6, which asserts the Restart Manager listed the application as a windowed holder and shut it down. **Generalization candidate:** the Windows fact belongs to any Tiger tool that replaces files something may be executing; the method — a probe must model the policy's holder, not the mutation's — is general. ## A baseline's dependency state is a servicing fact, so a row declares its premise and the driver checks it **Area:** the acceptance matrix (`TigerSetup-Validation.md` §5.2, `lab/Invoke-MatrixRows.ps1`), TigerWinLab baselines **Status:** Active **Symptom:** the 0.10.0 release validation failed W3 and W5 — the Windows 10 rows whose premise was "WebView2 absent, .NET prepared" — with the installer finding nothing to acquire: `TigerWinLab-Win10-Clean` had held the WebView2 Runtime on a fresh restore since the September 2026 servicing of 22H2 (19045.6456), delivered by the Edge update the lab's baseline maintenance installs. The rows had been true of the image when they were written. Worse than the two failures were the rows that passed: W1, W2, W4 and W6 were described as exercising WebView2 acquisition and had passed on a baseline where there was nothing to acquire — vacuous evidence that read as green. **Cause:** the matrix wrote a dependency state into each row as an assumption about the image rather than as a premise the run verifies, and a lab that follows current servicing — the right choice, because a pinned or stripped image stops representing the machines the product installs onto — will move that state from under a row without saying so. **Do not:** pin the image or remove a runtime Windows delivered to recreate an old state; keep a row whose premise is false as an "expected failure"; read a passing acquisition row as evidence of acquisition without checking what it started from; or describe a row by the state its baseline used to have. **Use instead:** the two WebView2-absent states are Server 2019's, the one supported platform Windows does not give the runtime to, and every row with that premise runs there (S1–S6); the client rows start from WebView2 only or, with .NET prepared, both. The driver declares each row's starting state (`$RowPremise`) beside its baseline and asserts it from the lab's own measurement at the start of the row's scenario step (`premise/dependency state`), so the next runtime servicing delivers fails the rows it invalidates on the first run instead of passing them quietly. **Prevented by:** `Test-TigerSetupDependencyPremise`, which fails a row whose measured state is not its declared one and warns when the lab did not measure it; a driver that refuses to start when the row and premise tables disagree; and `TigerWinLab-Requirements.md` §3, which records that the baselines are current servicing and which baseline provides which state. **Generalization candidate:** TigerWinLab could report the dependency state of a clean baseline as part of its catalogue, so a consumer's premise could be checked before the lease rather than at the row's first step; the rule — a fixture's state is a measurement, not a constant — belongs to every Lab consumer. ## A single lab row can carry an 8–10 s launch stall that is nobody's code **Area:** benchmark and acceptance timing on the clean Windows 11 baseline; any conclusion drawn from one row's install or uninstall time **Status:** Active (cause is a consistent hypothesis, not verified) **Symptom:** in the broad-corpus campaign (`benchmark/report-0.10.0-broad.md`, 45 rows, 90 timed commands) seven commands started 8–10 s late and then ran at the usual speed — in all three technologies, on payloads of four files and of twelve thousand, on install and on uninstall. TigerSetup's engine log places the whole delay *before* the engine's first event (`beforeSeconds` in the row record: 8.8 s on a WinMerge uninstall that then took 0.5 s, 10.0 s on a Notepad++ install that then took 1.0 s); Inno Setup's and NSIS's show as a 9 s process lifetime on sub-second work. The identical WinSCP Inno Setup row took 2.7 s in a smoke run and 10.6 s in the campaign. **Cause:** every stalled launch is of an executable freshly written to the VM — a staged installer, TigerSetup's extracted engine, Inno Setup's second-phase copy, NSIS's `Au_.exe` — on an online clean Windows 11 with Defender at its defaults, whose cloud-delivered first-sight check holds an unknown executable for up to 10 s. That fits every instance and nothing else does, but no row ran with the check disabled, so it is a hypothesis. It is not the per-file scan of *The first read of a file an installation just wrote is the scanner's*, which scales with what was written; this is a fixed delay on one process start. **Do not:** read a single row's 9 s as a regression in the engine, the loader, the journal or a competitor; re-measure only the stalled rows (re-rolling a random stall biases the set); or compare technologies on one small payload's row. **Use instead:** medians and quartiles over the corpus, which absorb it; for a TigerSetup row, the engine span and `beforeSeconds` from the log the job brings back, which say where the time went; and, when a stall must be excluded from a measurement rather than tolerated, a lab row whose baseline has the first-sight cloud check off — a lab capability question, not a product change. **Prevented by:** the broad report's engine table (`before / engine / after` per TigerSetup row) and the campaign notes naming the stalled commands; `Invoke-BroadBenchmarkLab.ps1` records process lifetime and completion wait separately, so a stall never hides inside a hand-off. **Generalization candidate:** TigerWinLab — a clean baseline's Defender cloud-check behaviour is a platform fact every consumer timing a first launch on it would want to know or switch off; the reading rule is general. ## An elevated process cannot borrow the desktop user's token, but it can ask the desktop's shell **Area:** launch after install from an elevated wizard (`win::interactive`); anything that must start a program as the signed-in user from an elevated process **Status:** Active **Symptom:** the obvious design — read Explorer's token and start the program with `CreateProcessWithTokenW`, as the SDK's "run as desktop user" sample does — works when the prompt was a consent prompt (the same account, elevated) and fails when an administrator approved a standard user's credential prompt, or when an installer is started elevated over another account's desktop, which is exactly what the lab's elevated session is. **Cause:** an interactive logon's token carries a DACL that grants the Administrators group `TOKEN_QUERY` (`SW`) and nothing else — full access belongs to the user, SYSTEM and the logon session — so an administrator of another account can read the shell's token but not duplicate it. `SeDebug` does not help; it is about process objects, not token objects. Checked on a real token's security descriptor with `GetKernelObjectSecurity`. **Do not:** borrow, duplicate or impersonate the shell's token; enable privileges to force it; or fall back to the elevated token when it fails. **Use instead:** ask the shell to start the program. `ShellWindows` (`{9BA05972-F6A8-11CF-A442-00A0C90A8F39}`) has `RunAs = Interactive User`, so from any account it reaches this desktop's Explorer; its desktop view hands out `Shell.Application`, and `IShellDispatch2::ShellExecute` runs inside Explorer, with Explorer's token and environment. `TOKEN_QUERY` is still enough to refuse an elevated shell (UAC off) before asking. The program's process id is not returned, so it is found by name among the processes that were not there before. **Prevented by:** `launch-uac` (credential prompt, another account) and `launch-elevated` (started elevated over the standard user's desktop) in `lab/Invoke-LaunchRows.ps1`, which assert the program's parent is `explorer.exe` and its token is the standard user's, unelevated. ## The lab agent reads the wizard's check box through MSAA, not the UIA Toggle pattern **Area:** lab rows that read a check box's state on the wizard (`guest/Invoke-ElevationAcceptance.ps1`, TigerKeyring's acceptance) **Status:** Active **Symptom:** the first launch row found the completion page's check box but its `toggleState` was empty, so "offered checked" failed while the launch itself — which only happens with the box checked — passed. The earlier elevation rows' code that "cleared the offer when it was On" had never matched anything for the same reason. **Cause:** the agent's `ui-find` element for the wizard's plain `BS_AUTOCHECKBOX` carries no UI Automation Toggle pattern, while its legacy MSAA description does carry the state. **Do not:** treat an empty `toggleState` as "unchecked", or leave a check that can never match looking like a working guard. **Use instead:** `Get-CheckState` in the guest script — the Toggle pattern where present, else the MSAA state's `STATE_SYSTEM_CHECKED` bit (`0x10`) — and a `WARN`, never a pass or a fail, when neither is readable; the behaviour the state leads to (a launch, or none) is asserted separately. **Prevented by:** the launch rows assert both the state and its consequence. ## A folder created as a side effect is a folder nobody removes **Area:** engine resources (`win::shortcut::write`), uninstall and rollback **Status:** Active **Symptom:** the first presence row for TigerSetup's own installer found the `TigerSetup` Start Menu folder still there, empty, after a clean uninstall. All three links were gone and every other residue check passed. The same happened after a rolled-back install. Every package that used `[[shortcuts]] folder` had shipped with the defect since the key existed. **Cause:** writing a link created its parent folder on the way (`create_dir_all`). That creation was never journaled or recorded, so no removal was ever planned for the folder. The unit and process tests checked the links and never the folder they sat in. **Do not:** let a resource create something on the way without deciding what removes it, and do not accept "the links are gone" as proof that the Start Menu is clean. Do not bound an upward folder walk by "inside a root" alone: the scope's roots nest (Startup lies inside Programs), and the first version of the fix would have deleted an empty Startup folder. **Use instead:** removing a link — forward, undoing a create, or resuming a removal whose link is already gone — removes the folders it leaves empty. It stops at any of the scope's shortcut folders, at a junction and at the first folder that still holds anything, and never fails the run over a folder (`resource::shortcut::remove_emptied_folders`). Restoring a link recreates its folder first. **Prevented by:** `a_start_menu_folder_goes_with_its_last_link` (`crates/tigersetup-setup/tests/features.rs`), `a_shortcut_folder_inside_another_is_never_removed` (the engine's unit tests), and the presence rows' `uninstall/Start Menu folder gone` check. ## A clean Windows 11 opens a `.md` file only after asking which app to use **Area:** installed documents, desktop lab rows **Status:** Active **Symptom:** a shortcut to an installed Markdown file does not open it on a clean Windows 11 (26200). Windows shows "Select an app to open this .md file", with Notepad listed first. A `.pdf` opens directly in Edge, behind Edge's one-time welcome card. **Cause:** the baseline has no association for `.md` at all (`assoc .md` finds nothing, `HKCR\.md` is absent). A TigerSetup shortcut can target only an installed file, so it cannot name Notepad itself. **Do not:** assume a document shortcut opens because it opens on a developer's machine, which usually has an editor registered for `.md`. When checking it on a clean guest, do not look for the picker as a new top-level window: the agent's window list never showed it. **Use instead:** find the picker by its "Just once" button, desktop-wide, through UI Automation, choose Notepad, and record the question as a WARN (`lab/guest/Invoke-PresenceAcceptance.ps1`). Open a document or a `.lnk` from a lab row with `cmd /c start "" ` (ShellExecute, what a click does). The agent's `start-process` is CreateProcess and fails with "not a valid Win32 application". ## Defender's machine-learning verdict can quarantine an anonymous test binary overnight **Area:** the local gate (`cargo test --workspace`), release readiness **Status:** Active **Symptom:** all fifteen loader tests failed with OS error 225 ("the file contains a virus or potentially unwanted software"). Microsoft Defender had quarantined the test fixture `fake-engine.exe` as `Trojan:Win32/Wacatac.C!ml` seconds after it was linked. It did so again on every rebuild, although the bytes (`/Brepro`) were the same ones that had passed the day before. **Cause:** a cloud machine-learning verdict against a tiny, unsigned console program with no version resource. The product binaries and the self-installer, which carry TigerSetup's VERSIONINFO, scanned clean. **Do not:** add a Defender exclusion to get the gate green. That changes the machine's security and hides exactly what a WinGet submission's scan will do. Do not read the failures as a loader regression either. **Use instead:** give every executable the build produces, test fixtures included, the version resource and the hardening flags the product binaries carry (`FAKE_ENGINE` in `crates/tigersetup-loader/build.rs`). Scan the release binaries and the installer explicitly (`Start-MpScan -ScanType CustomScan`) before a release, and read `Get-MpThreatDetection` whenever a file vanishes or cannot be opened. ## A hosted CI runner is not a developer shell, and its first failure hides the rest **Area:** `ci.yml`, every test that touches the token, a temporary path, Git identity or the TigerAiCore configuration **Status:** Active **Symptom:** the first real CI run failed three engine unit tests that passed on every developer machine, and because `cargo test` stops at the first failing test binary, no later binary and none of the later steps ran at all — the release-mode COM tests, the release build, `Test-LabScripts.ps1`, `Test-Release.ps1` and the final `git diff` were never observed. **Cause:** a GitHub `windows-latest` runner runs the job as the administrator `runneradmin` with an elevated token, its `%TEMP%` is the 8.3 spelling `C:\Users\RUNNER~1\AppData\Local\Temp`, it has no Git identity, and it has no `TigerAiCoreConfig`. A developer shell is unelevated, long-named, has an identity and a configuration. One engine test read the real token (`elevation::required`), two compared a path Windows had expanded (see the shortcut lesson above), and `Test-Release.ps1` failed on the missing identity (`git config --get user.name` exits 1 when there is none). Once reproduced, more elevated-only defects were visible by reading: four loader tests and two wizard launch tests assume the unelevated path, the machine-scope ACL test's elevated branch expected an explicit list on inheriting files, and the elevated `staging_directory()` resolved `%SystemRoot%` through a table that did not know it, staging relative to the current directory. **Do not:** push to find out. Do not special-case `GITHUB_ACTIONS` either: each failure was a test assumption or a product defect. **Use instead:** before a change reaches CI, reproduce the runner's differences locally — `TMP`/`TEMP` set to a short spelling of a real directory (the `ShortPath` of a folder with a long name), `GIT_CONFIG_GLOBAL` pointed at an empty file, `TigerAiCoreConfig` removed, and the `GITHUB_*` variables set — and run the workspace with `--no-fail-fast`. An elevated token cannot be produced without a consent prompt, so make token-dependent decisions take the token as a parameter (`elevation::required_for`, `staging_root`) and test both values, and read every `is_elevated` branch a test reaches before trusting a local pass. **Prevented by:** the tests above check the half that applies to the token they run with. The elevated half is the one difference no local run reproduces, so `ci.yml` runs `cargo test --workspace --no-fail-fast` on the runner, started by hand when a change touches a token-dependent path; a failing binary no longer hides the others. It does not run on every push: the whole gate there cost about twenty minutes a commit for evidence the local gate had already given. **Generalization candidate:** the runner facts are recorded in TigerAiCore's `docs/release-model.md`; the reproduction recipe is Rust- and project-specific. ## The working tree is not the committed text: compare shipped text with the commit, CRLF-exact **Area:** release validation, lab self-installer rows, Windows text files **Status:** Active **Symptom:** The 0.12.0 draft failed exactly one check in `user-nopath` and `machine-nopath`: the shipped `help\TigerSetup-Help.md` was "not" `docs\TigerSetup-Help.md`. The content was identical; the shipped copy was CRLF, the file on the developer machine LF. A plausible first fix — compare the two with line endings normalized away — would have made the row green and blind to exactly the defect the CRLF policy exists to catch. **Cause:** Windows text files are canonically CRLF, and that is intentional: AI agents and some editors write LF-only or, worse, mixed files on Windows. Nothing in the repository establishes it — there is no `.gitattributes` and no repository-level setting. It is Git for Windows' **system** `core.autocrlf=true`, on this machine and on the `windows-latest` runner: blobs are stored LF, and a checkout writes CRLF. A file written after the checkout keeps whatever its writer produced, and `git status` still calls it clean, because the comparison goes through the same conversion. So the working tree here held 255 LF files that Git reports as unchanged, while the release workflow's fresh checkout of the same commit packaged CRLF. The check hashed the working tree, which is neither the commit nor what a checkout writes. **Do not:** compare shipped text with working-tree bytes, and do not make the comparison line-ending-insensitive, trim, or normalize whitespace — each lets an LF-only or mixed file through. Do not "fix" the working tree by hand either; Git's checkout is the authority for the CRLF spelling of a commit. **Use instead:** the commit the artifact was built from (the release record's `sourceCommit`), rendered through Git's own checkout conversion — `git cat-file --filters :` — compared byte for byte, plus a separate CRLF-only check (no lone LF, no lone CR) on the shipped file. On a machine whose Git does not apply the policy, the expected bytes come out LF and the CRLF-only check fails loudly rather than passing. **Prevented by:** `Get-TigerSetupCommittedFile`, `Get-TigerSetupCommittedTextChecks` and `Get-TigerSetupLineEndings` in `lab/TigerSetupLab.psm1`, used by `lab/Invoke-SelfInstallerRows.ps1` (`-SourceCommit`); `lab/Test-LabScripts.ps1` exercises the verdicts — CRLF passes; LF-only, mixed, lone CR, changed content and trailing whitespace fail — and proves in a throwaway repository with its own `core.autocrlf=true` that the expected bytes follow the commit, not the working tree, and that a file committed mixed is not repaired on checkout. **Consequence to know:** `packages/tigersetup/Build-Package.ps1` copies the working tree's help as it is, so a local candidate built from an LF working tree ships LF Markdown and now fails the CRLF-only check. That is a true finding — such a candidate is not what the release builds — not a flaky row. ## An Inno Setup uninstaller's process and its registration both end before its uninstall does **Area:** legacy migration (`crates/tigersetup-engine/src/legacy.rs`, `TigerSetup-Design.md` §5.12) **Status:** Active **Symptom:** after TigerMarkView migrated from its Inno Setup installer, a later TigerSetup uninstall left the empty install directory behind — per user intermittently (one pass, five failures), for all users every time. The install log showed `legacy_found` to `legacy_uninstalled` in about half a second, and the uninstall removed one directory fewer than the install had created. **Cause:** the migration waited for the `unins000.exe` it started and then for the registration key to disappear, and both signals come early. `unins000.exe` copies itself to `%TEMP%\…-uninstall.tmp\_unins.tmp` and runs the copy as a second phase (`Setup.Uninstall.pas`); the copy undoes the install log in reverse, so the key — written last — goes first; `DeleteUninstallDataFiles` then sends `WM_KillFirstPhase`, `unins000.exe` exits 0, the copy sleeps 500 ms, deletes `unins000.exe` and only then removes the directories it could not remove while that was in them (`LoggedProcessDirsNotRemoved`), the install root among them. So the install root still existed when TigerSetup planned, the engine correctly recorded it as found rather than created, and correctly left it standing on uninstall. The comment in the old code ("re-launches itself and returns immediately") and an earlier lesson ("its lifetime is the uninstall's") each described half of this. **Do not:** treat a foreign uninstaller's exit, or its registration disappearing, as the end of its uninstall; poll for the old install root to disappear (that waits forever for a root holding a user's file, and papers over the missing completion condition); read a directory left after uninstall as an ownership bug before reading the migration's log. **Use instead:** run the uninstaller in a job object that allows no breakaway, created suspended and resumed only once assigned, and wait — bounded — for the job to be empty (`process::start_tree`); read the exit code and the key only then. A console process in the job brings its `conhost.exe` with it, so count "at least the two phases", not exactly two. An uninstaller that relaunches itself elevated through `ShellExecute runas` would escape the job; the migration never meets one, because it runs a machine-scope uninstaller from the elevated engine and a per-user one for a per-user installation. **Prevented by:** `crates/tigersetup-setup/tests/legacy.rs` — an Inno-shaped fake (`TigerSetupTestAction.exe --hand-off`) whose second phase removes the key and then waits at a gate the test holds shut, so "the key is gone but the uninstall is not over" is observed rather than raced, in both scopes and through `Setup.exe` to a later uninstall; lab row M16's `cycles` job repeats the real TigerMarkView Inno migration and the later uninstall in both scopes and asserts the install root is gone and the log's `processes=` count. **Generalization candidate:** no — the job-object completion rule is TigerSetup's; the Inno facts belong with the migration that depends on them. ## One `del` with an unreachable path deletes none of the others **Area:** the uninstaller copy's self-deletion helper **Status:** Active **Symptom:** After an uninstall through the copy Add/Remove Programs runs, the copy and the executable it had moved aside stayed in `%TEMP%\TigerSetup` indefinitely, although the helper kept retrying and nothing held either file: PowerShell deleted both at once from outside. **Cause:** the helper ran one `del` naming the moved-aside executable, the copy and the state directory's `uninstall.exe`, and the third path's directory — the state directory — was already gone. `cmd.exe`'s `del` given a path whose directory does not exist fails the whole command and deletes none of the files before or after it. The same script passed a unit test whose missing file sat in a directory that existed. **Do not:** combine several files in one `del`, or assume a missing file is a harmless no-op for its neighbours. Do not read a file that survives a retrying deletion as proof that something still holds it. **Use instead:** one `del` per file (`self_deletion_script` in `crates/tigersetup-setup/src/main.rs`). **Prevented by:** the unit test `the_self_deletion_script_waits_for_its_files_then_removes_empty_directories` names a file in a directory that does not exist; the process tests `a_committed_uninstall_through_the_copy_leaves_nothing_behind` and `the_uninstaller_copy_removes_the_product_and_its_state_directory` wait for the helper and assert that nothing of TigerSetup's is left in the temporary folder or the state root. **Generalization candidate:** none — it is `cmd.exe` behaviour, and only this helper scripts it. ## `cargo test` started from Git Bash fails the notices check in PowerShell's authorization **Area:** the verification gate; agent shells **Status:** Active **Symptom:** `the_notices_list_exactly_the_crates_the_release_binaries_link` failed twice in a row with "THIRD-PARTY-NOTICES.md is not current", although `eng\Update-ThirdPartyNotices.ps1 -Check` passed when run by hand. The output the assertion carries ended in `SecurityError: AuthorizationManager check failed.` **Cause:** the test runs `pwsh -File eng\Update-ThirdPartyNotices.ps1 -Check`. When `cargo test` itself was started from an agent's Git Bash shell, that `pwsh` failed its execution-policy authorization before the script ran, so the check failed without looking at the notices. The same `cargo test` started from PowerShell passes, every time. The notices were never stale. **Do not:** treat this failure as stale notices and regenerate them, and do not run the gate's `cargo test` from Git Bash; `pwsh -File` run directly from Git Bash is not affected, which is why the lab and release-tooling steps pass there. **Use instead:** run the verification gate's cargo steps from PowerShell; read the assertion's carried output before acting on its first line. **Prevented by:** nothing mechanical — it is the shell the gate is started from. **Generalization candidate:** possibly TigerAiCore's, for any project whose Rust tests run `pwsh`, once a second project shows it.