--- name: create-routine description: Create, edit or port a Mentra automated testing routine with English requirements, saved actions verified during authoring, and deterministic replay through the shared framework. To request routine runs or authoring from a PR, use select-pr-routines instead. --- # Create, edit or port a routine Routines should be **fast, reliable and easy to create or edit**. Keep English instructions, observable expectations and executable actions together in the private [Mentra-Automated-Testing repository](https://github.com/Mentra-Community/Mentra-Automated-Testing). Read its [porting guide](https://github.com/Mentra-Community/Mentra-Automated-Testing/blob/main/docs/ROUTINE-PORTING.md) when migrating old coverage; it links the deleted source and explains what to reuse. ## Put the behavior in the routine Inspect the selected harness revision's `routines/`, `framework/` and controller schemas. Start from the closest routine for the platform and glasses. Preserve proven product actions and fixtures; replace old executor/ownership wrappers. | Location | Responsibility | | --- | --- | | `routines//routine.ts` | Export `createRoutine(state)` using `defineRoutine` and `step`; English metadata, product setup/steps/teardown and fixtures | | `framework/drivers/` | Shared interactions and recorded observations; Mac uses `executeMacStep(action, context)` with the supplied `context.ui` | | `framework/platforms/`, `framework/glasses/` | Composed platform and glasses lifecycle providers | | `framework/authoring/`, `orchestration/` | Held sessions, jobs, lane ownership, repair and publication | Declare platforms, entry (`home` or `sign-in`), account, requirements, fixtures, stable ordered step IDs, `glasses.models` and required capability IDs in `requires`. Keep device identities, secrets and tool paths in private lane configuration. Confirm the installed lane offers those capabilities. Add a reusable provider once for missing shared functionality; do not hide host setup in product steps. Preflight the complete fixture contract before reserving hardware. Check tool roles, not just executable hashes: the Mac UI driver and app launcher are distinct pins. When a capability is missing, assign its shared provider work separately and use the other lane or already installed routines while it is built. No routine-name branches in workers, dispatch or catalog, and no hardcoded videos: source enrollment discovers definitions; published passing runs supply examples. ## Hold one session and verify the saved actions Start through the built-in machine `routine-work` create/edit job described in the harness [job guide](https://github.com/Mentra-Community/Mentra-Automated-Testing/blob/main/docs/ROUTINE-WORK.md) and [assigned-agent skill](https://github.com/Mentra-Community/Mentra-Automated-Testing/blob/main/.agents/skills/prepare-routine-work/SKILL.md). The supervisor owns the workspace, machine agent, reservation and held session. The following are inner operations for that assigned agent, using its provisioned `MENTRA_TEST_CLIENT_CONFIG` and the current `mentra-test` CLI; they are not a parallel coordinator authoring path: ```sh bun run mentra-test lane request @reservation.json bun run mentra-test lane wait @wait.json bun run mentra-test author start @start.json bun run mentra-test author command @command.json bun run mentra-test author inspect @scope.json bun run mentra-test lane give-back @give-back.json ``` Read harness `orchestration/README.md`, `orchestration/controller.ts` and `framework/authoring/session.ts` for current schemas and held-session behavior; `contracts/controller.ts` defines admission. Run the CLI from the harness checkout, not MentraOS. Do not invent IDs or use an old standalone author CLI. Use `ControllerClient`/these service endpoints for controller mutations, including diagnostic attachments. Never open `ControllerStore` against the live database to register evidence, change ownership or manufacture cleanup receipts. Read-only SQL can help inspect state; a missing public operation is framework work to assign. Reservation request supplies `requestId`, `laneId`, `purpose`, `admissionExpiresAt`. Wait with `{reservationId, afterGeneration, timeoutMs}` until granted. Start supplies the granted `reservationId`, `generation`, a stable `operationId`, selected `build` and canonical editable `sourcePath`. Default start performs setup and starts the original recorder; optional `setupMode: "manual"` exposes individual lifecycle actions. Each author command carries `{reservationId, generation, operationId, command}`. Nested commands use `op`: `steps`, `snapshot`, `step` with `stepId`, `actions` with `phase: "setup" | "test" | "teardown"`, `action` with setup/teardown `phase` and `actionId`, or `finish`. Use returned IDs and inspect `{reservationId, generation}` until each operation settles; a settled operation may contain a failed assertion. Use a new operation ID for each action; a lost response reuses its original ID to reconcile that call. Direct driver calls must retain the supplied owned context. For example, the inner product command is `{ "op": "step", "stepId": "saved-id" }`, inside `command`, not a separate CLI verb. Declare the complete flow before starting. Use computer use to discover controls, save each action and assertion, then execute that saved action through the same driver/helper replay will use. Traverse the **whole English flow** this way: a manual click does not prove a different script written afterward. Prefer the simplest supported interaction that works; verify outcomes rather than successful clicks. Complete every saved product step and normal teardown before ordinary replay. A partially successful held traversal or expired recording is not that boundary. On a settled step failure, inspect the actual error, edit that existing action and retry with a concrete `retryReason` from its current safe prerequisite state. Keep the same owner, recorder and passing prefix; do not reinstall or restart setup for an ordinary authoring mistake. Do not repeat an uncertain submission/firmware write. If returning to a prerequisite needs an already passed product action, inspect `{op: "actions", phase: "test"}` for eligibility and repeat that same saved action with an explicit `retryReason` describing the observed prerequisite. The controller must confirm its previous input settled; a source reload alone permits no repeat. For a completed navigation tap, wait for its observable destination before the next input. A delivered-but-rejected tap keeps its intent: reconcile the resulting page without tapping again. Keep those checks in the same saved action for replay. The held loader preserves `createRoutine(state)` state and original lifecycle while reloading existing product steps. Changing step IDs/order, lifecycle callbacks or metadata requires finishing the session first. Shared helper/native changes require the updated installed revision and a fresh session. Fix a broken app control rather than accumulating alternate input or focus algorithms. Prepare and compile changed shared source off hardware while other work uses the lanes. Once affected owners release, activate one frozen candidate; local iteration may use reviewed source before merge while retaining the PR's review/CI merge gates. Check the installed recorder's duration, byte limit and output allowance before a long flow, including held editing and accepted operation settlement; step deadlines do not extend capture. A Mac fixture with a recorded browser window declares `external-window` with `fixture-data` and uses the shared admitted policy for both recordings. A product update can continue after a recorder or client deadline; inspect that original operation and settle it normally rather than issuing another update or restarting the passing prefix. Request authoring reservations before waiting for the current run to finish, so the next queued job does not repeatedly displace ready authoring work. Before ordinary dispatch, confirm the installed executor source and enrolled definition revision agree; frozen requests do not change during a service upgrade. Enroll the intended source before submitting new work. A stale local request that never launched can be cancelled normally with a reason, then replaced with the same saved actions on the intended source. Preserve launched/nightly requests. Publish startup failures through normal evidence delivery without replay; inspect the exact export error when publication stalls rather than repeating the test. For UI transitions, verify the departing overlay disappears as well as the new page appears. Home controls can remain visible behind a miniapp. Use bounded postcondition observation; an acknowledged click is not a completed transition. Inspect current controls rather than copying old labels blindly: Android's radio icon can toggle while its label opens details, and an empty miniapp switcher opener can remain present on idle Home. Require the actual state before sending input. Before hardware, compare the old saved selector with the selected build's current component. Several URL editors can coexist: preserve OTA's specific manifest placeholder instead of selecting any editable field. After an uncertain typing response, observe the exact requested value before clearing or typing again; the empty placeholder disappears when input succeeded. Keep that observation in the same saved action, not a separate replay technique. Static headings may appear twice on a platform: require readable content, and use exact IDs/counts for the actionable controls that must be unique. On Android, use the supplied `ui.scroll(anchor, direction)` for a bounded gesture inside the observed scroll view, then resnapshot. Check `checked` for toggles rather than assuming a click changed them; public text replacement uses `clearText` before `type`. Use `hideKeyboard` for the actual IME. The optional `systemUi` retains the same ownership and permits only the enrolled system-dialog namespaces; normal `ui` remains scoped to the Mentra App. Reuse observations, not assumptions. Android can repeat a radio label on its parent and text child; count actual checkable controls. A tall option group can span viewports: accumulate all known checked/unchecked states under the same foreground group, reject contradictions and scroll toward an unobserved option. Validate visible preconditions again in the driver's final input-planning snapshot. After one acknowledged tap, observe its resulting state; don't repeat input because a receipt write or screenshot failed. Retain an answered-input flag before later assertions so a settled held retry continues observation rather than toggling again. Keep a failing native command's bounded original error/cause in diagnostics before iterating. Compare its failed expectation with the original AX/XML and recording: an observed offer with a different label is a selector mismatch, not proof the offer needs more time. Correct the saved observation in the held state first. If original diagnostics would be disposed when an authoring reservation returns, preserve their bounded safe failure summary in the existing operation receipt first. Distinguish an empty successful trace from a malformed/failed read; do not repeat setup merely to guess the missing cause. Compare a failed HTTP probe with the app's actual request contract before calling it a server outage. A missing positive log reply is an observation gap, not proof the product action failed. Use the assigned phone's current-process trace for BLE replies when camera logs flood the glasses' short tail. Reconcile the original request instead of resending it. For cleanup-only corrections, select the reviewed provider source through the public repair API while retaining the original resource/fixture inputs; a new checkout alone does not change the implementation used by repair. When a framework callback times out, retain its original operation/call identity, deadline and bounded queue/start/finish/send facts through existing diagnostics. A successful earlier snapshot does not prove callback settlement. Fill a diagnostic gap before another expensive reproduction; exclude callback inputs and credentials. For an apparent hosted recording fault, compare the same run's exact asset hash and decoded frame at the saved timestamp with its screenshot before changing native capture. The coordinator owns browser playback/seek diagnosis. For audio coverage, read the harness `framework/audio/witness.md` and current browser service reference before adding helpers. The optional native witness uses the existing audio grant: await actual capture readiness before the stimulus and evaluate completed pinned PCM. Keep challenge words/assertions in the routine; shared providers own exact device routes, children and guarded mute restoration. Browser device options require actual selected state, and RTP counters alone do not prove heard speech. Preserve sequential speech/mute-control coverage without claiming simultaneous duplex. Missing configured endpoints/tools are a precise prerequisite to coordinate, not reason to revive an old reservation or runner. Flag actual bugs and impossible/human-only requirements with the exact failed step. Routine code does not repair the harness. Finish runs the original teardown; then give back with `{reservationId, generation, requestId}` for ordinary boundary cleanup. ## Shared lifecycle and modified miniapps Shared providers install the selected Mentra App, establish requested account/entry, prepare applicable glasses/fixtures, record, settle resources and uninstall the owned app. Routine setup/teardown own only product-specific effects. Inspect the current miniapp's persistence before porting old teardown: app-local SimpleStorage is removed with shared app data, while backend fixtures need their own exact owned-ID cleanup. Uninstall does not delete cloud data. Cleanup must not wait for a product effect that failed to be created. Use the supplied account context (`account` on Mac, `credentials()` on Android) and optional `audio` or `fixtures` when the installed platform supports them. Routine fixture content stays in `routines//`; reusable capture/connection/audio and platform delivery belong to shared providers. Do not copy another lane's serial, account, audio route or firmware setup into the routine. To try a modified miniapp, build/pack it in its source repo, then from MentraOS run `bun scripts/load-authoring-miniapp.mjs --mac` (set `MENTRA_MAC_APP`) or `--android `. The installed app needs existing Super Mode and miniapp permissions. Keep the temporary server until loading completes, verify the changed saved step, then stop it. See the [miniapp CLI guide](../../../sdk/miniapp-cli/README.md#try-a-packed-miniapp-during-routine-authoring). ## Replay and publish After the full saved flow works, commit/enroll its exact source and replay the **same actions** through normal setup/test/teardown: ```sh bun run mentra-test source enroll @source-enrollment.json bun run mentra-test run submit @run-request.json bun run mentra-test run dispatch-once '{"id":"ACCEPTED_LOCAL_REQUEST_ID"}' bun run mentra-test run inspect '{"id":"ACCEPTED_LOCAL_REQUEST_ID"}' ``` Use `enrollRoutine`/platform enrollment and `localAdmissionId` helpers for source provenance and local admission. Activate a changed shared framework/native revision only after affected held sessions and runs have finished; do not replace their pinned source under active owners. Ordinary replay stops at its first failed product step, preserves remaining steps as `not-run` and still tears down; other runs stay independent. Preserve the original error if cleanup/publication also fails. Completion needs passing assertions, teardown, acknowledged evidence and working recording/step seeking. A composite fixture must preserve evidence errors returned by its shared providers even when their physical cleanup succeeded. Bound the whole recorded observation to its declared timeout; polling must not consume a bounded event journal by writing a marker on every read. Do lengthy external fixture preparation before entering the browser observation deadline, retaining its declared app/resource operation budget. A preparation refusal before input can use the existing original-owner settlement API only when it proves zero dispatch and unchanged idle state; unknown or answered inputs remain retained. Settle active media before changing its route. After media is settled, attempt independent cleanup of browser evidence, clipboard, receiver and network resources even if one fails; preserve the first error and subsequent failures. The coordinator owns deployed Admin/playback verification. Report exact source/build/platform, result URL, recording and timings. Do not add repeated qualification runs without changed code or unresolved failures. `run retry-publication` retries delivery without hardware replay. Dispose owned local payloads after acknowledgement; preserve shared tools and native Codex/Claude history. Keep frozen execution evidence immutable; late boundary/repair observations use their own existing operation diagnostics. Check new exports against the installed publisher's asset-count and envelope-size bounds before freezing. A rejected old export stays unchanged for normal delivery retry after its owning contract is fixed; do not filter attachments or fabricate acknowledgements to obtain a pass. Use focused checks and [codex-pr-review](../codex-pr-review/SKILL.md) for the PR; [select-pr-routines](../select-pr-routines/SKILL.md) selects relevant coverage labels.