--- name: marionette-flutter-drive-app description: Set up and drive a running Flutter app (debug or profile) with Marionette — an AI agent's hands and eyes for the app. Covers adding marionette_flutter, MarionetteBinding, MarionetteConfiguration for custom widgets, a LogCollector, and MarionetteDeviceConfig — then tapping/typing/scrolling, screenshots, logs, hot reload/restart, device sweeps, and custom extensions, via marionette_mcp or the marionette CLI. Use whenever asked to test, click through, smoke-test, automate, QA, or interact with a Flutter app; when integrating marionette_flutter or registering a custom extension; after a feature or UI/design-system change; to reproduce or verify a bug; to walk a behaviour-neutral refactor; to sweep for accessibility; to exercise form validation; or to capture before/after evidence for a PR. Also load whenever a Marionette call fails, can't see a widget, reports not connected or a version mismatch, or hits a single-binding assertion. Not for CI suites, release builds, or performance work — see When not to use. --- # marionette_flutter: drive an app Marionette lets an AI agent drive a **running** Flutter app (debug or profile — never release) the way a real user would: read the screen, tap, type, scroll, screenshot, read logs, hot reload. It has two halves: - **`marionette_flutter`** — the package you add to the app. It registers a VM service extension per action (`ext.flutter.marionette.*`). - **The bridge** — either **`marionette_mcp`** (MCP tools, e.g. `connect`, `tap`, `get_interactive_elements`) or **`marionette_cli`** (shell commands, e.g. `marionette tap --key ...`). Nearly the same capabilities, different transport — the CLI additionally offers `record-video`, with no MCP equivalent (see the capability table, below). Prefer the MCP tools when they're in your tool list; fall back to the CLI in restricted environments (enterprise policy, a shell-only agent) — run `marionette help-ai` once at the start of such a session and follow its reference for exact syntax. Everything shared between the two applies to either transport — action names differ only in casing (`get_interactive_elements` vs `get-interactive-elements`). ## Preparing the app `connect` (or any CLI command) only works if the target app initialized `MarionetteBinding`. If you're not sure whether it has, check `main.dart` (or ask) — skip the rest of this section if it's already there and you just need to drive the app, and jump to *When to use this*. ```bash flutter pub add marionette_flutter ``` ```dart void main() { if (kDebugMode) { MarionetteBinding.ensureInitialized(); } else { WidgetsFlutterBinding.ensureInitialized(); } runApp(const MyApp()); } ``` This is all a standard Material-widgets app needs. It runs inside `main()`, so any time you add or change it, **hot restart** — a hot reload doesn't re-execute `main()`, so the change will look like it did nothing. Gate every piece of Marionette wiring behind `kDebugMode` (or `!kReleaseMode` — see *Device-config sweeps*, below) so none of it ships in a release build: both are compile-time constants the release branch survives cleanly. ### Custom design system? Teach Marionette about it Zero-config setup only recognizes stock Material widgets (`ElevatedButton`, `TextField`, `Switch`, …) and extracts text only from `Text`, `RichText`, `EditableText`, `TextField`, and `TextFormField`. If the app wraps its own buttons, fields, or text — true of most production apps — `get_interactive_elements` will look sparse and `tap(text: ...)` will fail to find things that are clearly on screen. That's not a Marionette bug; it just doesn't know your widgets yet: ```dart MarionetteConfiguration( // Recognize your custom interactive widgets. isInteractiveElement: (element) => switch (element.widget) { MyButton() || MyTextField() => true, _ => false, }, // Extract their visible text for text-based matching and discovery. extractText: (element) { final widget = element.widget; if (widget is MyText) return widget.data; return null; }, ) ``` Pass it as `MarionetteBinding.ensureInitialized(MarionetteConfiguration(...))`. For a non-trivial design system, extract the configuration to its own file (e.g. `lib/debug/marionette_config.dart`) imported only behind `kDebugMode`, rather than growing it inline in `main.dart` — that keeps release builds free of design-system introspection helpers even before the compiler strips them. `isInteractiveElement` and `extractText` receive the `Element`, not just the `Widget`. Match with an `is` check or pattern on `element.widget` — that covers generic widgets (`MySelect()` matches `MySelect`) and subclasses (`MyButton()` matches `MyPrimaryButton extends MyButton`), and can read the widget's fields to decide per instance. `extractText` can also walk the subtree when a label is itself a widget rather than a plain string. If the app still uses the deprecated `isInteractiveWidget: (type) => ...` or `shouldStopTraversal: (type) => ...`, move it to `isInteractiveElement` / `shouldStopTraversalAtElement`: a `Type` only compares with exact `==`, so it silently misses generic and subclassed widgets. While both are set, their results are combined with OR. Custom-painted text and badges reach no `Text` widget at all, and a `WidgetSpan` is only a problem when its embedded content isn't itself built from `Text`/`RichText` (an icon, a custom-painted chip) — plain text nested inside one is still its own discoverable element. For genuinely non-text content, annotate with `Semantics(label: ..., value: ...)` instead; Marionette surfaces that as a `Semantics` element with the joined `'label: value'` string, and it costs nothing if `label`/`value` are absent (unlabeled `Semantics` nodes are silently skipped). Not every custom widget needs an entry: add one when it's a primary interactive primitive (button, field, toggle, tab, chip) or when `get_interactive_elements` is noticeably missing it; skip widgets that are only ever decorative, or ones that already wrap a primitive Marionette sees (a `MyCard` built on `GestureDetector` is already covered). The list doesn't need to be exhaustive up front — extend it as real use surfaces a missing target. One more `MarionetteConfiguration` field, `shouldStopTraversalAtElement`, is worth naming only to warn against it: it's tempting to add a scroll container there to "reduce traversal cost," but on a real production app that measurably *dropped* widget coverage from 25.8% to 17.8% — the agent lost visibility into everything nested below the cut, including the content it needed to reach. Leave it `null` unless a profiler has shown a real, measured cost, and never point it at a scrolling container. ### Logs for `get_logs` Without a `LogCollector`, `get_logs` returns a message explaining how to add one rather than actual logs. Pick whichever matches how the app already logs: - Uses the `logging` package → `flutter pub add marionette_logging`, pass `LoggingLogCollector()` as `logCollector`. - Uses the `logger` package → `flutter pub add marionette_logger`, pass `LoggerLogCollector()` (it doubles as a `LogOutput` for that package too). - Anything else → `PrintLogCollector` (ships in `marionette_flutter`) exposes a manual `addLog(message)` to call from wherever logs already flow — routing Flutter's own `debugPrint` through it is a common choice, tee'd into the existing implementation rather than replacing it, to keep `debugPrintThrottled`'s throttling. ### Screenshot size `take_screenshots` downscales captures to fit within 2000×2000 physical pixels by default, to keep base64 payloads manageable. Override `maxScreenshotSize` on `MarionetteConfiguration` if a screen needs more detail (`Size(3000, 3000)`, say), or set it to `null` to disable resizing entirely — but keep the default unless there's a concrete reason to raise it, since larger screenshots mean larger payloads on every call. ### Device-config sweeps `set_device_config` — sweeping text scale, bold text, or light/dark appearance without touching real OS settings — is the one capability that needs a widget in the tree, since Marionette won't insert one into an app behind its back: ```dart void main() { if (!kReleaseMode) { MarionetteBinding.ensureInitialized(); runApp(const MarionetteDeviceConfig(child: MyApp())); } else { WidgetsFlutterBinding.ensureInitialized(); runApp(const MyApp()); } } ``` Without it, `set_device_config` responds with these setup instructions instead of succeeding. Because it lives in `main()`, adding it also needs a hot **restart**. `!kReleaseMode` (rather than `kDebugMode`) is the better gate here because it additionally covers profile builds, where the VM service is still reachable — use the same gate for both the binding and this widget, since gating the widget more loosely than the binding gains nothing (it degrades to a no-op child-passthrough if the binding isn't there, so a missed gate costs one inert element, never a crash). ### The single-binding rule Flutter allows exactly one `WidgetsBinding` per process. If something else claims it first, `MarionetteBinding.ensureInitialized()` throws a clear error — except with one real offender, where it doesn't: - **`flutter test`** running `main()` under `kDebugMode` collides with `AutomatedTestWidgetsFlutterBinding`. Guard against it — `Platform.environment.containsKey('FLUTTER_TEST')` — or give tests a separate entrypoint that skips `MarionetteBinding` entirely. - **Sentry** is the known silent offender: `SentryFlutter.init()` always calls `WidgetsBinding.ensureInitialized()` inside its own `appRunner`-wrapping zone before your code runs, and it swallows the resulting binding error as if it were a reportable crash — so the app just hangs on the splash screen forever with no exception, no crash, and no log output. The fix is initialization order: call `MarionetteBinding.ensureInitialized()` **before** `SentryFlutter.init()`, not inside its `appRunner`. Sentry checks for an existing binding first and reuses it, so going first avoids the conflict entirely. The same ordering fix applies to any other plugin whose `init()` touches `WidgetsBinding` ahead of your `appRunner`. ### Version alignment The MCP `connect` tool checks that `marionette_mcp` and the app's `marionette_flutter` are on the same version, and fails with a clear message rather than a confusing runtime error if they aren't. The CLI does not perform this check — a stale `marionette_flutter` there can fail in less obvious ways instead of a clear mismatch error. Either way, don't chase a mismatch by retrying — align the versions: `flutter pub add marionette_flutter` in the app, and `dart pub global activate marionette_mcp`/`marionette_cli` for the bridge (or `dart pub add dev:marionette_mcp`/`dev:marionette_cli` if you'd rather pin it as a dev dependency instead of a global tool). ## When to use this Reach for Marionette any time the fastest way to know whether something works is to actually run the app and look, rather than reason about the code: - **After implementing a feature.** Connect, navigate to the new screen, and exercise the happy path — a working build isn't the same as a working feature. - **After a UI or design-system change.** Confirm the affected screens still render and respond as expected, not just that the app compiles. - **Reproducing a bug from a ticket.** Follow the reported steps, capture a screenshot and `get_logs` output as evidence the bug is real and what it looks like — before touching any code. - **Verifying a bugfix.** Reproduce with the same steps, apply the fix, `hot_reload`, and replay those exact steps again to confirm the symptom is gone (not a different, adjacent symptom). - **Walking a "behaviour-neutral" refactor.** Claims that a refactor changes nothing observable are exactly the claims worth checking — a quick regression walk of the touched screens catches what review alone misses. - **An accessibility sweep.** Use `set_device_config` to try `text_scale` around 2.0–3.0 (real devices top out near 3.0) and `bold_text: true` on the screen in question, then check `get_logs` for overflow errors — that combination is where cramped layouts usually break first. (Needs the app to have opted in via `MarionetteDeviceConfig` — see *Device-config sweeps*.) - **Form and validation work.** Enter deliberately invalid input and read the error strings back via `get_interactive_elements`/`take_screenshots` — it's the only way to confirm the *user-visible* message is right, not just that a validator returns an error code. - **Evidence for a PR or ticket comment.** A quick before/after `take_screenshots` pair is often more convincing, and faster to produce, than a written description of a UI change. For a multi-step flow rather than a single static change, `record-video` (CLI-only, no MCP equivalent — see the capability table) captures the whole interaction as a short clip, which shows off a new feature or a fixed flow better than a screenshot pair can. ## When not to use this - **CI or a repeatable regression suite.** Marionette's gestures are best-effort simulations of user input, not deterministic instrumentation — results can vary across platforms, custom widgets, and overlays. For a suite that needs to pass or fail the same way every time, use [Patrol](https://patrol.leancode.co) instead. - **Release builds.** Not a limitation to work around — Marionette relies on the Dart VM Service, which doesn't exist in a release binary. Debug and profile only. - **Performance work.** Marionette can *trigger* the interactions you want to profile, but measuring them is Flutter DevTools' job, not this one. ## Connecting Every action needs a VM service WebSocket URI (`ws://127.0.0.1:PORT/TOKEN=/ws`) — the developer running the app has it in their `flutter run` console output, or in a DevTools link. Ask for it and pass it to `connect` (MCP) or `--uri` / a registered `-i ` (CLI) before anything else. If more than one running instance is already in play — several apps registered with the CLI, or the user's message mentions more than one device or simulator — don't guess which one to use. Either infer the right one from what's already been said in the conversation (e.g. "the iOS one" when only one registered instance is iOS), or ask which to target. Getting the wrong instance is easy to miss until several steps later, so it's worth the one question up front. For the CLI, prefer `--uri` for a one-off session and `marionette register ` + `-i ` when you'll issue many commands against the same app in a row. When you're done for the session (MCP), call `disconnect` — it's not required for correctness (a fresh `connect` implicitly replaces it), but leaving a session dangling makes it easy to act on a stale connection later without noticing. Give `connect` a `session_title` (a short description of what you're testing, e.g. `"profile validation"` — keep it under ~60 characters, longer is truncated) — if the app enables session reports (`MarionetteConfiguration(enableSessionReports: true)`), it opens a fresh session directory that carries this run's step log and screenshots. One Marionette run is one session directory, always: there's no resuming, even if you pass the same title again on a later `connect` — that just opens another, separate directory. If a run gets interrupted partway through (a compaction, a dropped connection), its session is simply abandoned; don't try to pick up where it left off, and don't reuse its title. **Always give it in English**, regardless of the prompt's own language (see *The report contract*, below) — it gets slugified straight into the session directory's name, and that name should stay predictable and filesystem-friendly rather than following the run's language. See *Reporting what you found* for what that directory is for and what you're expected to do with it before disconnecting. ## What you can do once connected | Category | Actions | Notes | |-------------------|--------------------------------------------------------------------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | Inspection | `get_interactive_elements`, `take_screenshots`, `get_logs` | `get_interactive_elements` is how you "see" the screen — call it before guessing at a target, and again after any navigation you didn't drive step by step. `take_screenshots` returns images inline by default; pass `inline: false` once a screenshot is evidence to cite rather than something you need to look at now — see *Reporting what you found*. `get_logs` needs a `LogCollector` wired up app-side (see *Logs for `get_logs`*, above); otherwise it returns setup instructions instead of an error. On a screen that repeats the same subtree, pass `ancestor_keys` to list just one copy of it (see *Repeated keys*, below). | | Gestures | `tap`, `secondary_tap`, `double_tap`, `long_press`, `swipe`, `pinch_zoom`, `scroll_to`, `press_back_button` | Match by `key` › `identifier` › `text` › `type` › coordinates, in that preference order — see *Good practices*. When the same key repeats in several identical subtrees, add `ancestor_keys` — see *Repeated keys*, below. `secondary_tap` is desktop-only. | | Text input | `enter_text`, `press_key` | `enter_text` overwrites a field's value directly, and takes `ancestor_keys` like the gestures do. `press_key` sends a real key event (submit on `enter`, shortcuts via `modifiers`) but only edits in-place on desktop/web — on mobile, field editing still needs `enter_text`. | | Device config | `set_device_config` | Sweeps text scale / bold text / light-dark under the app's *current* screen, without touching OS settings. Needs the app to opt in — see *Device-config sweeps*, above — otherwise it returns setup instructions rather than failing. | | Custom extensions | `list_custom_extensions`, `call_custom_extension`, plus any first-class tool an app registered with a schema | See *Custom extensions*, below. | | Dev workflow | `hot_reload`, `hot_restart` | `hot_reload` preserves state; use `hot_restart` only for changes a reload can't pick up (main()/bootstrap edits, global singletons, state shape) — requires the app to be running via `flutter run`. | | Session | `connect`, `disconnect` | `connect` must be called before any other tool; a second `connect` implicitly disconnects the first. `connect` also opens a fresh session directory (`session_title`) that owns the step log and screenshots — see *Reporting what you found*. | | Video (CLI only) | `record-video` | Records a WebM video of the session (`-o/--output`, `-d/--duration`, `--width`/`--height`; needs `ffmpeg` on `PATH`). No MCP equivalent — use `take_screenshots` there instead. | ## Good practices - **Look before you act.** Call `get_interactive_elements` (or `take_screenshots` for a visual check) before the first gesture on a new screen. Don't infer what's on screen from memory of the source code — the whole point of Marionette is to check the *running* state. The listing is compact by default: it drops rendering details and reports `visible` only when an element is not visible. Pass `compaction: "none"` (`--compaction=none` in the CLI) when you need the full property dump. - **Selector priority: `key` > `identifier` > `text` > `type`/coordinates.** Keys (`ValueKey`) and Semantics `identifier`s survive copy changes, localization, and refactors; `text` breaks the moment a label is edited or translated; coordinates break on any layout shift, font-scale change, or screen-size difference. Treat coordinates as a last-resort diagnostic, not the normal way to target something — if nothing else works, that's a sign the widget needs a key or `Semantics(identifier: ...)`, not a reason to keep using coordinates. - **Repeated keys: scope the match with `ancestor_keys`.** Every tool takes the first match in the tree, so on a screen that repeats the same subtree (grid cells, repeated cards, a dev screen embedding several copies of the app) a plain `tap(key: "cell.joinButton")` always hits the first copy. Pass the keys of the wrappers around the one you mean, outermost first: `{"key": "cell.joinButton", "ancestor_keys": ["session_2", "grid.cell_3"]}`. Each key is looked up strictly *inside* the previous one, which is what lets you reach a wrapper whose own key also repeats (`grid.cell_3` exists in every session), and the target itself must be inside the last one. A single key is often not enough — if the wrapper repeats too, go one level up. - **Finding the chain:** `get_interactive_elements` is a flat list in tree order, so a keyed wrapper, when it is listed at all, appears just before the elements inside it. A wrapper whose centre catches no hit test isn't listed, so if a key you expect is missing, read it from the source or ask for it. To confirm a chain, list only that subtree: `get_interactive_elements(ancestor_keys: [...])`. A scope alone is enough there, the wrapper itself is listed too, and a wrong key fails by name. - **A key with no element fails the call**, naming it and where it was looked for (`Scope element with key "grid.cell_9" (ancestor_keys[1]) not found inside "session_2"`). It never falls back to searching the whole screen, so a typo can't silently act on the wrong copy. - **`scroll_to` works into lazy lists:** a scope that isn't built yet (e.g. `ancestor_keys: ["row_20"]` in a long `ListView`) is looked for again after every drag, and the list that scrolls may be inside the scope or above it. A scope key that exists nowhere is reported only after the list has been scanned, like a missing target. - `coordinates` and `focused_element` ignore `ancestor_keys`, and a scope on its own is not a selector for the gesture tools. - Prefer this over tapping by coordinates or baking indices into keys: the chain survives layout changes and keeps production keys clean. - **An element "not found" is usually a setup gap, not a bug.** Before concluding a widget can't be reached, suspect it's a custom widget type Marionette doesn't recognize yet, or custom-painted content with no `Semantics` annotation — both fixed under *Preparing the app*, above, not something to work around here. - **If coverage looks low across a whole screen rather than one missing element, don't reach for `shouldStopTraversalAtElement`.** It's tempting to suggest filtering a scroll container out of traversal to "reduce noise," but that's the one config change measured to make things worse (25.8% → 17.8% widget coverage on a real app) — see *Custom design system?*, above. - **Confirm side effects with `get_logs`, not just UI state.** A tap that looks like it worked isn't the same as the network call actually firing. - **Give the app's flows explicit context.** Marionette can see the widget tree and act on it, but it has no idea what the product's flows, naming conventions, or edge cases are. State which screen to reach, the expected keys/labels, any precondition the flow assumes (e.g. "assume the user is already logged in"), and the interaction goal in one sentence, rather than assuming that's inferable from the UI alone. - **Treat every gesture as best-effort.** Focus, text entry, and scrolling are simulations of user input and can behave differently across platforms, custom widgets, or overlays. A flow that's consistently flaky (not just occasionally) is a signal to expose a clearer key/identifier, flatten a deeply nested `GestureDetector` hit target, or extend `MarionetteConfiguration` — not to retry harder. ## Reporting what you found This section applies only when `connect`'s response actually includes an `Opened session:` path — check the real response, not just whether you expected one. Two different things can be missing it: the app never enabled session reports (`MarionetteConfiguration(enableSessionReports: true)`) — don't turn that on yourself unless the user asks for a report — or it did, but the session directory itself couldn't be created (a connect response saying so instead, e.g. "Could not open a session directory, so this run will not be logged"). Either way, there's no session directory, no `steps.md`, and nowhere to write `report.md` — see the note on this at the end of *The report contract*, below. `connect` opens a fresh session directory under `.marionette/sessions/` — the server's own record of the run, kept separate from your own context. `disconnect`'s response repeats that directory's path with a reminder to write `report.md`: treat that as a requirement, not a suggestion — it's the enforcement mechanism precisely because an instruction here, on its own, gets skipped. Two files in that directory are appended by the server, not you — don't write to them yourself: - **`steps.md`** — one line per tool call: the tool, a short selector, and the outcome. No payloads (`enter_text` values are redacted). - **`screenshots/`** — populated only when you call `take_screenshots` with `inline: false`, which is what you want once a screenshot is evidence you're citing rather than something you need to look at right now (each inline image costs real visual tokens). You own one more file there, which the server never writes: - **`report.md`** — written once, at the end, from your own context. Write it on either of two triggers: the run completed, or you're stopping early (a blocking failure, a budget/time limit, an ambiguous requirement you can't resolve alone). A report for a run stopped early states *why* it stopped and what was left untested, instead of silently reporting only what got covered — there's no resuming to pick it back up later, so this is the only record of what happened. Before drafting a single line, name the report language explicitly: the language of the prompt that started this session (the message that led to the first `connect` call) — not English by default, not the language of the app's UI under test, not whatever language your own tool-call narration has been in. This check is easy to skip because it's just a rule sitting in the contract below; doing it now, as a deliberate step before writing, is what actually prevents defaulting to English. See *The report contract* for the full rule. ### The report contract - **Write `report.md` — and the chat message — in the language the prompt that started the run was written in.** Proper nouns stay exactly as they are regardless of that language: tool names, file paths (`report.md`, `steps.md`, `screenshots/01.png`), and widget keys/identifiers (`login_emailTextField`) are never translated. Check this before writing, not after — see the reminder in *Reporting what you found*, above. **`session_title` is the one thing that's exempt in the other direction: give it in English always**, whatever language the prompt itself is in — see *Connecting*, above. - **The `Tested:` line is rendered from `steps.md`, never from memory or your own sense of what you did.** This is the anti-overstatement mechanism — read the step count and the actions taken back out of the file the server wrote, don't estimate. - **Every finding cites evidence** — a step, a `get_logs` line, or a screenshot path. No evidence means it's a suspicion, not a finding: say so explicitly rather than upgrading a hunch. - **Under six lines per finding, no prose paragraphs.** A severity tag, one-line summary, a repro path, and the evidence citation — that's the shape; if it doesn't fit, cut the finding down rather than making room. - **Prefix the headline and every finding with a status icon**, so severity is scannable without reading a word: `✅` clean (no findings), `⚠️` only suspicions, `🐛` at least one confirmed bug — on the headline, the worst severity present; on a finding, its own. `➖` marks something explicitly skipped (a `Not covered:` line, or an individual check you didn't run). Icons decorate the existing tags/wording — `[bug]`, `[suspicion]`, `Not covered:` — they never replace them: keep both, so the report still greps cleanly and reads fine wherever emoji don't render. - **The chat message after disconnecting follows a fixed, compact template — not free-form prose, and not a second copy of `report.md`:** ``` Marionette · — : ✓ ✗ [] (up to 5; beyond that: "+N more in report.md") ∅ Not covered: — → /report.md · steps, screenshots ``` `` is a fixed, four-word vocabulary — like `[bug]`/`[suspicion]`, never translated regardless of the prompt's language: `PASS` (nothing found, full coverage), `SUSPECT` (only suspicion(s), no confirmed bug), `FAIL` (at least one confirmed bug), `BLOCKED` (coverage was cut short by something that stopped you continuing — bad test data, a stuck flow, a missing prerequisite; use this over `FAIL`/`PASS` when *that's* the story, even if you also found a bug along the way). Each `✗` line's `[type]` is `bug`, `blocker`, `ux`, or `suspicion` — `bug`/`blocker`/`ux` corresponds to a `[bug]`-tagged finding in `report.md` (evidence-backed); `suspicion` corresponds to a `[suspicion]`-tagged one. Rows are optional beyond the headline and the `→` line: skip `✓` if nothing worked, skip `✗`/`∅` for a `PASS`. The step/screenshot counts on the `→` line come from `steps.md`, same as `Tested:` in `report.md` — not a re-estimate. This uses `✓`/`✗`/`∅`/`→`, not the `✅`/`⚠️`/`🐛`/`➖` from `report.md`, deliberately: this message prints as raw terminal text, where a plain Unicode symbol renders as one predictable, monochrome character everywhere, while a full-color emoji can render at an inconsistent width or not at all depending on the terminal. `report.md` is a file usually opened in an editor or on GitHub, where colored emoji render cleanly — different consumption context, different choice; don't mix the two sets. - **No session directory (see *Reporting what you found*, above)? There's no `report.md` and no contract to follow** — give the same shape of chat message anyway (verdict, `✓`/`✗`/`∅` lines), just drop the `→` line (no path to give) and don't invent a `Tested:`-style step count — you have no `steps.md` to render one from truthfully. This is what `report.md` itself looks like (shown in English here; write yours in whatever language the prompt used) — a confirmed bug, a clean run, then a suspicion: ``` 🐛 Marionette report — 2 findings · User Profile Tested: profile view, edit form, save flow (14 steps, 3 screenshots) 1. 🐛 [bug] Date of birth accepts future dates — no validation Repro: Profile → Edit → DOB = 2099-01-01 → Save Evidence: step 9, log "saved dob=2099-01-01" 2. 🐛 [bug] Save does not persist — changes lost after restart Repro: Profile → Edit → Name = "X" → Save → hot_restart → Profile Evidence: step 14, screenshots/02-after-restart.png ``` ``` ✅ Marionette report — no issues · Checkout Tested: cart, address form, payment, confirmation (18 steps) Checks: field validation, back navigation, dark mode, text scale 2.0 ``` A run with only a suspicion (no hard evidence of a real bug) gets `⚠️` instead, on both the headline and the finding itself: ``` ⚠️ Marionette report — no confirmed issues, 1 suspicion · Settings Tested: notification toggles, language picker, theme toggle (9 steps, 2 screenshots) 1. ⚠️ [suspicion] Dark mode toggle reverts to light after hot_restart Repro: Settings → enable Dark mode → hot_restart → Settings Evidence: screenshots/02-after-restart.png shows light theme again ➖ Not covered: account deletion, data export — out of scope for this run ``` The chat message for those three, plus a `BLOCKED` case none of them happen to show (this has a different shape from `report.md`'s own headline — denser, with a verdict word, since it has to stand alone as the whole message): ``` Marionette · User Profile — FAIL: 2 bugs ✓ Profile view, edit form ✗ [bug] Date of birth accepts future dates — no validation ✗ [bug] Save does not persist — changes lost after restart → .marionette/sessions/user-profile-20260922T1412/report.md · 14 steps, 3 screenshots ``` ``` Marionette · Checkout — PASS: no issues ✓ Cart, address form, payment, confirmation, dark mode, text scale 2.0 → .marionette/sessions/checkout-20260922T1801/report.md · 18 steps ``` ``` Marionette · Settings — SUSPECT: theme toggle may not persist ✓ Notification toggles, language picker ✗ [suspicion] Dark mode toggle reverts to light after hot_restart ∅ Not covered: account deletion, data export — out of scope for this run → .marionette/sessions/settings-20260922T1601/report.md · 9 steps, 2 screenshots ``` ``` Marionette · Password reset — BLOCKED: reset email never arrives ✓ Request form, validation messages ✗ [blocker] "Send reset link" succeeds but no email logged or received after 3 attempts ∅ Not covered: setting a new password, confirmation screen — flow can't proceed past the email step → .marionette/sessions/password-reset-20260922T1601/report.md · 11 steps, 1 screenshot ``` `.marionette/` is gitignored by default (the server writes that `.gitignore` itself the first time a session is created). Committing a session directory — attaching a report to a PR, say — is a deliberate choice you make explicitly, never something to do as a matter of course. ## Custom extensions Apps can expose their own actions via `registerMarionetteExtension` (route navigation, seeding test data, feature-flag toggles — anything the generic tools can't express). Call `list_custom_extensions` before assuming an action isn't possible. Extensions that declare a scalar `inputSchema` show up as first-class, individually named tools (sanitized to `[a-z0-9_-]`, e.g. `appNavigation.goToPage` → `app_navigation_go_to_page`) with validated arguments; schema-less ones are only reachable through the generic `call_custom_extension`, passing the *real* (unsanitized) name and key-value args. ## Troubleshooting | Symptom | Likely cause | Fix | |----------------------------------------------------------------------------------------------------|-----------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------| | "Not connected to any app" | No successful `connect` yet in this session | `connect` before anything else — see *Connecting*, above | | MCP `connect` fails with a version mismatch | `marionette_mcp` and `marionette_flutter` are different versions (the CLI doesn't check this) | See *Version alignment*, above | | Custom buttons/fields don't show up, or `tap(text:)`/`scroll_to(text:)` can't find a visible label | Widget type or text isn't recognized | See *Custom design system?*, above | | A gesture or `enter_text` hits the wrong copy of a repeated widget | The key repeats across identical subtrees; the first match wins | Add `ancestor_keys` — see *Repeated keys*, above | | `Scope element with key "…" (ancestor_keys[N]) not found` | That wrapper key is misspelled, not built yet, or not inside the previous one | Check with `get_interactive_elements(ancestor_keys: [...])`; off-screen rows need `scroll_to` first — see *Repeated keys*, above | | `get_logs` says no collector configured | No `LogCollector` wired up | See *Logs for `get_logs`*, above | | `set_device_config` returns setup instructions instead of succeeding | App hasn't opted in | See *Device-config sweeps*, above, then hot restart | | Binding assertion error on startup (often under `flutter test`) | Two `WidgetsBinding`s initialized | See *The single-binding rule*, above | | Session directory lands somewhere unexpected (e.g. not the project root) | No `MARIONETTE_SESSION_DIR` set and the server's own working directory isn't reliably the repo root | Pass `session_dir` to `connect` (`--session-dir` on the CLI) explicitly — see *Reporting what you found*, above | | Nothing above applies, and it's a release build | Marionette needs the VM Service | Not supported by design — see *When not to use* | ## CLI fallback When MCP tools aren't available, install `marionette_cli` and run its self-describing reference once per session before driving anything: ```bash dart pub global activate marionette_cli marionette help-ai ``` That prints every command's syntax, expected output, and exit codes — treat it as the authoritative low-level reference; everything in the capability table above maps onto it one-for-one (`tap` ↔ `tap --key/--identifier/--text/--type/--x/--y`, `get_interactive_elements` ↔ `get-interactive-elements`, etc.), plus `record-video`, which only exists on this side. `ancestor_keys` is `--ancestor-keys`, repeated once per key, outermost first: `tap --key cell.joinButton --ancestor-keys session_2 --ancestor-keys grid.cell_3`. Use `--uri ` for a one-off session and `register ` + `-i ` for repeated interaction with the same app; `marionette doctor` checks connectivity of every registered instance and `unregister` cleans up stale ones. Session reports work the same way here: pass `--session ` (the CLI's equivalent of `session_title`) consistently across a script's invocations to log a multi-step run into one session directory instead of a fresh, untitled one per command — see *Reporting what you found*, above. ## Keeping this skill in sync This file must track the package release it ships with, and the capability table above should never disagree with `marionette help-ai`'s own output — both are generated from the same command set. If a tool or CLI command is added, renamed, or changes behavior, update the table and any affected practice/troubleshooting entry, and check `marionette help-ai` alongside it: if one mentions a command the other doesn't, fix both together rather than picking one to trust going forward.