--- name: smoke-tests description: Use when running VS Code smoke tests or working on smoke-test CI steps. Covers npm run smoketest / smoketest-no-compile, grep filtering tests, and a temporary repeat-loop technique for tracking down flaky smoke tests in CI. --- # Running Smoke Tests Smoke tests live in `test/smoke/` and drive a full VS Code instance (Electron, web, or remote) through end-to-end user flows. ## Scripts - `npm run smoketest` โ€” compiles the smoke tests first (`test/smoke`), then runs them. - `npm run smoketest-no-compile` โ€” runs the already-compiled smoke tests. CI uses this after an explicit compile step. Both forward extra arguments after `--` to the runner (`test/smoke/test/index.js`). ## Common options | Option | Description | |--------|-------------| | `-g ` (alias `-f`) | Grep filter on test/suite titles (mocha `grep`). | | `--build ` | Run against a packaged build instead of the compiled-from-source dev build. | | `--tracing` | Capture Playwright traces (and screenshots on failure). | | `--web` | Run the browser smoke tests instead of Electron. | | `--headless` | Headless browser (used with `--web`). | | `--remote` | Run the remote smoke tests. | ```bash # Run everything (Electron, from source) npm run smoketest # Run only a subset of suites by name, with tracing (replace with your suite, e.g. "Agents Window") npm run smoketest -- -g "" --tracing # Run against a packaged build (CI style) npm run smoketest-no-compile -- --tracing --build "/path/to/VSCode-darwin-arm64/Code - OSS.app" ``` The `-g` pattern matches against test/suite titles. For example, `-g "Agents Window"` matches all three Agents Window suites (`Agents Window`, `Agents Window (local AgentHost)`, and `Agents Window (local AgentHost, SDK sandbox)`); use whatever substring identifies the suite(s) you care about. The runner exits non-zero if any test fails, so a `0` exit code means every selected test passed. ## Temporarily looping a suite to hunt flaky CI tests When a smoke test fails intermittently only in CI, a useful technique is to **temporarily** run the suspect suite many times in a row and fail on the first failure. This reproduces the flake under the real CI environment and captures its traces/screenshots, instead of waiting for it to recur naturally across unrelated PRs. This is a debugging aid, **not a permanent CI fixture**: - Add it on a throwaway branch, push, and let CI run it. Iterate until you reproduce (and then fix) the flake. - **Remove the loop before merging** โ€” leaving it in would add ~an hour per platform to every run. - It is **not specific to any one suite**. Point the `-g` filter at whichever suite you are investigating (the examples below use `"Agents Window"`, but substitute your own). ### Where to add it Drop the loop next to the existing Electron smoke step, gated on the same condition, in the test step(s) for the platform(s) where the flake reproduces: **GitHub PR workflows** (run from source, no `--build`): - `.github/workflows/pr-linux-test.yml` (bash; sets `DISPLAY: ":10"`) - `.github/workflows/pr-darwin-test.yml` (bash; no `DISPLAY`) - `.github/workflows/pr-win32-test.yml` (PowerShell) **Azure DevOps test steps** (run against the packaged build via `--build`): - `build/azure-pipelines/linux/steps/product-build-linux-test.yml` - `build/azure-pipelines/darwin/steps/product-build-darwin-test.yml` - `build/azure-pipelines/win32/steps/product-build-win32-test.yml` ### Shape Loop N iterations (e.g. 20) and abort on the first failing run. Give it a generous timeout โ€” N sequential runs of a ~3-minute suite can take roughly an hour. Bash (Linux/macOS): ```yaml # TEMPORARY: loop the suite to reproduce a flaky failure. Remove before merge. # Replace with the suite you're investigating (e.g. "Agents Window"). - name: ๐Ÿงช Smoke test flakiness probe (TEMPORARY) if: ${{ inputs.electron_tests }} timeout-minutes: 60 run: | for i in $(seq 1 20); do echo "::group::Smoke probe run $i/20" npm run smoketest-no-compile -- --tracing -g "" || { echo "::error::Smoke test failed on run $i/20"; exit 1; } echo "::endgroup::" done ``` PowerShell (Windows) checks `$LASTEXITCODE` after each run and `exit 1` on failure. The AzDO variants use `set -e` (bash) / `$LASTEXITCODE` (pwsh) for fail-fast and append `--build ""`. ### Why fail-fast The loop is a probe: the first failure is the signal. Stopping immediately preserves the failing run's traces/screenshots (under the logs artifact) and avoids burning ~an hour of agent time finishing a run that has already proven flaky. ## Debugging CI smoke failures Both CI systems publish the smoke runner's per-platform logs (the `.build/logs` directory) as a downloadable artifact. The artifact's internal layout is identical on both โ€” only the artifact name and the download tool differ. ### Downloading the logs artifact #### GitHub Actions The GitHub PR workflows upload the artifact as `logs----`, where `` is `linux` / `macos` / `windows`, `` is `electron` / `browser` / `remote`, and `` is the run attempt (e.g. `logs-macos-arm64-electron-1`). The run id is the number in the run/job URL โ€” for `โ€ฆ/actions/runs//job/` use ``. Download with the `gh` CLI: ```bash # A specific artifact into ./logs gh run download -n logs---- -D ./logs # Or every artifact from the run gh run download ``` `gh run view ` lists the run's jobs/artifacts; the run summary page in the browser also has an **Artifacts** section at the bottom. #### Azure DevOps The artifact name depends on which pipeline produced it: - **Product build** (`product-build-.yml`): `logs---` โ€” no suite segment, e.g. `logs-macos-arm64-1`. - **Suite-split CI build** (`product-build--ci.yml`): `logs----` โ€” the `` segment is `lower(VSCODE_TEST_SUITE)` (e.g. `electron`), so e.g. `logs-macos-arm64-electron-1` (same shape as GitHub). `` is `linux` / `macos` / `windows`, `` is `x64` / `arm64`, and `` is `$(System.JobAttempt)`. Download with the Azure CLI: ```bash az pipelines runs artifact download \ --org --project \ --run-id --artifact-name \ --path ./logs ``` For the VS Code build that is `--org https://dev.azure.com/monacotools --project Monaco`; see the `azure-pipelines` skill for finding the ``. ### Inside the artifact Under `smoke-tests-/` (`smoke-tests-electron/`, `smoke-tests-browser/`, or `smoke-tests-remote/`, matching the suite that ran): - `smoke-test-runner.log` โ€” the mocha driver output plus, for suites that use the mock LLM server, its verbose request/response bodies (look for `request body:`). On a Copilot CLI / Copilot session failure it also carries a tail of the captured Copilot runtime logs (see below). - `_suite_/copilot-runtime-logs/process-*.log` โ€” the Copilot runtime (`@github/copilot` CLI) process logs, captured by `dumpFailureDiagnostics` when a Copilot-runtime session fails. **Check these first for a hang or "Timed out waiting for response"**: they are the SDK/CLI's own account of what it did (startup, auth, model request, turn lifecycle, and any panic / out-of-order event / protocol error) and explain a timeout that the test error alone does not. **Agent Host** sessions (Agents Window / local AgentHost) write a full log run at `trace` (`chat.agentHost.copilotSdk.logLevel`). Chat Sessions editor (Copilot CLI / Claude) and Local sessions run the SDK in-process and write only a minimal startup log here (`Server started, waiting for requests`) โ€” enough to tell whether the runtime came up; their detailed model/turn diagnostics are in the `GitHub Copilot Chat.log` below. (Claude / Codex sessions use a different runtime and are not captured here.) - `_suite_/window2/exthost//โ€ฆlog` โ€” per-suite extension-host logs (e.g. `GitHub.copilot-chat/GitHub Copilot Chat.log`). Many diagnostics are gated behind a setting the suite enables in its `before` hook, so check the suite's setup if an expected log line is missing. - `_suite_/playwright-screenshot-*.png` โ€” last-frame screenshot captured when a test fails (only when the suite ran with `--tracing`). `` is the mocha suite title with non-word characters replaced by `_`. See also the `code-oss-logs` skill. ## Distinction from other test types - **Unit tests** (`.test.ts`) โ†’ `scripts/test.sh` / `runTests` tool (see the `unit-tests` skill). - **Integration tests** (`.integrationTest.ts` + extension tests) โ†’ `scripts/test-integration.sh` (see the `integration-tests` skill). - **Smoke tests** (`test/smoke/`) โ†’ `npm run smoketest` โ€” full end-to-end UI flows.