Your AI-powered QA engineer. Crawl docs, explore your app, learn flows, and generate & execute tests — from a Chrome side panel or via a Model Context Protocol server.
Pathfinder ships in two flavors. Use one or both — they share the same engine.
Side-panel UI in Chrome. Install once, point at any web app, click through the tabs. Best for hands-on QA work and ad-hoc exploration.
Standalone Node.js process exposing 20+ tools over the Model Context Protocol. Lets Claude Desktop, Cursor, Cline, or any custom AI agent drive Pathfinder programmatically — no browser required.
Click the ⚙ Settings icon → choose your provider (OpenAI / Anthropic / Google) and paste an API key. Defaults to the latest models — GPT-5, Claude Sonnet 4.6, or Gemini 3 Pro.
Go to the Knowledge tab → paste your help/docs URL → click Crawl. Pathfinder reads every page, embeds the content, and builds a vector knowledge base used to ground test generation.
Open the page you want to test → Explore tab → Start Exploration. AI navigates your app and builds a page graph. Then Learn Flows from Exploration extracts user journeys. Finally, Generate in the Tests tab to write executable cases.
Paste the root URL of your help site or docs. Pathfinder crawls every linked page, extracts text, and stores embeddings.
Set depth (1–5), or flip on Single page only to scan just the active tab URL. Hit Start Exploration. Pathfinder clicks visible elements, opens modals, captures forms, fields, wizards, data tables, API calls, and accessibility issues.
After exploring, click Learn Flows from Exploration. Pathfinder partitions the graph by sub-app (e.g. /admin vs /learner), runs LLM extraction in parallel, and merges the results.
Click Generate on any flow. AI writes positive, negative & edge-case tests grounded in observed form fields. Constraint-based boundary tests (max-length+1, required-empty, format violations) are added deterministically.
Per-step status, screenshots on failure, healing attempts, network requests, and full error messages.
.spec.tsThree on-demand analyses on top of your captured data:
Every field is documented in the settings reference below.
Save named presets bundling planning mode, target origin, headers, and storage state. Apply a preset before "Run All" to switch between staging / prod / sandbox in one click.
Five features that exist for one reason: a test suite you cannot trust is worse than no suite at all. Each one targets a specific way UI tests waste your time.
When a step cannot find its element, Pathfinder tries to repair the selector instead of failing:
The vision tier only runs when a screenshot exists and the three cheaper tiers have failed, so it costs nothing on a normal run. It is what rescues icon-only buttons and canvas widgets that have no useful DOM text. Every heal is recorded on the result with the tier that fixed it.
How to use it: click the repeat icon on any test row in the Tests tab. The test runs three times back to back.
The last distinction matters: a test that always fails is a real bug report, and hiding it would be the wrong call. Only genuinely non-deterministic tests get quarantined. Clear the badge on the row to bring a test back into normal runs.
How to use it: open a failed result. The failure banner offers Resume from step N, and every row in the execution timeline has its own resume button — so you can bisect a long flow instead of only retrying the failure point.
{{placeholder}}How to use it: expand a test in the Tests tab and paste CSV into Attach data set. The header row names the variables; reference them in any step value as {{column_name}}.
email,password,expect
alice@acme.com,hunter2,success
bad@acme.com,wrong,error
loop_index and loop_iteration are reservedCSS-in-JS frameworks generate class names that change on every build — sc-1e593sq-0, css-1x2y3z, Button_root__a1b2c, jss42. A selector built on one is guaranteed to break at the next deploy.
Pathfinder rejects them at three points: when generating a selector, when healing one, and in the prompts the model is given. Nothing to configure. The testability report flags any element that has only a hashed class as a gap that will break, and sorts those to the top — that is your list of elements to give a data-testid.
Everything Pathfinder produces can leave the browser. Exports live on the Results tab toolbar.
| Export | What you get | Use it for |
|---|---|---|
| JSON | Full result objects — steps, timings, healing attempts, network | Custom dashboards, archival |
| HTML report | Self-contained page with per-step evidence | Sharing a run with someone who has no extension |
| JUnit XML | Standard CI test-report format | Jenkins, GitLab CI, GitHub Actions test summaries |
| Playwright | A real .spec.ts file | Checking tests into your repo and running them in CI |
| Screen recording | Video of the run | Bug reports, reproducing a race |
| TestRail | Status, duration, error and failure screenshot on each case | Teams whose source of truth is TestRail |
getByTestId, getByRole, or a CSS locator, following the same preference order the executor uses. No AI is involved, so the same run always produces the same file.
Every setting in the panel, what it changes, and what it defaults to.
| Setting | What it does | Default |
|---|---|---|
| AI Provider | Which API your key belongs to: OpenAI, Anthropic, or Google AI. Switching it resets the model fields and clears the verified model list. | OpenAI |
| API Key | Your own key. Stored in chrome.storage.local and sent only to the provider you chose — never to any Pathfinder server, because there isn't one. | — |
| Test key & load models | Asks the provider which models this key can call. Doubles as the key test: a provider that returns its catalogue has accepted the credential. Turns the two model fields into dropdowns of exactly what you are entitled to use. | — |
| Model | The chat model used for planning, generation and healing. Free text until you test the key, so proxies and preview models still work. | gpt-5 |
| Embedding mode | Local runs all-MiniLM-L6-v2 on your machine — free, no key, ~23 MB once. API uses your provider's embedding endpoint. | Local |
| Embedding Model | Which embedding model to call in API mode. Anthropic has no embedding API — use Local, or OpenAI/Google. | text-embedding-3-small |
| Setting | What it does | Default |
|---|---|---|
| AI-Guided Exploration | Agent mode. The model ranks which elements are worth clicking on each page — one extra AI call per page, noticeably better exploration graphs. Turn it off to explore cheaply. | On |
| Explore Depth | How many links deep the crawler follows from the start URL. | 5 |
| Max Crawl Pages | Hard ceiling on pages visited in one crawl — your protection against a site with infinite pagination. | 200 |
| Describe Images (Vision AI) | Sends images found while crawling to a vision model for a text description. Adds AI cost per image, so it is off unless your app's meaning is carried by images. | Off |
| Setting | What it does | Default |
|---|---|---|
| Planning mode | Auto tries interactive first and falls back to single-shot on retry. Interactive walks the live app step by step before writing the plan — slowest, most accurate. Single-shot generates every step from one DOM snapshot — fastest, cheapest. | Auto |
| Test concurrency | How many tests run at once, each in its own tab. Raise it for speed; keep it at 1 if your app cannot tolerate parallel sessions. Capped at 4. | 1 |
| Test personality | Biases generation toward a style — balanced, happy path, aggressive edge, security focused, accessibility first, performance minded, or custom. Each changes temperature, the positive/negative/edge ratio and tests generated per flow. Custom takes up to 200 characters of your own instructions; empty custom behaves as balanced. | Balanced |
| Execution presets | Named run profiles bundling start URL, persona label, setup steps, authentication checks and whether a logged-in session is required. Apply one before a run to switch between staging / prod / sandbox. | — |
| Setting | What it does | Default |
|---|---|---|
| Webhook URL / Trigger | POST run summaries to Slack, Linear or Jira. Trigger on every run, on failures only, or never. | Off |
| TestRail | Host, email, API key and a default run id. Test connection verifies them before you rely on them. | — |
| Retain API response bodies | Debug aid for contract findings: keeps redacted response bodies for 24 hours so you can see why a check fired. Off by default because retention is a decision, not a convenience — switching it off deletes what was already stored. | Off |
| Theme | Dark or light. Toggle with the sun/moon icon in the header. | Light |
| Clear all data | Wipes crawled pages, flows and results from this browser. Irreversible, and it does not touch your API key. | — |
pathfinder_* in chrome.storage.local, flows and results in the IndexedDB database pathfinder_db. Pathfinder has no backend. The only outbound requests are to the AI provider you chose, the app you are testing, and any integration you configured yourself.
chrome://extensions.
includeDangerous: true in the START_EXPLORATION message (or wire a UI toggle if you've added one).
pathfinder_* in chrome.storage.local; flows live in IndexedDB pathfinder_db.
The MCP package lets external AI agents call Pathfinder's engine over the Model Context Protocol. It runs as a local Node.js process and exposes 20+ tools.
# From the Pathfinder repo root
cd MCP
# Install Playwright + dependencies
npm run setup
# Start MySQL (used for KB + results)
docker run -d \
--name pathfinder-mysql \
-e MYSQL_ROOT_PASSWORD=pathfinder \
-e MYSQL_DATABASE=pathfinder \
-e MYSQL_USER=pathfinder \
-e MYSQL_PASSWORD=pathfinder \
-p 3307:3306 \
mysql:8.0
# Configure .env
cat > .env <<'EOF'
AI_PROVIDER=anthropic
ANTHROPIC_API_KEY=sk-ant-...
MYSQL_HOST=127.0.0.1
MYSQL_PORT=3307
MYSQL_USER=pathfinder
MYSQL_PASSWORD=pathfinder
MYSQL_DATABASE=pathfinder
EOF
# Build
npm run build
Claude Desktop — edit ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"pathfinder": {
"command": "node",
"args": ["--env-file=.env", "/absolute/path/to/Pathfinder/MCP/dist/index.js"],
"cwd": "/absolute/path/to/Pathfinder/MCP"
}
}
}
Cursor — Settings → MCP → Add new server, point to the same command.
Cline / custom agents — register the server in your client's MCP config; it speaks stdio JSON-RPC out of the box.
Once registered, ask the AI: "Use Pathfinder to crawl https://docs.example.com, then explore https://app.example.com, learn flows, and run smoke tests headless." The agent will chain crawl_knowledge → explore_app → learn_flows → run_one_liners automatically.
Approximate per-flow and per-test-run costs for a mid-sized site (≈240 flows / 1,440 generated tests, weekly runs):
| Model | In $/1M | Out $/1M | Generation | 4× weekly run | Total / month |
|---|---|---|---|---|---|
| GPT-5-mini | 0.25 | 2 | ~$2.5 | ~$25 | ~$28 |
| GPT-5 | 1.25 | 10 | ~$12 | ~$140 | ~$150 |
| Gemini 3 Pro | 2 | 10 | ~$15 | ~$165 | ~$180 |
| Claude Sonnet 4.6 | 3 | 15 | ~$26 | ~$230 | ~$255 |
| Claude Opus 4.7 | 15 | 75 | ~$130 | ~$1,120 | ~$1,250 |
Levers that change the bill: planningMode: 'preplan' drops execution cost ~70%; useLocalEmbeddings: true zeros embedding spend; describeImages: true adds 2–4× via vision calls. Constraint-based tests (up to 25 per flow) cost zero.
all-MiniLM-L6-v2 for those.
/login, /signin, /auth, /sso, /oauth, /session. If it gets bounced to one of these mid-run, it aborts immediately with a clear message instead of looping on the login form. Sign back in and re-run.
storage_state_path JSON; capture it once via capture_auth or import_chrome_cookies, then point every run at it.
export_knowledge / export_explore.
pathfinder_* in chrome.storage.local and IndexedDB pathfinder_db. MCP: clear_knowledge + clear_explore.