How to use Pathfinder

Your AI-powered QA engineer. Crawl docs, explore your app, learn flows, and generate & execute tests — from a Chrome side panel or via a Model Context Protocol server.

On this page

  1. Two ways to use Pathfinder
  2. Quick start (browser extension)
  3. Tab-by-tab guide
  4. Reliability & debugging features
  5. Export & integrations
  6. Complete settings reference
  7. Advanced features
  8. MCP server — for AI agents
  9. Full MCP tool reference
  10. Cost & model selection
  11. FAQ

Two ways to use Pathfinder

Pathfinder ships in two flavors. Use one or both — they share the same engine.

🧩 Browser extension

Side-panel UI in Chrome. Install once, point at any web app, click through the tabs. Best for hands-on QA work and ad-hoc exploration.

  • Visual side panel with tabs
  • Live exploration with screenshots
  • One-click test generation & runs
  • Self-healing selector retries
  • All data stays in your browser

🤖 MCP server

Standalone Node.js process exposing 20+ tools over the Model Context Protocol. Lets Claude Desktop, Cursor, Cline, or any custom AI agent drive Pathfinder programmatically — no browser required.

  • Headless Playwright execution
  • MySQL-backed knowledge & results
  • Auth via captured Playwright sessions
  • 20+ tools: crawl, explore, run, analyze
  • Designed for agentic workflows
Quick Start (Browser extension — 3 steps)
1

Set up your AI

Click the ⚙ Settings icon → choose your provider (OpenAI / Anthropic / Google) and paste an API key. Defaults to the latest models — GPT-5, Claude Sonnet 4.6, or Gemini 3 Pro.

⚙ Settings → Provider → Save key
2

Crawl your documentation

Go to the Knowledge tab → paste your help/docs URL → click Crawl. Pathfinder reads every page, embeds the content, and builds a vector knowledge base used to ground test generation.

3

Explore → Learn flows → Generate & run tests

Open the page you want to test → Explore tab → Start Exploration. AI navigates your app and builds a page graph. Then Learn Flows from Exploration extracts user journeys. Finally, Generate in the Tests tab to write executable cases.

How the pieces connect
💡
Pipeline thinking Knowledge crawl + app exploration → flow extraction → test generation → execution → results & analysis. Each tab is one stage. Most stages can be exported and re-imported, so you can crawl in one environment and run tests in another.
📚 Docs
🌐 Knowledge
🔍 Explore
🛡 Analysis
✅ Results
🧪 Tests
📋 Flows
Tab-by-tab guide
📚 Knowledge

Build your RAG knowledge base

Paste the root URL of your help site or docs. Pathfinder crawls every linked page, extracts text, and stores embeddings.

  • Up to 1,500 pages per crawl
  • Re-crawl is incremental — unchanged pages skipped via content hash
  • IVF vector index for fast similarity search
  • Export & import the KB for sharing
  • Optional local embeddings (free, on-device MiniLM)
🔍 Explore

Map your app autonomously

Set depth (1–5), or flip on Single page only to scan just the active tab URL. Hit Start Exploration. Pathfinder clicks visible elements, opens modals, captures forms, fields, wizards, data tables, API calls, and accessibility issues.

  • Single page only toggle — scope to current URL, no link following
  • Up to 1,500 elements detected, 75 click targets per page
  • 6-minute per-page time budget
  • Walks shadow DOM & same-origin iframes
  • Auth-wall detection — aborts if session expires mid-run
  • Export & import the page graph
📋 Flows

AI-learned user journeys

After exploring, click Learn Flows from Exploration. Pathfinder partitions the graph by sub-app (e.g. /admin vs /learner), runs LLM extraction in parallel, and merges the results.

  • Hard guarantee: ≥ 2 flows per page, no page missed
  • Auto-generated coverage fillers for under-represented pages
  • Multi-step wizard awareness (Next Step / Continue)
  • Bulk select + Generate (N) or Generate All
  • Bulk delete (selected or all)
🧪 Tests

Generate & run test cases

Click Generate on any flow. AI writes positive, negative & edge-case tests grounded in observed form fields. Constraint-based boundary tests (max-length+1, required-empty, format violations) are added deterministically.

  • Planning modes — pre-plan or live interactive replan
  • Test personalities — 6 presets plus your own custom prompt
  • Self-healing selectors with retry & replan
  • One-line tests — paste & expand manual checks
  • CSV bulk import / export, and TestRail run import
  • Attach a data set — run one test once per CSV row
  • Stability check — run N times to catch flaky tests
📊 Results

Pass / fail with evidence

Per-step status, screenshots on failure, healing attempts, network requests, and full error messages.

  • Step-by-step trace with timing
  • Screenshots on failure, plus screen recordings
  • Export JSON, HTML, JUnit XML, or Playwright .spec.ts
  • Push results to TestRail
  • Resume from any step instead of replaying the whole test
  • Slack / Linear / Jira webhook integrations
🛡 Analysis

Coverage, accessibility, contracts

Three on-demand analyses on top of your captured data:

  • API Coverage — which endpoints are tested vs untested (uses HAR from CDP)
  • Accessibility — WCAG audit of the active tab via the CDP a11y tree
  • API Contracts — validate captured responses against an OpenAPI spec
⚙ Settings

Provider, models, behavior

  • AI Provider — OpenAI, Anthropic, or Google AI
  • Test key & load models — verifies the key and lists the models it can call
  • Model / Embedding Model — picked from that verified list
  • Agent mode — AI-ranked clicks during explore
  • Planning mode — auto / single-shot / interactive
  • Test concurrency, Explore depth, Max crawl pages
  • Webhook and TestRail integrations

Every field is documented in the settings reference below.

🧱 Execution Presets

Reusable run profiles

Save named presets bundling planning mode, target origin, headers, and storage state. Apply a preset before "Run All" to switch between staging / prod / sandbox in one click.

Reliability & debugging features

Five features that exist for one reason: a test suite you cannot trust is worse than no suite at all. Each one targets a specific way UI tests waste your time.

🩹 Self-healing selectors

Four tiers, cheapest first

When a step cannot find its element, Pathfinder tries to repair the selector instead of failing:

  • 1. Alternatives — other selectors generated for the same element
  • 2. Similarity — nearest match in the current DOM by role, text and position
  • 3. AI — the model proposes a selector from the live DOM
  • 4. Vision — the model reads the failure screenshot

The vision tier only runs when a screenshot exists and the three cheaper tiers have failed, so it costs nothing on a normal run. It is what rescues icon-only buttons and canvas widgets that have no useful DOM text. Every heal is recorded on the result with the tier that fixed it.

🔁 Stability check

Catch a flaky test before it costs you a morning

How to use it: click the repeat icon on any test row in the Tests tab. The test runs three times back to back.

  • Stable — every attempt passed
  • Unstable — mixed results → the test is quarantined and skipped by "Run all"
  • Failing — every attempt failed → not quarantined

The last distinction matters: a test that always fails is a real bug report, and hiding it would be the wrong call. Only genuinely non-deterministic tests get quarantined. Clear the badge on the row to bring a test back into normal runs.

⏭ Resume from step

Stop replaying the first 27 steps

How to use it: open a failed result. The failure banner offers Resume from step N, and every row in the execution timeline has its own resume button — so you can bisect a long flow instead of only retrying the failure point.

  • Skipped steps are recorded as skipped, never as passed
  • Resuming past a step that captured a value fails with a clear message rather than typing a literal {{placeholder}}
🧮 Data-driven runs

One test, many rows

How to use it: expand a test in the Tests tab and paste CSV into Attach data set. The header row names the variables; reference them in any step value as {{column_name}}.

email,password,expect
alice@acme.com,hunter2,success
bad@acme.com,wrong,error
  • One result per row, each labelled with its row
  • Column names must be valid identifiers; loop_index and loop_iteration are reserved
  • Parse errors list every problem with its line number — you fix the file once, not five times
🧱 Build-hash guard

Selectors that survive a deploy

CSS-in-JS frameworks generate class names that change on every build — sc-1e593sq-0, css-1x2y3z, Button_root__a1b2c, jss42. A selector built on one is guaranteed to break at the next deploy.

Pathfinder rejects them at three points: when generating a selector, when healing one, and in the prompts the model is given. Nothing to configure. The testability report flags any element that has only a hashed class as a gap that will break, and sorts those to the top — that is your list of elements to give a data-testid.

Export & integrations

Everything Pathfinder produces can leave the browser. Exports live on the Results tab toolbar.

ExportWhat you getUse it for
JSONFull result objects — steps, timings, healing attempts, networkCustom dashboards, archival
HTML reportSelf-contained page with per-step evidenceSharing a run with someone who has no extension
JUnit XMLStandard CI test-report formatJenkins, GitLab CI, GitHub Actions test summaries
PlaywrightA real .spec.ts fileChecking tests into your repo and running them in CI
Screen recordingVideo of the runBug reports, reproducing a race
TestRailStatus, duration, error and failure screenshot on each caseTeams whose source of truth is TestRail
🎭
Playwright export — how to use it Run your tests, then click Playwright on the Results toolbar. You get a spec built from the steps that actually ran — including any selector that was self-healed mid-run — not from the original plan. Locators map to getByTestId, getByRole, or a CSS locator, following the same preference order the executor uses. No AI is involved, so the same run always produces the same file.
⚠️
Dropped items are shown, never hidden A few things have no Playwright equivalent — API-called / API-status assertions, for example. Rather than silently omitting them, the export lists exactly what was dropped in a banner above the results, so you know what the generated spec does not check.
🧾
TestRail — how to use it Add your host, email and API key under Settings → TestRail, then Test connection. Import a run's cases from Tests → Import → TestRail run id; re-importing the same run updates rather than duplicates. After a run, click TestRail on the Results toolbar to push status, elapsed time, the error message and the failure screenshot back to each case. If some cases fail to push, the rest still land and you are told exactly which ones did not.
🔔
Webhooks Set a webhook URL and a trigger (every run, on failure only, or off) in Settings. Payloads are formatted for Slack, Linear and Jira.
Complete settings reference

Every setting in the panel, what it changes, and what it defaults to.

AI configuration

SettingWhat it doesDefault
AI ProviderWhich API your key belongs to: OpenAI, Anthropic, or Google AI. Switching it resets the model fields and clears the verified model list.OpenAI
API KeyYour own key. Stored in chrome.storage.local and sent only to the provider you chose — never to any Pathfinder server, because there isn't one.—
Test key & load modelsAsks the provider which models this key can call. Doubles as the key test: a provider that returns its catalogue has accepted the credential. Turns the two model fields into dropdowns of exactly what you are entitled to use.—
ModelThe chat model used for planning, generation and healing. Free text until you test the key, so proxies and preview models still work.gpt-5
Embedding modeLocal runs all-MiniLM-L6-v2 on your machine — free, no key, ~23 MB once. API uses your provider's embedding endpoint.Local
Embedding ModelWhich embedding model to call in API mode. Anthropic has no embedding API — use Local, or OpenAI/Google.text-embedding-3-small
📐
Switching embedding mode invalidates your index Local vectors are 384-dimensional and API vectors are 1536. If you already crawled with one, clear all data and re-crawl after switching — the two cannot be compared.

Exploration & crawling

SettingWhat it doesDefault
AI-Guided ExplorationAgent mode. The model ranks which elements are worth clicking on each page — one extra AI call per page, noticeably better exploration graphs. Turn it off to explore cheaply.On
Explore DepthHow many links deep the crawler follows from the start URL.5
Max Crawl PagesHard ceiling on pages visited in one crawl — your protection against a site with infinite pagination.200
Describe Images (Vision AI)Sends images found while crawling to a vision model for a text description. Adds AI cost per image, so it is off unless your app's meaning is carried by images.Off

Test generation & execution

SettingWhat it doesDefault
Planning modeAuto tries interactive first and falls back to single-shot on retry. Interactive walks the live app step by step before writing the plan — slowest, most accurate. Single-shot generates every step from one DOM snapshot — fastest, cheapest.Auto
Test concurrencyHow many tests run at once, each in its own tab. Raise it for speed; keep it at 1 if your app cannot tolerate parallel sessions. Capped at 4.1
Test personalityBiases generation toward a style — balanced, happy path, aggressive edge, security focused, accessibility first, performance minded, or custom. Each changes temperature, the positive/negative/edge ratio and tests generated per flow. Custom takes up to 200 characters of your own instructions; empty custom behaves as balanced.Balanced
Execution presetsNamed run profiles bundling start URL, persona label, setup steps, authentication checks and whether a logged-in session is required. Apply one before a run to switch between staging / prod / sandbox.—

Integrations & data

SettingWhat it doesDefault
Webhook URL / TriggerPOST run summaries to Slack, Linear or Jira. Trigger on every run, on failures only, or never.Off
TestRailHost, email, API key and a default run id. Test connection verifies them before you rely on them.—
Retain API response bodiesDebug aid for contract findings: keeps redacted response bodies for 24 hours so you can see why a check fired. Off by default because retention is a decision, not a convenience — switching it off deletes what was already stored.Off
ThemeDark or light. Toggle with the sun/moon icon in the header.Light
Clear all dataWipes crawled pages, flows and results from this browser. Irreversible, and it does not touch your API key.—
🔐
Where your data lives Everything is local: settings and knowledge under pathfinder_* in chrome.storage.local, flows and results in the IndexedDB database pathfinder_db. Pathfinder has no backend. The only outbound requests are to the AI provider you chose, the app you are testing, and any integration you configured yourself.
🔓
Host permissions are asked for, not assumed The extension ships with no site access. It asks for a specific origin the first time it needs one — the app you are crawling, your TestRail host, your AI provider when you test a key. You can review and revoke these any time in chrome://extensions.
Advanced features
🧭
Single-page exploration Toggle "Single page only" in Explore to scan just the current tab URL. Useful for capturing a single complex form (e.g. a wizard) without following links across the rest of the app.
⚠️
Aggressive / destructive exploration (opt-in) By default, buttons labelled Delete / Remove / Logout / Cancel subscription are skipped to protect your data. To capture destructive flows on a sandbox account, send includeDangerous: true in the START_EXPLORATION message (or wire a UI toggle if you've added one).
🔁
Re-explore a single page On the page graph, click the refresh icon next to any node to re-scan that page only — useful when you've changed a form and want fresh field metadata without re-running the full crawl.
📦
Export & import everything Knowledge base, exploration graph, and learned flows can each be exported to JSON and imported back into another browser, environment, or teammate's install. Keys are scoped to pathfinder_* in chrome.storage.local; flows live in IndexedDB pathfinder_db.
🛡
Test personalities Pick Settings → Test Personality to bias generation toward a style. Each one changes the AI temperature, the positive/negative/edge mix, and how many tests a flow yields: Balanced (default), Happy path, Aggressive edge, Security focused, Accessibility first, Performance minded, or Custom — up to 200 characters of your own instructions, injected into every generation prompt. Leave Custom empty and generation falls back to Balanced.

MCP server — for AI agents

The MCP package lets external AI agents call Pathfinder's engine over the Model Context Protocol. It runs as a local Node.js process and exposes 20+ tools.

🤖
What "MCP" means Model Context Protocol is an open standard from Anthropic that lets AI agents call tools running on your machine. Once Pathfinder MCP is configured, any MCP-compatible client (Claude Desktop, Cursor, Cline, custom Python/TS agents) can crawl docs, explore apps, run tests, and read results — without you opening a browser.

1. One-time setup

# From the Pathfinder repo root
cd MCP

# Install Playwright + dependencies
npm run setup

# Start MySQL (used for KB + results)
docker run -d \
  --name pathfinder-mysql \
  -e MYSQL_ROOT_PASSWORD=pathfinder \
  -e MYSQL_DATABASE=pathfinder \
  -e MYSQL_USER=pathfinder \
  -e MYSQL_PASSWORD=pathfinder \
  -p 3307:3306 \
  mysql:8.0

# Configure .env
cat > .env <<'EOF'
AI_PROVIDER=anthropic
ANTHROPIC_API_KEY=sk-ant-...
MYSQL_HOST=127.0.0.1
MYSQL_PORT=3307
MYSQL_USER=pathfinder
MYSQL_PASSWORD=pathfinder
MYSQL_DATABASE=pathfinder
EOF

# Build
npm run build

2. Wire up your MCP client

Claude Desktop — edit ~/Library/Application Support/Claude/claude_desktop_config.json:

{
  "mcpServers": {
    "pathfinder": {
      "command": "node",
      "args": ["--env-file=.env", "/absolute/path/to/Pathfinder/MCP/dist/index.js"],
      "cwd": "/absolute/path/to/Pathfinder/MCP"
    }
  }
}

Cursor — Settings → MCP → Add new server, point to the same command.

Cline / custom agents — register the server in your client's MCP config; it speaks stdio JSON-RPC out of the box.

3. Try it from your agent

Once registered, ask the AI: "Use Pathfinder to crawl https://docs.example.com, then explore https://app.example.com, learn flows, and run smoke tests headless." The agent will chain crawl_knowledge → explore_app → learn_flows → run_one_liners automatically.

Full MCP tool reference (20 tools)

Test execution

run_one_liners
Expand and execute one-liner test descriptions against a target app. Returns HTML report.
run_csv
Run or expand tests from CSV (one per line, or with title/type/context/start_url headers).
expand_tests
Expand one-liners into detailed steps without executing.
cancel_operation
Cancel an ongoing crawl, explore, run, or all of them.

Knowledge base

crawl_knowledge
Crawl a docs site (depth + max_pages) into the RAG knowledge base.
export_knowledge
Export the KB as JSON (file_path or inline).
import_knowledge
Replace the current KB with a previously exported snapshot.
clear_knowledge
Wipe the KB. Use before re-crawling or switching projects.

App exploration

explore_app
Headless Playwright explore — maps pages, forms, modals, navigation. Supports storage_state_path for auth.
export_explore
Export the page interaction graph as JSON.
import_explore
Import a previously exported graph — run tests against any environment using exploration from another.
clear_explore
Wipe the exploration graph.
get_graph
Return the current graph in human-readable format.

Flows & results

learn_flows
Extract user workflows from the exploration graph using AI.
get_flows
List all learned flows.
get_results
Fetch test run results + HTML report for a specific run_id.

Authentication

capture_auth
Open a real browser so you can log in manually; saves cookies+localStorage to a JSON file for headless reuse.
import_chrome_cookies
Read cookies from your existing Chrome / Brave / Arc profile into a Playwright storage state file.

Agent memory & introspection

remember
Store a persistent memory (selector heal, step pattern, tenant quirk, timing profile, test outcome).
recall
Search persistent memories by keyword or category.
pathfinder_version
Return server, protocol, and per-tool schema versions for client/server drift detection.

Cost & model selection

Approximate per-flow and per-test-run costs for a mid-sized site (≈240 flows / 1,440 generated tests, weekly runs):

Model In $/1M Out $/1M Generation 4× weekly run Total / month
GPT-5-mini0.252~$2.5~$25~$28
GPT-51.2510~$12~$140~$150
Gemini 3 Pro210~$15~$165~$180
Claude Sonnet 4.6315~$26~$230~$255
Claude Opus 4.71575~$130~$1,120~$1,250

Levers that change the bill: planningMode: 'preplan' drops execution cost ~70%; useLocalEmbeddings: true zeros embedding spend; describeImages: true adds 2–4× via vision calls. Constraint-based tests (up to 25 per flow) cost zero.

Common questions
✅
Do I need an API key? Yes for chat-driven steps (explore agent mode, flow learning, test generation, planning). Embeddings can be local (free) by enabling Local (Free) in Settings — Pathfinder uses on-device all-MiniLM-L6-v2 for those.
🔄
Can I re-crawl after content updates? Yes — Pathfinder hashes each page's content. Unchanged pages are skipped instantly; only updated pages are re-embedded. Re-crawls of large doc sites usually finish in seconds.
🛑
What happens if my session expires mid-explore? Pathfinder detects URLs matching /login, /signin, /auth, /sso, /oauth, /session. If it gets bounced to one of these mid-run, it aborts immediately with a clear message instead of looping on the login form. Sign back in and re-run.
🌐
Will it crawl pages behind a login? Browser extension uses your live browser session — log in first, then start the crawl/explore. MCP server uses a Playwright storage_state_path JSON; capture it once via capture_auth or import_chrome_cookies, then point every run at it.
🔧
Why does Chrome show "Pathfinder started debugging this browser"? During Explore, Pathfinder attaches the Chrome DevTools Protocol (CDP) to capture HAR entries, accessibility-tree audits, and live screenshots. The debugger detaches the moment exploration finishes. No page data leaves your device.
📤
How do I share work with my team? Use the export/import buttons in Knowledge, Explore (graph), and Flows. Each exports a self-contained JSON your teammate can import into a fresh install. The MCP equivalents are export_knowledge / export_explore.
🧹
How do I reset everything? Browser: open the Explore tab → Delete (clears graph + flows only; KB and tests are kept). For a full wipe, uninstall the extension and reinstall — storage keys live under pathfinder_* in chrome.storage.local and IndexedDB pathfinder_db. MCP: clear_knowledge + clear_explore.