# AginxBrowser API Reference [English](API.md) | [中文](API.zh-CN.md) > Complete HTTP API reference. One binary, local install, every capability a POST away. ## Quick Start ```bash # Build and start cargo build --release ./target/release/aginxbrowser # Verify the service curl http://127.0.0.1:8089/health # → {"status":"ok","engine":"diting","version":"0.3.1","commit":"a1b2c3d",...} # Fetch a page curl -sS -X POST http://127.0.0.1:8089/fetch \ -H "Content-Type: application/json" \ -d '{"url":"https://example.com"}' # Create an interactive session curl -sS -X POST http://127.0.0.1:8089/session/create \ -H "Content-Type: application/json" \ -d '{"url":"https://example.com"}' ``` --- ## HTTP API Listens on `0.0.0.0:8089` by default; override via the `AGINXBROWSER_BIND` environment variable. ### GET /health Health check. Also the build-identity call: `version` + `commit` answer "which source is this binary" (compare against the release tag to verify doc/tag/binary/source are the same commit), `v8` is the JS kernel version actually executing scripts (the engine truth behind the UA string — "15.0.274.2" doesn't depend on which persona a session carried), and `ua`/`tls` say what the instance presents to sites — the UA browser traffic carries (`AGINXBROWSER_UA` override, else the pinned persona; imported sessions keep the copied request's own UA by design) and the default TLS fingerprint (`"off"` in non-stealth builds). `commit` is `"unknown"` for git-less builds. ```bash curl http://127.0.0.1:8089/health ``` Response: ```json { "status": "ok", "engine": "diting", "version": "0.3.1", "commit": "a1b2c3d", "v8": "15.0.274.2", "ua": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/145.0.0.0 Safari/537.36", "tls": "chrome145", "capabilities": { "screenshot": true, "stealth": true, "captcha_solver": false, "ffmpeg": { "version": "8.1.1" } } } ``` --- ### POST /fetch Fetch a page and return its content. Supports tiered rendering, automatic Cloudflare bypass, TLS fingerprint switching, and JS data extraction. **Request fields:** | Field | Type | Required | Default | Description | |------|------|------|------|------| | url | string | ✅ | — | Target URL | | format | string | | `"markdown"` | Output format: `markdown` / `html` / `text` | | selector | string | | `null` | CSS selector; only extract the matching region | | wait_secs | u64 | | `null` | Extra seconds to wait after page load (let JS rendering finish) | | use_proxy | bool | | `false` | Route through the `AGINXBROWSER_PROXY` proxy. Set `true` for overseas sites | | cookies | string[] \| object[] | | `[]` | Cookies injected before navigation: `"name=value"` strings (may carry `; Domain=…; Path=/; Secure` attributes) or cookie objects `{"name","value","domain","path","secure","httpOnly","sameSite"}`. Entries declaring `Domain=` anchor at that domain, so sibling-domain login state (`.taobao.com` / `.tmall.com` style) survives injection instead of being dropped by RFC 6265 domain checks | | max_chars | usize | | `50000` | Truncate `content` to this many characters. `0` = unlimited | | auto_bypass_challenge | bool | | `true` | Automatically detect and bypass Cloudflare Turnstile challenges | | render_tier | string | | `"auto"` | Rendering strategy (see below) | | tls_fingerprint | string | | `null` | TLS fingerprint (stealth mode), see below | | js_extract | object | | `null` | JS data extraction (see below) | | sanitize | bool | | `true` | Strip prompt-injection carriers from the text/markdown output (see below) | | capture_xhr | string[] | | `null` | Return the page's script-initiated XHR/fetch response bodies as a first-class `xhr` field. Entries are URL substrings; `[]` = every XHR/fetch. Forces the browser tier | **`render_tier` options:** | Value | Description | |----|------| | `auto` | Direct HTTP fetch first; automatically fall back to the browser when content is insufficient (**recommended**, default) | | `http` | Pure HTTP, no browser. Fastest but cannot capture JS-rendered content | | `browser` | Force the full JS browser. Slowest but most reliable. `obscura` is still accepted and is not the name to send | > **JS-heavy SPA sites** (cls.cn, juejin-class feeds, WeChat articles) render their content client-side: with `auto` the HTTP pass returns a small skeleton and the sufficiency gate usually catches it, but a site that returns a *plausible-looking* stub defeats the heuristic. When you know the target is an SPA, pass `"render_tier":"browser"` explicitly — you skip a wasted HTTP round-trip and get the rendered page directly. The `tier` field in the response tells you which path served a given fetch. > > **WeChat article text** (`mp.weixin.qq.com`, no explicit `selector`): the article body is server-rendered inside `#js_content` but hidden until WeChat's own JS reveals it, so plain `body` text is just the title/byline shell. `format:"text"`/`"markdown"` reads the `#js_content` container and returns the fuller of the two extractions — the full article comes back even when the reveal script never finishes. **`tls_fingerprint` options (requires `--features stealth`):** | Value | Description | |----|------| | `null` | Default Chrome145 | | `"chrome145"` | Chrome 145 | | `"firefox133"` | Firefox 133 | | `"firefox147"` | Firefox 147 | | `"safari17_5"` | Safari 17.5 | | `"safari18"` | Safari 18 | | `"safari26"` | Safari 26 | | `"edge145"` | Edge 145 | **`js_extract` format:** ```json { "expression": "JSON.stringify(window.__INITIAL_STATE__)", "timeout_ms": 5000 } ``` | Field | Type | Default | Description | |------|------|------|------| | expression | string | — | JS expression evaluated in the page context | | timeout_ms | u64 | `5000` | Timeout waiting for a non-null result (milliseconds) | **Response fields:** | Field | Type | Description | |------|------|------| | url | string | Final URL (after redirects) | | title | string? | Page title | | content | string | Fetched content (markdown/html/text) | | truncated | bool | Whether `content` was truncated by `max_chars` | | tier | string? | Which path served the page: `"http"` (plain HTTP + conversion, ~100ms) or `"browser"` (V8 render) — present under `render_tier: "auto"` too, so callers can see why a fetch was fast or slow | | redirected_from | string[]? | The redirect trail: `redirected_from[0]` is the URL you asked for, `url` is where the content actually came from (absent when no redirect happened) | | js_extract_result | any? | JS extraction result (only present when `js_extract` is set) | | sanitize_report | object? | What the injection stripper removed (only present when `sanitize` fired — see below) | | xhr | object[] | Script-initiated response bodies (only present when `capture_xhr` is set — see below) | | captcha_event | object? | CAPTCHA event (only present when a CAPTCHA is detected; covers Cloudflare/Google/Baidu challenge pages plus Taobao/Tmall risk-control signals — `punish` redirects, `x5sec`, and MTop `FAIL_SYS_USER_VALIDATE`/`RGV587` replies even when they arrive as HTTP 200) | **`sanitize` — injection stripping (default on):** Page text is untrusted input, and a reading tool owes its caller content that doesn't carry instructions aimed at the reader. On `text`/`markdown` output (raw `html` is never touched) three carriers are stripped: - **Zero-width/steganographic characters** (`​` family) — never legitimate page prose. - **Hidden-span text** — the live DOM is probed for elements a human can't see (`opacity:0`, sub-4px font) whose text `innerText` happily carries. Whole-hidden containers (SSR-pending reveals, WeChat `#js_content`) are guarded: when hidden text is more than half the extraction, nothing is removed. - **Instruction-shaped lines** — lines matching curated injection phrasings (EN/CN: "ignore previous instructions" class, chat markup like `<|im_start|>`) are dropped whole, because the payload continues past the matched phrase. Removal is observable, never silent: when anything fires, `sanitize_report` says what — `{"zero_width_removed": 1, "hidden_spans_removed": 1, "patterns_hit": {"ignore_previous_instructions": 1}}`. It's a heuristic, not a firewall; to study the payload itself, pass `"sanitize": false`. The `selector` parameter is the CSS-narrowing half of the story: restrict the extraction to the content region and the chrome's injected noise never enters the text at all. **`capture_xhr` — the page's own API face:** Usually the cleanest read of a JS-heavy page isn't the rendered DOM but the JSON APIs the page itself calls. `capture_xhr` returns those response bodies alongside the text: ```bash curl -sS -X POST http://127.0.0.1:8089/fetch \ -H "Content-Type: application/json" \ -d '{"url":"https://spa.example.com/list","capture_xhr":["/api/"],"wait_secs":2}' ``` Each row is `{"url", "method", "status", "mime", "body", "body_truncated", "request_headers"}`. `request_headers` is the request's outbound header set (lowercased) — the signed customs the page's JS put on the call. Bodies are capped at `min(max_chars, 8000)` chars per entry (at most 20 entries) so a page's API traffic can't flood the context; binary (base64) bodies are skipped. `wait_secs` matters here: the page's fetches need a moment to land in the network log. **`captcha_event` format:** | Field | Type | Description | |------|------|------| | engine | string | Name of the search engine that triggered the CAPTCHA (empty for `/fetch`) | | captcha_type | string | `cloudflare_turnstile` / `recaptcha_v2` / `hcaptcha` / `slider` / `unknown` | | url | string | URL that triggered the CAPTCHA | | detected_at | u64 | Wall-clock time of detection (unix seconds) | | hit_count | u32 | Consecutive CAPTCHA hits for this engine — the backoff-ladder step driving the suspension duration (`1` → 5 min, `2` → 10 min, `3` → 30 min, `4+` → 1 h). Always `1` on the `/fetch` path | | auto_solve_attempted | bool | Whether auto-solve was attempted | | auto_solve_succeeded | bool | Whether auto-solve succeeded | **Example — basic fetch:** ```bash curl -sS -X POST http://127.0.0.1:8089/fetch \ -H "Content-Type: application/json" \ -d '{"url":"https://example.com"}' ``` ```json { "url": "https://example.com/", "title": "Example Domain", "content": "# Example Domain\n\nThis domain is for use in illustrative examples...", "truncated": false, "tier": "http" } ``` The `tier` field reports which strategy served the request: `"http"` (plain HTTP + conversion) or `"browser"` (V8 render) — including under `render_tier: "auto"`, so callers can see why a fetch was fast or slow. **Example — extract structured data from an SPA:** ```bash curl -sS -X POST http://127.0.0.1:8089/fetch \ -H "Content-Type: application/json" \ -d '{ "url": "https://spa-site.example.com", "js_extract": { "expression": "JSON.stringify(window.__INITIAL_STATE__)", "timeout_ms": 3000 } }' ``` **Example — extract a specific region (CSS selector):** ```bash curl -sS -X POST http://127.0.0.1:8089/fetch \ -H "Content-Type: application/json" \ -d '{"url":"https://github.com/trending","format":"text","selector":"article","use_proxy":true}' ``` **Caching**: `/fetch` has an in-process cache (key includes url/format/selector/cookies/use_proxy/max_chars/render_tier/tls_fingerprint); TTL is controlled by `AGINXBROWSER_CACHE_TTL_SECS` (default 600s, `0` disables). Repeat fetches of the same URL hit the cache (~0.01s vs ~1s on first fetch). **Security**: built-in SSRF protection (blocks non-http(s) schemes and private-network/loopback IPs), DNS rebinding protection, and tracker blocking (stealth mode). An RFC 9309 robots.txt checker ships built in but is off by default (`AGINXBROWSER_HONOR_ROBOTS=1` to opt in). --- ### POST /click Load a page and click the specified element (`element.click()`), returning the page text after the click. **Request fields:** | Field | Type | Required | Default | Description | |------|------|------|------|------| | url | string | ✅ | — | Target URL | | selector | string | ✅ | — | CSS selector | | wait_secs | u64 | | `null` | Extra seconds to wait after page load | | use_proxy | bool | | `false` | Route through a proxy | | cookies | string[] \| object[] | | `[]` | Cookies injected before navigation (`"name=value"` strings or cookie objects, same semantics as `/fetch`) | | tls_fingerprint | string | | `null` | TLS fingerprint (stealth mode) | **Response fields:** | Field | Type | Description | |------|------|------| | url | string | Final URL | | selector | string | The selector used | | clicked | bool | Whether the click succeeded | | text_after | string? | Page text after the click | **Example:** ```bash curl -sS -X POST http://127.0.0.1:8089/click \ -H "Content-Type: application/json" \ -d '{"url":"https://example.com","selector":"a"}' ``` --- ### POST /eval Execute arbitrary JavaScript on the page and return the result. Supports `async`/`Promise`. **Request fields:** | Field | Type | Required | Default | Description | |------|------|------|------|------| | url | string | ✅ | — | Target URL | | script | string | ✅ | — | JS expression or async IIFE | | wait_secs | u64 | | `null` | Extra seconds to wait after page load | | use_proxy | bool | | `false` | Route through a proxy | | cookies | string[] \| object[] | | `[]` | Cookies injected before navigation (`"name=value"` strings or cookie objects, same semantics as `/fetch`) | | tls_fingerprint | string | | `null` | TLS fingerprint (stealth mode) | **Response fields:** | Field | Type | Description | |------|------|------| | url | string | Final URL | | result | any | JS execution result | > The `script` parameter of `/eval` supports **async functions**: returned Promises are awaited automatically. Ideal for dynamically rendered React/Vue-style pages — wait for rendering to finish, then extract data. **Example — async script (wait for dynamic rendering):** ```bash curl -sS -X POST http://127.0.0.1:8089/eval \ -H "Content-Type: application/json" \ -d '{ "url":"https://github.com/trending", "script":"(async()=>{await new Promise(r=>setTimeout(r,4000));return Array.from(document.querySelectorAll(\"article.Box-row\")).slice(0,5).map(a=>a.querySelector(\"h2 a\")?.textContent?.trim())})()", "use_proxy":true }' ``` --- ### POST /search Native aggregated search with optional automatic content fetching. Agents go from "search" to "read" in one step. **Request fields:** | Field | Type | Required | Default | Description | |------|------|------|------|------| | q | string | ✅ | — | Search query | | fetch_top | usize | | `0` | Fetch full page content for the top N results. `0` = snippets only | | categories | string | | `"general"` | Search category, comma-separated: `general` / `images` / `news`. `images` returns direct image links | | language | string | | `"zh-CN"` | Language | | max_results | usize | | `10` | Maximum number of results | | max_chars_per | usize | | `4000` | Per-result content truncation in characters. `0` = unlimited | | wait_secs | u64 | | `3` | Seconds to wait for JS rendering per page while fetching content | | use_proxy | bool | | `false` | Whether to route content fetching through a proxy (overseas sites) | | engines | string[] | | `[]` | Restrict to these engine names (e.g. `["baidu"]`). Empty = all engines serving `categories`. Unknown names → 400 with the valid list | | time_range | string | | — | Freshness window: `day` / `week` / `month` / `year`. Honored by engines with dated results (bing_news passes a server-side freshness window); others ignore it. Invalid values → 400 | > Unknown request fields are rejected with a 400 — a typo'd parameter name (`count`, `limit`, `num`) fails loudly instead of being silently ignored. **Built-in search engines** (`GET /engines` returns this list with live health): | Engine | Categories | Description | |------|------|------| | baidu | general | Baidu HTML SERP (browser UA; the old `tn=json` API is wall-dead) | | bing | general | Bing HTML parsing | | sogou | general | Sogou web search | | sogou_wechat | general, news | Sogou WeChat article search | | duckduckgo | general | DuckDuckGo HTML (stealth fingerprint) | | wikipedia | general | Wikipedia MediaWiki search API (language edition follows `language`; CN-blocked → direct-first/proxy-retry) | | bing_news | general, news | Bing News infinite-scroll fragment; `time_range` via server-side freshness window | | baidu_images | images | Baidu Images `acjson` JSON | | bing_images | images | Bing Images `images/async` | | arxiv | academic | arXiv API (full-text matching — academic queries only, not the general pool) | | openalex | academic | OpenAlex works API — cross-publisher literature (DOI links, inverted-index abstracts); occasional anonymous 429s degrade to a skipped engine, arXiv still serves the category | | huggingface | general, ai | Hugging Face models search | | github | general, code | GitHub repository search | | stackexchange | general, code | StackExchange API | | mdn | code | MDN Web Docs v1 search API (code queries only — CJK fuzzy-match noise keeps it out of the general pool) | | npm | general, packages | npm registry | | pypi | general, packages | PyPI | | rubygems | packages | RubyGems registry (no-offset API — paged client-side; CN-blocked → direct-first/proxy-retry) | | hn | general | Hacker News via the Algolia API (story links with points/comment counts; `time_range` maps to a created_at filter; Ask-HN threads fall back to the HN item URL) | `meilisearch` additionally registers when `AGINXBROWSER_MEILI_URL` + `AGINXBROWSER_MEILI_INDEX` are set (private index). Google search is not included — it requires a proxy and steady CAPTCHA clearance from mainland China; use `duckduckgo`/`bing` instead. Engines are queried concurrently and results merged with deduplication: identical URLs (after normalization) merge into a single entry, `engines` lists the source engines, and `score` accumulates. `GET /engines` returns every engine's name, categories, and live suspension state — call it to discover valid names for the `engines` filter. **CAPTCHA progressive backoff**: when an engine triggers a CAPTCHA it pauses automatically, with the pause duration escalating on consecutive hits (5 min → 10 min → 30 min → 1 h) and resetting after a successful search. Set the `CAPTCHA_SOLVER_API_KEY` environment variable to enable automatic CAPTCHA solving. **Response fields:** | Field | Type | Description | |------|------|------| | query | string | Search query | | number_of_results | usize | Number of results returned (equals `results.length`) | | results | array | Result list | | captcha_events | array | List of CAPTCHA events (see `captcha_event` format above — each carries `detected_at` + `hit_count`) | | engine_errors | object | Per-engine reason an engine contributed nothing: CAPTCHA suspension (with resume countdown), transient fetch/parse failure, or a task panic. Absent when every eligible engine answered | **Each entry in `results`:** | Field | Type | Description | |------|------|------| | title | string | Title | | url | string | Link | | snippet | string | Search snippet | | engines | string[] | Source engines | | score | float | Combined score | | content | string? | Page content (only present within the `fetch_top` range) | | content_truncated | bool | Whether the content was truncated | | fetch_error | string? | Reason content fetching failed | | image_url | string? | Direct link to the image binary (downloadable straight to jpg/png with `curl -o`). Only with `categories=images` | | source_url | string? | URL of the page hosting the image (provenance/copyright) | | width | u32? | Image width (px) | | height | u32? | Image height (px) | > With `categories=images`, the `url` field equals `image_url` (the direct image link, convenient for immediate download); `snippet` is empty. Baidu Images prefers `objURL` (original image, highest quality) and falls back to a CDN-proxied direct link when unavailable. **Example — search + fetch top 3 pages:** ```bash curl -sS -X POST http://127.0.0.1:8089/search \ -H "Content-Type: application/json" \ -d '{"q":"macbook 价格","fetch_top":3,"max_chars_per":2000}' ``` **Example — restrict to WeChat articles, today's news only:** ```bash curl -sS -X POST http://127.0.0.1:8089/search \ -H "Content-Type: application/json" \ -d '{"q":"A股 午评","engines":["sogou_wechat","bing_news"],"categories":"news","time_range":"day","fetch_top":2}' ``` If `engines` names an engine the server doesn't know, the response is a 400 carrying the valid list: ```json {"error":"unknown engine \"wechat\"; valid engines: baidu, baidu_images, bing, bing_images, sogou, sogou_wechat, ... (also check GET /engines)"} ``` **Example — image search (returns direct links, downloadable straight from curl):** ```bash curl -sS -X POST http://127.0.0.1:8089/search \ -H "Content-Type: application/json" \ -d '{"q":"蔚来ES8 酒红内饰 后排视角","categories":"images","max_results":10}' ``` ```json { "query": "蔚来ES8 酒红内饰 后排视角", "number_of_results": 20, "results": [ { "title": "蔚来ES8 酒红内饰后排实拍", "url": "https://n.sinaimg.cn/.../img.jpg", "engines": ["baidu_images"], "score": 20.0, "image_url": "https://n.sinaimg.cn/.../img.jpg", "source_url": "https://auto.sina.com.cn/...", "width": 1920, "height": 1080 } ] } # Download the image curl -sL -o cabin_ref.jpg "" ``` --- ### GET /engines Discover the search-engine vocabulary plus live health. Returns every registered engine with the categories it serves and its current suspension state — call this before using /search's `engines` filter, or to see who is currently benched by a CAPTCHA. ```bash curl -sS http://127.0.0.1:8089/engines ``` ```json { "engines": [ { "name": "baidu", "categories": ["general"], "suspended": false, "captcha_count": 0 }, { "name": "sogou_wechat", "categories": ["general", "news"], "suspended": true, "suspend_remaining_secs": 184, "captcha_count": 1 } ] } ``` `suspended: true` means the engine hit a CAPTCHA and is in progressive backoff — `/search` skips it (the response's `engine_errors` says so, with the resume countdown) until the suspension expires. --- ### POST /download Stream a file from a URL to disk. Unlike `/fetch` (which returns page content for reading), `/download` saves the raw bytes — use it for binaries, archives, datasets, documents. The body streams chunk-by-chunk to disk (never buffered in memory), with SHA-256 computed incrementally so integrity is verifiable in one call. **Request fields:** | Field | Type | Required | Default | Description | |------|------|------|------|------| | url | string | ✅ | — | File URL (`http`/`https` only) | | filename | string | | auto | Output filename. Auto resolution: `Content-Disposition` header → URL path tail → `"download"` | | resume | bool | | `false` | Continue an interrupted download when a local partial file exists. Server support is probed via `Range: bytes=N-`: `206` appends, `200` restarts | | use_proxy | bool | | `false` | Route through proxy (auto-enabled for known blocked domains like github.com) | | cookies | string[] \| object[] | | `[]` | Cookies to send (`["name=value", ...]` or cookie objects) for gated downloads | **Response fields:** | Field | Type | Description | |------|------|------| | url | string | Final URL after redirects | | path | string | Absolute path of the completed file on disk | | filename | string | Resolved filename | | size_bytes | u64 | Bytes written by this call (append counts only appended portion) | | content_type | string? | Response Content-Type | | sha256 | string | SHA-256 over the complete file content | | resumed | bool | Whether an existing partial file was continued via Range/206 | **Behavior notes:** - Files land in `AGINXBROWSER_DOWNLOAD_DIR` (default: current working directory). In-flight data is written to `.part`, then renamed on success. - Same SSRF policy as `/fetch`: loopback / private / link-local targets are rejected unless `AGINXBROWSER_ALLOW_PRIVATE_NETWORK=1` (all-or-nothing) or `AGINXBROWSER_ALLOW_NETWORK=` (scoped allowlist — listed ranges open, cloud-metadata endpoints stay blocked). - Redirects are followed (up to 20 hops), each hop re-validated against SSRF. - A 30s stall timeout aborts if no bytes arrive (dead connection instead of hang). Hard cap: 4 GB per call. - Filenames are sanitized (path traversal stripped, length capped). **Example — download and verify:** ```bash curl -sS -X POST http://127.0.0.1:8089/download \ -H "Content-Type: application/json" \ -d '{"url":"https://github.com/obsidianmd/obsidian-releases/releases/download/v1.5.3/Obsidian-1.5.3-macOS.dmg","resume":true}' ``` --- ### POST /screenshot Render the page's post-JS DOM into a PNG screenshot (returned as base64). **Requires building with `--features screenshot`** (not included by default; see the build section). Does not use `/fetch`'s tiered rendering — it always drives the JS browser through full execution, then renders the result with the built-in diting engine (CSS cascade + Taffy box layout + CPU paint, no Chromium). Pass `"engine": "blitz"` to opt into the Blitz reference pipeline for comparison renders — that requires building with `--features blitz-reference` (blitz is not compiled in by default). **Request fields:** | Field | Type | Required | Default | Description | |------|------|------|------|------| | url | string | ✅ | — | Target URL | | width | u32 | | `1280` | Viewport width (CSS px) | | height | u32 | | `800` | Viewport height (CSS px; serves as a lower bound when `full_page` is set) | | scale | f32 | | `1.0` | Device pixel ratio; higher is sharper but yields larger PNGs | | full_page | bool | | `true` | Capture the entire scrollable page (tracks content height, capped at 16000px) | | wait_secs | u64 | | `null` | Extra seconds to wait after load (for JS rendering) | | selector | string | | `null` | CSS selector; capture the **specified element region** instead of the full page (see below) | | selector_all | bool | | `false` | Used with `selector`: skip cropping and return coordinates of **all matches** | | use_proxy | bool | | `false` | Route through the `AGINXBROWSER_PROXY` proxy | | cookies | string[] \| object[] | | `[]` | Cookies injected before navigation (`"name=value"` strings or cookie objects, same semantics as `/fetch`) | | tls_fingerprint | string | | `null` | TLS fingerprint (stealth mode) | **Response fields:** | Field | Type | Description | |------|------|------| | url | string | Final URL (after redirects) | | title | string? | Page title | | width | u32 | Actual PNG pixel width rendered (differs from the requested value when `full_page` tracks content height or `selector` crops) | | height | u32 | Actual PNG pixel height rendered | | image_base64 | string | base64-encoded PNG. Decode with `base64 -d`, or use directly as `data:image/png;base64,...` | | format | string | Always `"png"` | | selector_rects | object[]? | Present only when the request includes `selector`. One `{x, y, width, height}` per element, in **CSS px with the page's top-left corner as origin** (not viewport coordinates) | **Selector mode (element-level screenshot + coordinates):** - `selector` + `selector_all=false` (default): the image is cropped to the border box of the first matching element; `selector_rects` contains exactly one entry (the cropped region). - `selector` + `selector_all=true`: the image renders as a normal full page; `selector_rects` returns coordinates for **every match** — the agent can consume just the coordinates without the image. - Coordinates come from diting's post-layout Taffy border boxes, accumulated along the layout tree into absolute page coordinates. > ⚠️ **Inline element limitation**: inline elements containing only text (e.g. `文字`) have no standalone Taffy box — crop mode errors out advising you to pick a block-level ancestor, and `selector_all` mode returns `0x0`. Inline elements containing block-level or replaced content (`` etc.) fall back to the union of their descendants' boxes. Selectors targeting **block-level containers** (div/section/li, etc.) yield reliable coordinates. **Example — screenshot a Baidu search:** ```bash curl -sS -X POST http://127.0.0.1:8089/screenshot \ -H "Content-Type: application/json" \ -d '{"url":"https://www.baidu.com/s?wd=蔚来ES8","full_page":true,"wait_secs":2}' \ | jq -r .image_base64 | base64 -d > baidu.png ``` **Example — crop the first search result + get coordinates of all results:** ```bash # Crop just the first .result under #content_left curl -sS -X POST http://127.0.0.1:8089/screenshot \ -H "Content-Type: application/json" \ -d '{"url":"https://www.baidu.com/s?wd=蔚来ES8","selector":"#content_left .result"}' \ | jq -r .image_base64 | base64 -d > first-result.png # Skip the image — just want page coordinates for all 9 results curl -sS -X POST http://127.0.0.1:8089/screenshot \ -H "Content-Type: application/json" \ -d '{"url":"https://www.baidu.com/s?wd=蔚来ES8","selector":"#content_left .result","selector_all":true,"full_page":false,"width":100,"height":100}' \ | jq -c '.selector_rects' # [{"x":150,"y":2843,"width":608,"height":153}, {"x":150,"y":3016,"width":608,"height":69}, ...] ``` ```json { "url": "https://www.baidu.com/s?wd=蔚来ES8", "title": "蔚来ES8_百度搜索", "width": 1280, "height": 800, "image_base64": "iVBORw0KGgo...", "format": "png" } ``` > Screenshots are the agent's "visual input" — CSS rendering on complex sites is approximate (not Chromium pixel-perfect). Sub-resources such as images are not fetched separately (`` may be missing from screenshots); text and layout are reliable. --- ### POST /video Render a page's animation timelines to an MP4 video (returned as base64). **Requires building with `--features screenshot` and ffmpeg on the server's PATH.** The page's scripts must expose their timelines in `window.__timelines` — objects with `duration()` and `pause(t)` (a paused GSAP timeline registered there works as-is): ```js const tl = gsap.timeline({ paused: true }); tl.from("#box", { opacity: 0, x: -200, duration: 2, ease: "power2.out" }); window.__timelines = { main: tl }; ``` Each frame seeks every registered timeline to `t = i/fps` and paints the viewport — the frame values carry no wall clock, so output is deterministic across runs. The full path runs in-process: seek → viewport band paint → raw RGBA piped into ffmpeg → H.264/yuv420p MP4. **Request fields:** | Field | Type | Required | Default | Description | |------|------|------|------|------| | url | string | ✅ | — | Target URL (must populate `window.__timelines`) | | fps | f64 | | `24` | Frames per second | | width | u32 | | `1280` | Viewport width (CSS px; floored to even — yuv420p) | | height | u32 | | `720` | Viewport height (CSS px) | | hold_tail_secs | f64 | | `0.5` | Freeze the final timeline state for this many extra seconds | | max_duration_secs | f64 | | `120` | Safety cap on timeline + hold tail (longer timelines error instead of encoding) | | wait_timelines_ms | u64 | | `10000` | How long to wait for `window.__timelines` to appear | | use_proxy | bool | | `false` | Route through the `AGINXBROWSER_PROXY` proxy | | cookies | string[] \| object[] | | `[]` | Cookies injected before navigation (same semantics as `/fetch`) | | tls_fingerprint | string | | `null` | TLS fingerprint (stealth mode) | | narration | object[] | | `[]` | Voiceover clips: `{url, start_secs, volume}` — each fetched, delayed to its start time, all mixed into one AAC track. Any TTS output works (mp3/wav/ogg/m4a, probed by content). Fetch failure is an error, never a silent video | | audio | object | | `null` | Background music: `{url, volume, fade_out_secs, loop_audio}` — looped to cover the video, volume-scaled, faded out at the tail | | subtitles_srt | string | | `null` | Inline SRT text muxed as a soft (toggleable) mov_text track — no libass needed for muxing. Cap 64 KiB | | subtitles_language | string | | `null` | ISO language tag for the subtitle track ("eng", "zh") | **Response fields:** | Field | Type | Description | |------|------|------| | url | string | Final URL | | title | string? | Page title | | frames | u32 | Frames written to the encoder | | timeline_secs | f64 | Longest registered timeline, seconds | | duration_secs | f64 | Total video length = timeline + hold tail | | width / height | u32 | Encoded pixel size | | video_base64 | string | base64-encoded MP4 (H.264, yuv420p). Decode with `base64 -d`, or use directly as `data:video/mp4;base64,...` | | has_audio | bool | Whether an audio track (BGM and/or narration) was muxed in | | has_subtitles | bool | Whether a soft subtitle track was muxed in | | format | string | Always `"mp4"` | Audio details: audio URLs are fetched through the page's own HTTP client (http/https only, ≤16 MiB, 8 s timeout, same SSRF gate as subresources). One audio source rides a plain `-af volume/adelay/afade` chain; two or more (e.g. BGM + narration clips) are mixed with `amix=normalize=0` so narration isn't halved for having quiet music under it. With ≥2 audio inputs the subtitle stream is mapped explicitly. Burned-in captions need no engine support — register a second `__timelines` entry that writes caption text/opacity in `pause(t)`, and it seeks along with everything else. **Example:** ```bash curl -sS -X POST http://127.0.0.1:8089/video \ -H "Content-Type: application/json" \ -d '{"url":"https://example.com/anim.html","fps":20,"width":800,"height":450}' \ | jq -r .video_base64 | base64 -d > anim.mp4 ``` Errors carry the failure reason verbatim: no `__timelines` before `wait_timelines_ms`, zero timeline duration, the duration cap, ffmpeg missing from PATH, or ffmpeg exiting non-zero (with its stderr tail). --- ### POST /pdf Cut a rendered page into a set of pages and package as PDF (default), per-page PNGs, or an image-based PPTX/DOCX. **Requires building with `--features screenshot`.** Two slicing modes, picked by whether `selector` is set: - **Print** (no `selector`): fixed-height pages over the whole document (default 794×1123, A4 @96dpi), breaking at top-level block boundaries where possible — the break lands on the deepest block bottom that fits, with a half-page floor so pages never collapse to slivers. A remainder shorter than 64px merges into the previous page instead of a near-blank tail. - **Slides** (`selector` set): one page per CSS-selector match, sized to that element. Build the deck as plain HTML with one `.slide` div per slide; each match becomes a page at its own height. Each page is painted as a viewport band off the live tree's layout — same primitive the timeline pump rides, no Chromium. The PDF is image-based: per-page JPEG (`jpeg_quality`) embedded via DCTDecode, one page object per page with its own MediaBox in points (px→pt at 96dpi), so variable-height pages need no normalization. PPTX and DOCX are the same pages re-packaged: PPTX as one slide per page (deck-sized to the largest page, images anchored top-left), DOCX as one page-sized section per page with zero margins — Word sizes every section independently, so each page keeps its exact height. Both containers are written by hand (stored-ZIP, fixed timestamps — byte-deterministic), zero new dependencies. `format="pptx-native"` is the editable tier: instead of rasterizing pages, each `selector` match is walked element-by-element off the live tree (gBCR + computed style from the same layout cache the band paints ride) and mapped to native DrawingML — text becomes real text runs (`` with font family/size/weight/color/alignment), background boxes become shapes (solid fill, `border-radius` → roundRect, CSS gradients → `gradFill` with the angle converted), `` becomes a `p:pic` with the fetched bytes as a media part. Slides mode only: `selector` is required, one slide per match, and the deck is sized to the largest slide. What a text box can't express (per-glyph inline styling, z-index reordering, transforms, borders) degrades by omission — the element still lands as an editable shape. Note: the engine doesn't expand the CSS `background` shorthand into `background-image` yet, so gradient slides must use the `background-image` longhand. **Request fields:** | Field | Type | Required | Default | Description | |------|------|------|------|------| | url | string | ✅ | — | Target URL | | format | string | | `"pdf"` | `"pdf"` (base64 PDF), `"png"` (one base64 PNG per page), `"pptx"` (one slide per page), `"pptx-native"` (editable: element-level DrawingML; requires `selector`), or `"docx"` (one page-sized section per page) | | width | u32 | | `794` | Page width (CSS px) | | height | u32 | | `1123` | Page height (CSS px) — print pagination only; slides size each page to its element | | selector | string | | `null` | CSS selector; set → slides mode, unset → print mode | | max_pages | usize | | `50` | Safety cap on emitted pages (more pages errors instead of rendering) | | jpeg_quality | u8 | | `90` | JPEG quality 1-100 for page embedding in PDF/PPTX/DOCX (PNG format ignores it) | | use_proxy | bool | | `false` | Route through the `AGINXBROWSER_PROXY` proxy | | cookies | string[] \| object[] | | `[]` | Cookies injected before navigation (same semantics as `/fetch`) | | tls_fingerprint | string | | `null` | TLS fingerprint (stealth mode) | **Response fields:** | Field | Type | Description | |------|------|------| | url | string | Final URL | | title | string? | Page title | | pages | usize | Page count | | width / height | u32 | Requested page size (slides pages vary in height — each PNG self-describes) | | pdf_base64 | string? | base64 PDF (set when `format="pdf"`). Decode with `base64 -d`, or use directly as `data:application/pdf;base64,...` | | pages_base64 | string[] | One base64 PNG per page (non-empty when `format="png"`) | | pptx_base64 | string? | base64 PPTX, one slide per page (set when `format="pptx"` or `"pptx-native"`) | | docx_base64 | string? | base64 DOCX, one page-sized section per page (set when `format="docx"`) | | format | string | `"pdf"` / `"png"` / `"pptx"` / `"pptx-native"` / `"docx"` | **Examples:** ```bash # Print mode: paginate an article into A4 pages curl -sS -X POST http://127.0.0.1:8089/pdf \ -H "Content-Type: application/json" \ -d '{"url":"https://example.com/long-article.html"}' \ | jq -r .pdf_base64 | base64 -d > article.pdf # Slides mode: one .slide per page, as PNGs curl -sS -X POST http://127.0.0.1:8089/pdf \ -H "Content-Type: application/json" \ -d '{"url":"https://example.com/deck.html","selector":".slide","format":"png"}' \ | jq -r '.pages_base64[0]' | base64 -d > slide-0.png # Slides mode → PowerPoint deck (HTML .slide divs become real slides) curl -sS -X POST http://127.0.0.1:8089/pdf \ -H "Content-Type: application/json" \ -d '{"url":"https://example.com/deck.html","selector":".slide","format":"pptx"}' \ | jq -r .pptx_base64 | base64 -d > deck.pptx # Slides mode → editable PowerPoint (text runs, shapes, gradients, images) curl -sS -X POST http://127.0.0.1:8089/pdf \ -H "Content-Type: application/json" \ -d '{"url":"https://example.com/deck.html","selector":".slide","format":"pptx-native"}' \ | jq -r .pptx_base64 | base64 -d > deck.pptx ``` Errors carry the reason verbatim: a selector that matches nothing, more pages than `max_pages`, or a document with no content height. --- ### POST /v1/scrape (Firecrawl-compatible) [Firecrawl](https://github.com/mendableai/firecrawl)-compatible endpoint. Existing Firecrawl clients can migrate by simply changing the base URL. **Request fields:** | Field | Type | Required | Default | Description | |------|------|------|------|------| | url | string | ✅ | — | Target URL | | formats | string[] | | `["markdown"]` | Output formats: `["markdown"]` / `["html"]` / `["markdown","html"]` | | onlyMainContent | bool | | `false` | Main content only (parameter accepted, not yet implemented) | | waitFor | u64 | | `null` | Milliseconds to wait for JS rendering | | timeout | u32 | | `null` | Timeout in milliseconds (parameter accepted) | | actions | object[] | | `[]` | Actions to perform before scraping (see below) | | selector | string | | `null` | CSS selector | | tls_fingerprint | string | | `null` | TLS fingerprint (stealth mode) | **`actions` format:** ```json [ {"type": "click", "selector": "button.accept"}, {"type": "wait", "milliseconds": 1000} ] ``` | type | Fields | Description | |------|------|------| | `click` | `selector` | Click an element (anchor links navigate to the target page) | | `wait` | `milliseconds` | Wait the given number of milliseconds | | `screenshot` | — | Screenshot the rendered page, returned as a base64 data-URI (requires the `screenshot` feature) | | `scroll` | — | Scroll the page | | `writeText` | `text`, `selector` | Type text into the matching element | | `pressKey` | `key` | Press a key (Enter submits the enclosing GET form) | When any `actions` are present, `/v1/scrape` follows a **single-page session flow**: navigate once → execute actions in order → extract from the final state of that page. All actions operate on the same page. When the request includes a `screenshot` action (or `formats` contains `"screenshot"`), the response's `data.screenshot` carries a `data:image/png;base64,...` screenshot; the field is omitted when the `screenshot` feature is not enabled. **Response (Firecrawl format; HTTP 200 for both success and failure):** ```json { "success": true, "data": { "markdown": "...", "html": "...", "metadata": { "title": "Example Domain", "sourceURL": "https://example.com/", "description": "...", "statusCode": 200 } } } ``` --- ## Session API (Interactive Browser Sessions) Persistent browser sessions with indexed interaction. Each session gets its own V8 runtime + page context, and is reclaimed automatically after 8 minutes of inactivity. Lets AI agents browse the web the way a human does: open a page → inspect state → click/type → collect results. ### POST /session/create Create an interactive browser session. **Request fields:** | Field | Type | Required | Default | Description | |------|------|------|------|------| | url | string | | `null` | Initial URL (optional); `start_url` is accepted as an alias — creation navigates there before the session id is handed back | | use_proxy | bool | | `false` | Route through a proxy | | cookies | string[] \| object[] | | `[]` | Cookies injected before navigation (`"name=value",...` or cookie objects) so the session starts already logged in | | persistent | bool | | `false` | Persist login state to the server-side store: if the session idles out or the server restarts, the same `session_id` revives logged-in on the next call (`session/{id}/close` drops the snapshot; idle expiry keeps it) | | account | string | | `null` | Run as a named login identity (see [Named Accounts](#named-accounts-multi-login)): a private cookie jar seeded from the account record, written back after every action — concurrent logins on different accounts never clobber each other, and none touch the anonymous shared jar. The identity's device persona (own UA + hardware fingerprint) rides every session | **Response:** ```json {"session_id": "s_1", "url": "https://example.com/"} ``` ### POST /import/curl Turn a DevTools **"Copy as cURL"** command into a logged-in browser session — the credential-transfer path that works on a stock headless engine (no extension, no debug port). The human does the hard part of a login (CAPTCHA, SMS, slider) in their own Chrome, opens DevTools → Network, right-clicks any authenticated request → *Copy* → *Copy as cURL*, and pastes the command here. The engine parses out the Cookie header / `-b` jar, injects it into a fresh session anchored at the copied request's URL — the agent continues from where the human left off, no password or second login. bash, PowerShell and cmd copy flavors all parse. **Request fields:** | Field | Type | Required | Default | Description | |------|------|------|------|------| | curl | string | ✅ | — | The copied cURL command | | use_proxy | bool | | `false` | Route the session's traffic through the `AGINXBROWSER_PROXY` proxy | | account | string | | `null` | Attach the session to a named account: the imported login lands in the account's private jar and is written back under its name — one import per identity, no clobbering. On first import the copied command's User-Agent becomes the identity's persona (its device UA, reused by every later session) | **Response:** ```json { "session_id": "s_1", "url": "https://example.com/member/home", "host": "example.com", "cookie_count": 12, "method": "GET", "has_body": false, "user_agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) Chrome/145.0.0.0", "authorization_prefix": "Bearer eyJhbGciOi…", "expires_in_secs": 480, "note": "login state injected; navigate with session tools" } ``` The `session_id` is an ordinary session — drive it with the session API. `authorization_prefix` previews only the first 16 chars (the full header is a credential; the response never echoes whole cookies). `method`/`has_body` report what kind of request was copied. Cookies anchor on the copied request's host; other subdomains of the site may re-authenticate — that is the site's device binding, not a lost credential. **Tips**: an XHR on the logged-in page usually carries the most complete cookie set (document requests sometimes miss the `httpOnly` API-session pair). `-b FILE` cookie jars are rejected — the server cannot read your disk. Treat a pasted command like a password: it carries the login state verbatim. ### POST /session/{id}/clone Derive a new session carrying the source's full login state — cookies, `localStorage`/`sessionStorage`, viewport pin, dialog policy, proxy and keepalive flags — with the source left untouched. Use it to snapshot a logged-in state before risky actions, or to run the same login in parallel sessions. The manual `cookies` → `session/create` round-trip this replaces is where a hand-edited cookie string clobbers a working login. **Response:** ```json {"session_id": "s_2", "cloned_from": "s_1", "url": "https://example.com/dashboard", "viewport": {"width": 390, "height": 844, "mobile": true}, "expires_in_secs": 431} ``` `viewport` is `null` when the source never pinned one. ### GET /session/list List live sessions — the discovery twin of `/session/create` (reuse an idle session instead of spawning a fresh V8 thread per step). Entries carry idle age and the eviction budget; most recently active first. **Response:** ```json {"count": 1, "sessions": [{"session_id": "s_1", "idle_secs": 19, "expires_in_secs": 460}]} ``` Sessions are process-global and shared across callers — that's what makes "one instance per machine, every agent shares it" work. ### GET /sessions The full listing — every live session with its current page URL, so an agent resuming work can find "the session that was on the publish page" without guessing ids. Each entry asks its session thread for the URL with a 250ms budget, so a session pinned mid-eval reports `"url": null` instead of stalling the listing. **Response:** ```json { "count": 2, "sessions": [ {"session_id": "s_1", "url": "https://example.com/publish", "idle_secs": 12, "expires_in_secs": 468, "keepalive": false, "persistent": true}, {"session_id": "s_2", "url": null, "idle_secs": 3, "expires_in_secs": 477, "keepalive": false, "persistent": false} ] } ``` `expires_in_secs` is absent for `keepalive` sessions (they never auto-evict); `persistent` marks sessions whose login state is snapshotted and revivable. Close entries with `POST /session/{id}/close`. `/session/list` remains the lightweight sibling (id + idle age only, no per-session round-trips). **Restarts and upgrades** — on SIGTERM/SIGINT every live session (persistent or not) is flushed to the store before the process exits, and the next process revives them all under their own ids before it opens its listener: a binary swap costs no login state and no session ids, and `GET /sessions` right after a restart already shows the fleet. Restored sessions keep their original flags — a non-persistent session that then idle-expires takes its restart snapshot with it (a later touch of the id is `SESSION_NOT_FOUND`, not a revival with stale cookies). The flush is budgeted (60s overall, 10s per session): a session stuck mid-command keeps whatever its last per-command snapshot held. An explicit `close` still means done. ### Named Accounts (multi-login) Anonymous traffic shares one process-global cookie jar — deliberate (repeat-visitor cookies cut CAPTCHA rates) but it makes multi-account work impossible: two logins on one site clobber each other's session cookies, last writer wins. A **named account** is the fix — Chrome-profile semantics: one jar per identity, not a site×account matrix. `taobao-scraper` and `taobao-publisher` are two accounts; each keeps its own cookies and login state, written back to the store after every action, surviving idle eviction and server restarts. Sessions on the *same* account share one live jar (two tabs, one profile); different accounts never touch each other — and none touch the anonymous shared jar, since shared cookie history is itself a risk-control linkage signal. Attach a session with `account` on `POST /session/create` or `POST /import/curl` (1-64 chars of `[a-zA-Z0-9_-]`). The account is the persistence — account sessions don't need `persistent: true`. **Per-account device persona** — an identity is one stable *device*, not just one cookie jar. On first use an account draws a persona — a User-Agent from the Windows/macOS Chrome-145 pool plus a hardware seed that drives its `screen`/devicePixelRatio/GPU/canvas fingerprint — remembers both in its record, and reuses them for every later session, restarts included. Jar isolation alone can't stop risk-control linkage: two logins sharing one fingerprint read as "one device with two accounts". `import_curl` seeds the persona from the copied command's real User-Agent (the device the site already saw alongside those cookies). The pool is pinned to Chrome 145 so UA and the default TLS handshake never disagree; the hardware seed itself stays server-side — `GET /accounts` reports only the persona UA (`persona_ua`). **List accounts:** ```bash curl -sS http://127.0.0.1:8089/accounts ``` ```json { "count": 2, "accounts": [ {"name": "taobao-publisher", "domains": ["taobao.com", "tmall.com"], "cookie_count": 34, "updated_at": 1762934400, "persona_ua": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/145.0.0.0 Safari/537.36", "verify_url": "https://www.taobao.com/", "verify_predicate": "!!document.querySelector('.user-nick')", "verify_last": {"logged_in": true, "url": "https://www.taobao.com/", "checked_at": 1762934401}}, {"name": "taobao-scraper", "domains": ["taobao.com"], "cookie_count": 29, "updated_at": 1762933900, "persona_ua": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/145.0.0.0 Safari/537.36"} ] } ``` Metadata only — cookie values are credentials and never leave the server. **Check login (teach-once):** ```bash curl -sS -X POST http://127.0.0.1:8089/account/verify \ -H 'content-type: application/json' \ -d '{"name": "taobao-scraper", "url": "https://www.taobao.com/", "predicate": "!!document.querySelector(\".user-nick\")"}' ``` ```json {"name": "taobao-scraper", "logged_in": true, "url": "https://www.taobao.com/", "checked_at": 1762934455, "verify_url": "https://www.taobao.com/"} ``` The first call teaches the spec — `url` (a page that shows login state) and `predicate` (a JS expression truthy when logged in). It's remembered on the account; later calls can be bare (`{"name": "taobao-scraper"}`) and rerun the spec. The probe runs in a scratch session *as the account* — private jar, same egress — so a verify doubles as a cookie refresh. A probe failure (network, predicate throw) returns an error and is not a logout verdict; the last successful verdict rides in `verify_last`. **Log an account in (login wizard):** ```bash curl -sS -X POST http://127.0.0.1:8089/account/login \ -H 'content-type: application/json' \ -d '{"name": "xhs-main", "url": "https://www.xiaohongshu.com/login"}' ``` ```json {"name": "xhs-main", "status": "waiting", "needs": ["qr"], "markers": ["qr:login-qrcode-img"], "verdict": "login", "url": "https://www.xiaohongshu.com/login", "session_id": "s_42", "expires_in_secs": 1793, "handoff": "human step (qr) — open /live?session=s_42 ...; when done, re-call account_login with the same arguments (the account's shared jar already holds the finished login)"} ``` One call opens the login page *as the account* (implicitly created on first use; private jar + persona + recorded egress from the start) and probes the generic gates the page shows — `password` / `sms` / `qr` / `slider` — with `markers` carrying the evidence strings. The three statuses: - **`logged_in`** — no gate detected and the passed `predicate` fired: the automatic bounce happened, and the verify spec (`url` = the post-login landing URL, `predicate` as given) is stamped. Cookies are already written back (every successful action captures the account jar). - **`waiting`** — a human step is outstanding (`needs` non-empty). The session stays open (30 min ttl) for the `/live?session=` handoff or scripted `session_wait`/`session_input`. When the human finishes, **re-call with the same arguments** — the account's shared jar already holds the login, so the second call probes clean, the predicate fires, and verify is stamped. Idempotent by jar semantics, no state machine. - **`opened`** — no `predicate` given: page is open, gates reported; drive it with `session_input`/`session_wait` on `session_id` and finish with `account_verify` to stamp the spec. `use_proxy` seeds a fresh account's egress; an account with an existing record always reuses its recorded egress (one identity, one exit — changing it is a risk-control linkage signal). `timeout_ms` (default 60000, clamped 1000..120000) is the wait budget for the automatic bounce only — it is never spent while a human step is outstanding. The wizard never fills credentials or beats risk control; when a site rejects the environment, `needs`/verdict say so and `import_curl` (human logs in their own Chrome) remains the fallback. **Delete an account:** ```bash curl -sS -X DELETE http://127.0.0.1:8089/accounts/taobao-scraper ``` `{"deleted": "taobao-scraper"}` — stored record AND live jar. Cookie values are credentials; delete means gone. Unknown names get HTTP 404. ### POST /session/{id}/navigate Navigate to a new URL. **Request fields:** | Field | Type | Required | Description | |------|------|------|------| | url | string | ✅ | Target URL | **Response:** ```json {"url": "https://example.com/page2", "title": "Page 2"} ``` When the navigation lands on an anti-bot wall (punish page / `_____tmd_____/punish`), a `challenge` field joins the response — risk-control pages answer like ordinary pages, the flag is the machine-readable verdict: `{"url": "https://punish.taobao.com/", "title": "...", "challenge": "punish"}`. When HTTP redirects happened in flight, `redirected_from` joins the response with the full trail (issue #102): `redirected_from[0]` is the URL you asked for, `url` is where the document actually came from. The trail alone separates a login bounce (`login.example.com` → `example.com/dashboard`), a parameter-error rewrite (`…?err=invalid` destination) and a direct landing — same final URL, three different situations. Absent when the navigation involved no redirect. When a navigation exceeds its deadline (`AGINXBROWSER_NAV_TIMEOUT_MS`, default 30s), the error names the phase that died instead of a bare timer: `document fetch in flight` / `document received, not yet committed` / `document committed, scripts/settle still running` / `loaded after the deadline raced`, plus the **active document URL** and the navigation generation — `navigation exceeded 30000ms deadline (document fetch in flight; active document: https://old.example/; requested: https://new.example/; navigation #3 — /network rows carry the generation as "nav")`. If the attempt never committed, the session's url rolls back to the document still running, so `navigate`'s failure report, `/state` and `eval` describe the same page (#101). ### POST /session/{id}/preload Replace the session's document-start preload group (issue #96). Sources run before each new document's own scripts — **including inline `