--- name: extract description: Crawl an existing website (capped, multi-page) and seed stardust/current/ with PRODUCT.md, DESIGN.md, DESIGN.json, a per-page inventory, and the consolidated brand surface — the captured design system, palette, typography, motifs, and voice of the live site. Use when the user wants to analyze an existing site's design, extract or reverse-engineer its design system or brand, capture design tokens from a live site, import a website as the starting point for a redesign, capture the current state before a migration, or invokes `$stardust extract` (`/stardust:extract` in Claude Code). Trigger phrases include "analyze this site", "extract the design tokens", "capture the brand", "crawl the site", "reverse engineer the design". Not for scraping page data or content for its own sake (it captures design evidence, not datasets), and not for the redesign itself — extraction is descriptive; direction and prototyping happen downstream. license: Apache-2.0 compatibility: Requires Node 22+, Playwright with Chromium resolvable from the project, playwright-cli on PATH, and the impeccable skill (github.com/pbakaus/impeccable) installed alongside stardust. --- # stardust:extract Crawl an existing website, parse each page, extract the brand surface, and produce a stardust-formatted snapshot of the current state under `stardust/current/`. The output describes what the site **is**; later sub-commands consume it to decide what it **should be**. This skill is **descriptive**: it does not invent direction, it does not critique, and it does not modify the live site. It writes only under `stardust/current/` and updates `stardust/state.json`. ## Inputs - `` — required. The origin to crawl. Examples: `https://example.com`, `https://example.com/shop`. A path narrows the same-origin crawl to that subtree. - `--cap ` — optional. Override the default 5-page cap. The cap is intentionally small — a 5-page sample (home + four IA pillars/templates) is enough for cross-page brand aggregation, system-component detection, and the brand-review HTML; lift it (e.g. `--cap 25`) when a deeper crawl is genuinely needed. - `--all` — optional. Lift the cap entirely; extract every discovered page after junk filtering. Equivalent to `--cap 0`. Use when the user spontaneously asks for a full crawl. - `--pages ` — optional. Restrict the crawl to specific paths (slugs derived per `reference/ia-extraction.md`). Bypasses the cap. - `--refresh ` — optional. Re-extract one page that already exists in `state.json`. - `--single` — optional. Equivalent to `--cap 1`. Useful for testing. - `--wait ` — optional. Wait strategy per page. Default `medium`. See `reference/playwright-recipe.md` § Wait modes. - `--no-junk-filter` — optional. Disable the default junk-page filter in discovery (see `reference/ia-extraction.md` § Filtering). - `--no-consent-dismiss` — optional. Skip the pre-flight consent / cookie banner dismissal (see `reference/playwright-recipe.md` § Pre-flight: consent dismissal). Use when the redesign scope includes the consent surface or the dismissal's side-effects (script activation that wouldn't otherwise run) must be avoided. Default is to dismiss, keeping screenshots, voice aggregation, and per-section style unpolluted by the banner. - `--dynamics` — optional, **migration-bound**. Record per-page reach signals of the dynamic surface (data endpoints, forms, modal triggers, player ids) in each page JSON `dynamic` section and roll them up in `_crawl-log.json#dynamicSurface`. Set by `prepare-migration`, `replica` and `migrate`'s safety net; never by a bare extract, `uplift` or `audit` — dynamics is a migration concern. Depth and classification belong to the stardust `dynamics` skill. - `--concurrency ` — optional. Parallel browser contexts for the per-page capture loop. Default 4; sane range 4–8. See § Concurrency. - `--brand-source ` — optional, repeatable. An additional **same-brand** origin whose brand surface enriches the primary extraction (shallow capture: home + up to 2 nav-linked pages). See § Cross-site brand sources. - `--design-source ` — optional. Design-donor origin: its design system is captured to `stardust/canon-source/` and becomes the fixed redesign target while the primary origin supplies content. See § Cross-site brand sources. - `--prep` — optional. Run in **migrate-prep mode**: lift the cap, type each page, detect module candidates, capture typed content slots, emit the prep summary. See § Prep mode below. Typically invoked via the `prepare-migration` orchestrator skill rather than directly. ## Setup Run the master skill's setup procedure first (`skills/stardust/SKILL.md` § Setup): impeccable dep check, context loader, state read. Additional checks for this sub-command: 1. **Playwright availability.** The extraction step needs a real browser. Detect Playwright in this order: a Playwright MCP server, then a project-importable `playwright` module. **The `npx playwright` probe is NOT sufficient** — it confirms the CLI (which resolves a global install) but the recipe and `scripts/crawl.mjs` do `import { chromium } from 'playwright'`, and ESM module resolution does **not** honour a global install or `NODE_PATH` — the import throws `ERR_MODULE_NOT_FOUND` even where `npx playwright --version` succeeds. Verify the module is import-resolvable from the project root (probe: `node -e "import('playwright').then(()=>process.exit(0))"`); if it isn't, run `npm i -D playwright pixelmatch pngjs cheerio --legacy-peer-deps` (devDependencies, never `--no-save` — #125) or use the Playwright MCP server, before crawling. The `--legacy-peer-deps` flag is required on `aem-boilerplate` targets (their pinned `eslint@8` makes a plain `npm i` exit `ERESOLVE` before playwright is even considered). Don't trust the CLI probe alone. **`--no-save` installs are ephemeral.** Any later real `npm i` (e.g. a setup step adding a devDependency) prunes non-manifest packages, silently removing playwright mid-pipeline. Every downstream skill that renders (prototype, migrate, deploy, diff) must re-run the import-resolvability probe — and re-install on failure — at the start of its own run, not assume extract's install survived. **Script location matters.** ESM resolves `import 'playwright'` from the *script's* directory and the plugin tree ships no `node_modules`: copy every script byte-identical into the project (`stardust/scripts/…`) and run the copy. **Bundled crawler.** `skills/extract/scripts/crawl.mjs` is a runnable reference implementation of this whole sub-command — browser config + bot-management fallback, consent dismissal, wait + scroll, the capture list, per-page full-page screenshots (`assets/screenshots/.png`, consumed by the Phase 2.5 vision gate), response validation, and the § Capture-hygiene hardening (visibility filter, interstitial drop, SPA-shell flag, modal `textContent` capture, tracking-pixel discounting, cross-page duplicate detection). Prefer invoking it over hand-rolling a Playwright script per run; extend its in-page `capture()` to cover any recipe field it doesn't yet emit. 2. **Origin collision.** If `stardust/state.json` already records `site.originUrl` and the new `` is a different origin, stop and ask before clobbering. Stardust does not silently mix two sites in one project. **Flow guard (migration asks only).** If the ask carries migration intent ("migrate", "to EDS", "re-platform", "1:1", "replica") and `stardust/state.json` exists — or is about to be created — without `flow`, hand back to the master skill § Two migration flows before crawling: the flow is chosen and stamped there, and a keep-design ask enters through `replica` (which invokes this skill with `--prep` itself). A bare `extract ` for a redesign, audit or uplift is unaffected. (Recorded: `extract` on a raw URL as the entry of a same-design migration; the agent then built its own importer beside `replica`.) 3. **Browser contexts.** Open a fresh `BrowserContext` per capture worker (§ Concurrency; default 4). Run the **consent dismissal pre-flight** per `reference/playwright-recipe.md` § Pre-flight: consent dismissal *unless* `--no-consent-dismiss` is set. Cookies persist within a context but **not across contexts** — re-establish consent state per context (re-run the dismissal on the worker's first page, or clone the probe context's `storageState`). Record the resolved method in `_crawl-log.json#consent.method`. 4. **Bot-management probe.** When the first navigation in the run returns `ERR_HTTP2_PROTOCOL_ERROR` or `ERR_QUIC_PROTOCOL_ERROR`, or hangs through the entire hard-cap on what should be a fast origin, **do not retry headless** — switch to headed Chrome per § Failure modes → Bot-management block and record the switch in `_crawl-log.json#discovery.fetchTechnique` so re-runs start in headed mode without rediscovering the issue. ## Procedure ### Phase 1 — Discovery Discover the page inventory before crawling. Procedure in `reference/ia-extraction.md`. In summary: 1. Fetch `/sitemap.xml`, then `/sitemap_index.xml`, then check `robots.txt` for `Sitemap:` directives. 2. If no sitemap is reachable, run a same-origin BFS crawl from ``, depth-limited to 3, link-extracting from rendered HTML. 3. Filter the discovered URL list: same origin only, exclude `mailto:`, `tel:`, anchor-only links, query-only variations, common asset paths (`.css`, `.js`, `.pdf`, image extensions). 4. De-duplicate trailing-slash variations. 5. Apply the junk-page filter (`reference/ia-extraction.md` § Junk-page filter) unless `--no-junk-filter` is set. Surface the filtered list to the user as overridable. 6. Apply the cap (default 5, or `--cap`, or `--all` for no cap) and **proceed silently**. Print an informational summary of what was kept and what was cut — but do **not** gate on user confirmation. Users who want different scope set it spontaneously at command time: ``` $stardust extract https://example.com # default 5 pages $stardust extract https://example.com --cap 25 # bump to 25 $stardust extract https://example.com --all # lift the cap $stardust extract https://example.com --pages home,about,pricing $stardust extract https://example.com --single # just the entry URL ``` The agent reads spontaneous scope intent from the user's prompt (e.g. "extract all pages", "look at just the home and pricing", "do a full crawl") and applies the equivalent flag. No re-confirmation needed once intent is clear. Informational output (not a prompt — proceed immediately): ``` Discovered 38 pages on https://example.com (sitemap.xml). Filtered as likely junk (5): /test/, /sample-page/, /holiday1/, ... Selecting 5 highest-priority pages: - / (home) - /about - /pricing - /products - /contact Cut (28 pages, --all to lift): /blog/post-1, /blog/post-2, ... Extracting... ``` Selection heuristic: page-type checklist first, then score-based ranking (home + IA-pillar keywords + sitemap priority − archive / version markers). See `reference/ia-extraction.md` § Page selection and § Priority for the cap. The English-only keyword list is a known limitation for localized sites. 7. Write the discovered list to `stardust/current/_crawl-log.json` (created if absent) with `_provenance` and the full discovery reasoning, including `filteredAsJunk[]` and `userChoice`. This is an audit trail, not a state file. ### Phase 2 — Per-page extraction For each page in the cap-respecting list, render with Playwright following `reference/playwright-recipe.md`. Captures run **concurrently** per § Concurrency. The recipe is mandatory per page — in particular, do not skip the wait, scroll, or capture-list steps: - Viewport 1440 × 900 @ 2× DPR - Wait per the configured wait mode (default `medium`; see § Wait modes in `reference/playwright-recipe.md`) - Disable animations via `prefers-reduced-motion: reduce` - After the wait resolves, scroll to bottom in 4 viewport-height steps with 300 ms pauses, then return to top — this is required to trigger lazy-load and IntersectionObserver-driven content - Record `waitMs` and `waitMode` in the per-page `_provenance` Capture per page (full schema in `reference/current-state-schema.md`): - Page metadata (title, meta description, OG tags, theme-color) - Semantic structure: heading outline, landmark roles, sections - **Hero headline + lede (resolved)** — `heroHeadline` / `heroLede` picked by font-size × hero-region with a junk/hidden-state filter and a clean meta-description fallback (per `reference/playwright-recipe.md` § Capture list 5-bis). Required for JS-rendered sites whose document-order headings surface modal / promo / count junk instead of the real tagline. - Content: visible text per section (full innerText, **no truncation** per `reference/playwright-recipe.md` § Capture list 7), structured paragraphs (`body[]`), lists, FAQ Q/A pairs, and review/testimonial quotes per § Capture list 7-bis. Without these structured fields, every body region under a heading falls back to placeholder signature at migrate time. - CTA labels and href targets, link inventory (internal vs external) - Per-section computed style summary: dominant colors, font families in use, spacing rhythm, border-radius, shadows - Media inventory: img with `currentSrc`/`srcset` captured **with query strings intact** plus a `resolves` flag (HEAD/GET with browser UA + Referer), intrinsic dimensions, inline SVG count, video/iframe presence, `cssBackgrounds[]` (including pseudo-element `::before`/ `::after` walks per § Capture list 11) so `background-image` heroes and motifs do not silently disappear and 404ing CDN images are flagged before migrate ships `about:error`. - Font files captured via network-intercept (per § Capture list 16): every `woff2`/`woff`/`ttf`/`otf` response saved under `assets/fonts/` and recorded in `_brand-extraction.json#type.files[]` with licensing flag. - Icon-font detection (per § Capture list 17): when the page uses `[class^="icon-"]` with non-default `::before` font-family + codepoint, capture the family, save the file, and record the `iconClass → codepoint` table in `_brand-extraction.json#iconFont`. - Interactive elements: forms (with field types), buttons, modals detected by ARIA roles - Full-page screenshot to `assets/screenshots/.png` — script-captured by the bundled crawler after the wait/scroll settle (viewport-only fallback on extremely tall pages; mode in `_signals.screenshotMode`, relative path in the page JSON `screenshot` field) - **Dynamic surface (only with `--dynamics`)** — per-page reach signals: endpoints, third-party script hosts, forms, modal triggers, player ids, hydration hints, in the page JSON `dynamic` section and `_crawl-log.json#dynamicSurface` (schema in `reference/current-state-schema.md § Dynamic`). Evidence only; the stardust `dynamics` sub-skill probes archetypes in depth and decides. Save to `stardust/current/pages/.json` with `_provenance` as the first key. **The bundled crawler also saves the settled rendered DOM verbatim as `stardust/current/pages/.html`** (`page.content()` after the wait/scroll settle; path in the record's `renderedHtml` field). Capture once, parse offline: every downstream importer or sibling generator iterates its extraction against this artifact — free, reproducible, and provenance — instead of re-running live probes per selector guess (recorded: 4+ live round-trips per page family before the switch). Live probes stay for what the static DOM cannot answer: geometry and computed styles. Save referenced media to `stardust/current/assets/media/` preserving basename plus a short content hash. **Live-render evidence (synthesis is forbidden).** Refuse to mark a page `extracted` in `state.json` unless its `_provenance` contains `renderedBy: "playwright"`, an ISO-8601 `fetchedAt`, a positive integer `waitMs`, a `waitMode` from the recipe, and a final `httpStatus` in the 2xx/3xx range. These five fields are the contract enforced by `reference/current-state-schema.md` § Live-render evidence and read back by every downstream phase via `validateProvenance()` per `skills/stardust/reference/state-machine.md` § Provenance validation. Synthesizing a page record from `_brand-extraction.json` plus URL patterns plus captured photos — the 2026-04-30 e-commerce shortcut — is the failure mode this guard exists to prevent. When the agent (or a delegated sub- agent) cannot satisfy the contract for a page, treat the page as a Phase 2 failure: record under `_crawl-log.json#crawl.failures[]` with `errorClass: "ProvenanceMissing"` and continue. Mark the page `extracted` in `state.json` immediately after each successful page write. If a page fails, record the error in `_crawl-log.json` and continue — extraction is best-effort per page. ### Phase 2.5 — Vision verification Before anything downstream is authored, **look** at each captured page's screenshot (`assets/screenshots/.png`) — the multimodal model reads the image — and verify it against the extracted record: - Does the recorded hero (headline + asset) match what the pixels show? - Is the extracted palette plausible against the pixels? - Does a `cssBackgrounds: []` record look believable, or is imagery visibly present — a silent capture failure? - Is the logo captured? - Is the page actually rendered — not a consent wall, bot-block page, or blank SPA shell? Read the whole page through its thumbnail — `node stardust/scripts/thumb.mjs stardust/current/assets/screenshots/.png --width 480` (`skills/extract/scripts/thumb.mjs`, copied beside `crawl.mjs`; a box-filter downscale, so 1-px rules and hairline borders survive; its `--max-bytes` cap, 150 KB by default, re-encodes a heavy page narrower first and crops only as a last resort, never below 60 % of the page) — never the full-resolution screenshot, and read any detail as a crop of the full capture. The check is of the WHOLE page: read the script's stdout line before the image, and when it prints a share below 100 (`cropped at px of = %`) run the same command with `--offset ` and read `-thumb-.png` too, repeating while its `rows -px` line ends short of ``. A recorded hands-off run's thumbnails kept a median 28 % of each page and the check read them as the page. **A thumbnail enters the context only after `thumb.mjs` has written it (≤ 150 KB, or the floor thumbnail it named on stderr, exit 1) — or the vision check runs in a subagent that reads the thumbnails and returns one line per page** (`: ok | recaptured | suspect — `); a recorded hands-off run read eight `*-thumb.png` files of 268–608 KB each into one context — 3.2 MB of pixels for eight verdicts. On mismatch, re-run that page's capture with the escalation ladder before proceeding: bump the wait mode one step (`reference/playwright-recipe.md` § Wait modes), then headed Chrome (§ Bot-management fallback), then a fresh browser context. Record the outcome per page in `_crawl-log.json#visionCheck[]`: ```json { "slug": "pricing", "verdict": "recaptured", "notes": "record said zero CSS backgrounds; screenshot shows a full-bleed photo hero" } ``` `verdict` is `"ok" | "recaptured" | "suspect"` — `suspect` means the mismatch survived the ladder; downstream phases treat that record as unreliable. Vision is the **authoritative** capture check; the heuristic defenses (low-media flag, `spaShellSuspect`, duplicate hash) remain as cheap early signals but no longer gate alone. ### Phase 3 — Brand-surface extraction Run after the capture phases (2–2.5). Aggregation **may proceed incrementally** as concurrent page captures complete (§ Concurrency), but the written file must reflect every extracted page — including brand-source pages per § Cross-site brand sources. Produces `stardust/current/_brand-extraction.json` per `reference/brand-surface.md`. Some fields are home-only (logo, voice samples, register heuristic); the visual tokens that drive DESIGN.md (palette, radius, shadow, type) are aggregated across **all extracted pages** to avoid the home-page bias documented in `brand-surface.md` § Aggregation scope. **Computed-style census — run the shipped instrument; author no probe.** `node /skills/extract/scripts/style-census.mjs` (copied into the project like `crawl.mjs`, § Setup) measures every captured page at 1440 and writes `stardust/current/_computed-styles.json`; add `--width 360` when the breakpoints include it. One more live pass over every page — run it ONCE, after the crawl, in the background; it opens the live side through `live-session.mjs` like every other live instrument. Read the palette, type, motif and hover values from the file's `aggregate` and cite each from its `sources[]` (`url`, `selector`, `prop`); the per-page detail sits under `pages[url][width]`. **Content-cap model — run the shipped probe; author no lift (#124).** `node stardust/scripts/replica/cap-probe.mjs … --write-design stardust/current/DESIGN.json` (replica's script; copy it with `live-session.mjs`) renders each archetype at 1440 and at a DERIVED wide width — max(2560, largest cap × 1.25); a probe AT a cap's width cannot see it — and records every box that stops growing (kind `shell` / `content` / `module`, px, authored via, tier) into `extensions.breakpoints`: `containerMaxWidth` (the `--max-width` source; null when fluid), `probeWidth`, `caps[]`, `modules[]`. One navigation per archetype after the census (never beside it), once DESIGN.json is authored (it merges). A cap on the shell or on every archetype is design intent whatever its value; `capRegister` lists single-module or archetype-divergent caps for the inconsistency register. The cap model never records layout breakpoints: a redesign target keeps its own two; replica keeps the source's, or maps them to `--target-breakpoints`. Captures: - **Logo** by the v1 priority chain: inline SVG → `` with logo-ish class/id → `apple-touch-icon` → `og:image` → favicon → synthesized placeholder. Save to `stardust/current/assets/logo.`. - **Favicon** — ALWAYS captured as its own asset (independent of the logo chain) to `stardust/current/assets/favicon.`, per `reference/playwright-recipe.md` § Favicon capture. Downstream, `prototype` embeds it in the proposed page head and `deploy` ships it to the Edge Delivery site. - **Palette** — aggregate computed colors across **all extracted pages** (background, text, accents, borders, hovers). Frequency-sort, cluster near-duplicates, emit a role-named list (background, surface, text, primary, secondary, accent). - **Type** — font families in use with their weights, sizes, and computed line-heights. Identify the heading family vs body family. Run the modular-scale audit (`brand-surface.md` § Modular-scale audit) and emit `scaleAudit.kind = "modular" | "ad-hoc"`. - **Motifs** — signature border-radius (cross-page mode of non-zero values, weighted by element count), shadow stack (top 3 distinct, cross-page), gradient inventory, common patterns (chip, badge, card, hero-with-image). When the home-only mode disagrees with the cross-page mode, surface the divergence in `_provenance.notes`. - **Voice samples** — first paragraph of body copy, the hero headline, 3 representative CTA labels, a representative link list. Used by `direct` later but extracted now so the network round-trip is over. - **Hero image** — elevate the home page's primary visual asset to `voice.heroImage` (per `reference/brand-surface.md` § heroImage resolution), so downstream prototype picks the live hero rather than the `og:image` from the raw media list. - **Hero medium (signature)** — when the hero/first viewport carries a *moving* asset (background `