` wrapper, server-rendered iframes → drop the wrapping ``, Gravity Forms script blocks → strip on `gform`-mention, etc.
### Phase 7 — Copy `robots.txt`
Not linked from HTML; fetch it explicitly. Adjust the `Sitemap:` reference if the deployed sitemap path differs from the source.
### Phase 8 — Verify locally
Serve from `output/` with `python3 -m http.server`, then run the verify checklist in `AGENTS.md`:
1. Every URL in `urls.txt` resolves to a file (no missed pages).
2. No remaining `https:///` outside the canonical / `og:url` / JSON-LD allow-list.
3. No broken internal links from `wget --spider`.
4. Spot-check the homepage and a deep page in a browser. Watch srcset images, sidebar widgets, and the header banner — those break silently if missed.
### Phase 9 — Deploy
Host-specific recipes in `AGENTS.md`:
- **Cloudflare Pages** — `_redirects`, `_headers`, "no build command, no output directory" defaults.
- **Netlify** — same `_redirects` / `_headers` syntax, plus `netlify.toml`.
- **Vercel** — `vercel.json` with `redirects` / `headers`.
- **Generic** — nginx `try_files`, Apache `Options +MultiViews`.
## Gotchas
These are the things that bit us. Don't repeat them.
1. **Cloudflare bot protection 403s the default `Wget/1.x` UA.** Always set a real browser UA + `Accept` / `Accept-Language` headers (recipe). If you see `403 Forbidden` after a burst of requests, that's it — back off, switch UA, retry.
2. **Cross-page link rewriting only works in a single wget invocation.** wget's `-k` only rewrites to local paths it sees in the current run. If a page was downloaded in a separate invocation (e.g. to recover from a 403 on one URL), its links to the rest stay absolute. Solution: redo the full scrape once you have the right UA. Don't piecemeal it. If you're scraping at scale (10K+ URLs) and can't fit in one run, scrape in batches and re-run `scripts/rewrite-paths.py` afterwards as the canonical pass — `-k`'s output is then redundant.
3. **Default publish directory by host.** Cloudflare Pages serves the repo root when no build command is configured. Netlify and Vercel also default to root. If you scraped into `output/`, either move files to the repo root (`git mv output/* .`) or configure the host to publish from `output/`. Symptom of the wrong setup on Pages: every URL 404s with R2-style headers (`access-control-allow-origin: *`, `cache-control: no-store`) instead of a Pages-branded 404.
4. **WordPress Offload Media plugins** route `/wp-content/uploads/` to R2 / S3 buckets. wget may successfully fetch an image even when later direct access 404s (intermittent or partial bucket sync). Trust your local copy — that's why we scrape and self-host.
5. **Sitemaps and the Yoast XSL aren't linked from HTML.** wget `-p` won't find them. Fetch explicitly in Phase 1.
6. **Filenames with `?ver=...` query strings.** wget keeps these as literal filenames; HTML uses `%3F` encoding. Standard servers (Pages, Netlify, Vercel, `python -m http.server`) URL-decode and serve correctly. Don't try to "clean these up" unless something actually breaks.
7. **`og:url`, canonical, JSON-LD stay absolute.** They identify the canonical resource and are correct as-is when redeploying to the same domain. Only rewrite if changing domains.
8. **`sed -i ''` is macOS / BSD only.** GNU sed needs `sed -i` (no empty-string argument). Recipes in `AGENTS.md` flag the macOS-isms; default to the Python scripts where there's a choice — they're portable.
## Output structure
```text
/
index.html ← homepage
/index.html ← one per URL from sitemap
wp-content/ ← assets (themes, uploads, plugins)
wp-includes/ ← block library CSS, et al.
avatars/ ← self-hosted Gravatars (Phase 6)
sitemap_index.xml
page-sitemap.xml ← + any other child sitemaps
wp-content/plugins/wordpress-seo/css/main-sitemap.xsl
robots.txt
_redirects ← optional, host-specific
_headers ← optional, host-specific
```
Push to a git host and connect to the static host with **no build command** and **no build output directory** — defaults work.