--- name: ar-io-gateway-operator description: Operate any AR.IO node deployment — architecture, daily ops, diagnostics, and recurring pitfalls that apply to every operator. Use whenever the user asks about gateway health, indexing lag, ClickHouse pipeline, GraphQL routing, ArNS resolution, data retrieval and caching, bundle unbundling, webhook fan-out, container restarts, image rebuilds, disk pressure, HTTPSig overhead, or anything operationally generic to running an AR.IO node — even when they don't name a specific service ("is it healthy?", "why is it slow?", "did the build pick up?", "how full is the disk?"). For deployment-specific knowledge (domains, repo paths, sidecar config, this-machine numbers), pair this with the deployment-overlay skill (e.g. `vilenarios-gateway`) if one exists. --- # AR.IO Gateway Operator (generic) You are operating an AR.IO node — an Arweave gateway that indexes, caches, serves, and signs blockchain data. Deployments share the same architecture but differ in domains, paths, and which sidecars are present. **Run `scripts/health-check` first** every session: it prints a one-screen snapshot you can compare against the prior state before deciding to act. If a deployment-specific overlay skill is loaded (e.g. `vilenarios-gateway`), it provides the local repo paths, domains, and which sidecars exist. Treat the overlay as authoritative for those facts. ## Run first: `scripts/health-check` ```bash # from the ar-io-node repo root .claude/skills/ar-io-gateway-operator/scripts/health-check ``` What it checks and why: | Section | What "good" looks like | If wrong, jump to | |---|---|---| | containers | All core ar-io-node containers `Up` | "Bring everything up" | | disk | `/` < 80%, `/data` (if separate) has headroom | "Disk pressure" | | gateway healthcheck | `{"status":"ok",...}` | "When something breaks" | | indexer height vs chain head | lag ≤ 5 blocks | "Indexing falling behind" | | ClickHouse rows | If streaming enabled: `new_blocks > 0` and growing | "Streaming tables empty" | | core errors last 5 min | < ~10 (operational noise: peer fetches, aborts) | grep deeper | | image SHAs | If on a dev branch, must differ from `:latest` digest | "Building from source" | The script only covers the core ar-io-node containers defined in this repo's docker-compose. Sidecar checks belong in a deployment-overlay skill that calls this script then adds its own checks on top. ## How the gateway works (the model an operator needs) Knowing the data flows tells you where to look when symptoms appear. Source of truth is `src/system.ts` (DI wiring) and `docs/`; this section captures the operationally-load-bearing facts. ### Request flow ```text client → (optional edge: nginx/CDN) → envoy :3000 → core :4000 → response with X-AR-IO-* trust headers ↘ webhook fan-out → optional sidecars ``` Every served data response carries trust headers (`X-AR-IO-Verified`, `X-AR-IO-Trusted`, `X-AR-IO-Stable`, `X-AR-IO-Hops`, plus `X-Arweave-*` tag passthrough). Those headers are what HTTPSig signs — see `src/middleware/httpsig.ts` `TRIGGER_HEADERS`. Endpoints with no trust trigger (`/ar-io/info`, `/ar-io/healthcheck`, `/graphql` errors) are deliberately not signed. ### Data retrieval — composite source with fallback chain Configurable via `.env`. Defaults: - `ON_DEMAND_RETRIEVAL_ORDER=trusted-gateways,ar-io-network,chunks-offset-aware,tx-data` — what `/raw/` and `/` walk through on a cache miss - `BACKGROUND_RETRIEVAL_ORDER=chunks` — used by the unbundler when fetching bundle bodies Each source's failures are logged with `class` so you can grep (`GatewaysDataSource`, `TxChunksDataSource`, `ChunksDataSource`, etc.). Persistent failures from one tier are normal — the chain is designed to fall through. Look for `Failed to fetch chunk data from any source` / `all sources failed` as the *real* failure signal, not individual tier warnings. Three response headers worth knowing: `X-Cache: HIT|MISS` (gateway-side contiguous cache), `X-AR-IO-Hops` (how many AR.IO peers were traversed before this gateway served), `X-AR-IO-Trusted` / `X-AR-IO-Verified` (whether the data has been hash-verified against the chain). ### Indexing pipeline Two write paths run side-by-side: **1. SQLite primary indexer (always on)** — `BlockImporter` worker imports blocks one at a time into `new_blocks`/`new_transactions` (the unstable head, ~last 18 blocks), and a separate process promotes them to `stable_blocks`/`stable_transactions` once they're past the reorg horizon. ANS-104 bundle unbundling is gated by `ANS104_UNBUNDLE_FILTER` (often narrow). Database access is in a worker thread (`StandaloneSqlite`) — never call SQLite synchronously from the main thread. **2. ClickHouse pipeline (optional, `--profile clickhouse`)** — Stable rows are exported to parquet then imported by `clickhouse-auto-import` (hourly cycles) into `staging_*` and finally `transactions`. **Streaming pipeline** (PR #699 / `clickhouse-streaming-unstable-head` branch) mirrors the unstable-head SQLite writes directly into ClickHouse `new_blocks` / `new_transactions` so GraphQL can query the head without waiting for the parquet round-trip. Activated by `CLICKHOUSE_STREAMING_ENABLED=true`. Reorgs trigger bounded `ALTER TABLE ... DELETE WHERE` on the `new_*` tables. ### GraphQL routing `/graphql` reads from up to three legs that get merged: stable `transactions` table in CH, streaming `new_transactions` in CH, and the SQLite `new_*` fallback. With `CLICKHOUSE_GQL_SKIP_SQLITE_READS=true`, the SQLite leg is dropped to a tight-timeout fallback only. With `GRAPHQL_ON_DEMAND_RESOLUTION_ENABLED=true` (default), `transaction(id)` queries that miss the local DB will fall back to extracting from ANS-104 bundle binaries on demand — single-ID only, capped by `GRAPHQL_ON_DEMAND_RESOLUTION_MAX_CONCURRENT`. ### ANS-104 unbundling pipeline Whenever a bundle TX matches `ANS104_UNBUNDLE_FILTER`, it walks a four-queue chain. Knowing the chain is essential for diagnosing throughput regressions — a single queue at cap can mask everything downstream of it. ```text bundleDataImporter ─→ ans104Unbundler ─→ ans104DataIndexer ─→ dataItemIndexer ─→ SQLite writer download workers parser workers post-parse buffer pre-write buffer single-writer (ANS104_DOWNLOAD_…) (ANS104_UNBUNDLE_) (ANS104_DATA_…_SIZE) (DATA_ITEM_…_SIZE) ``` Tune knobs (all env-overridable): | Knob | Default | What it controls | |---|---:|---| | `ANS104_DOWNLOAD_WORKERS` | 5 | Parallel bundle downloads. Network-bound; overprovisioning is cheap | | `ANS104_UNBUNDLE_WORKERS` | 1 | Parallel parser worker_threads. CPU + network bound | | `DATA_ITEM_INDEXER_WORKER_COUNT` | 1 | Fastq concurrency for `indexDataItem`; SQLite is single-writer so gains are bounded | | `MAX_DATA_ITEM_QUEUE_SIZE` | 100000 | **Soft gate**: `shouldUnbundle()` returns false above this on either indexer queue | | `DATA_ITEM_INDEXER_QUEUE_SIZE` | 500000 | Hard cap on `dataItemIndexer`; items past this get dropped silently | | `ANS104_DATA_INDEXER_QUEUE_SIZE` | 500000 | Hard cap on `ans104DataIndexer` | **The buffer between soft gate and hard cap is where data loss happens.** With defaults the gate engages at 100 K but the hard cap is 500 K — a 5× buffer. Under heavy load, in-flight parser workers can flood items into the queue *after* the gate engages but *before* the indexer drains, exceeding the hard cap and silently dropping items. Symptom: `data_items_dropped_total{queue_name="dataItemIndexer"}` climbs while `bundles_unbundled_total` keeps increasing — bundles look "indexed" but their items aren't searchable. Mitigations: - Tighten `MAX_DATA_ITEM_QUEUE_SIZE` so the gate engages well before the hard cap (e.g., 25 000 with 500 000 hard caps, giving a 20× buffer). - Raise the hard cap if memory permits (each item slot ≈ 1–2 KB). - Don't add parser workers when `data_items_dropped_total` is non-zero — more workers = bigger floods past the gate. Pipeline diagnostic: a single `curl -sf .../ar-io/__gateway_metrics | grep -E 'queue_length|data_items_dropped|bundles_unbundle_in_flight'` tells you which queue is the bottleneck. ### Caching layers (in order, fastest first) | Layer | Where | TTL | Notes | |---|---|---|---| | In-process kv | core memory | varies per resolver | ArNS names cache, etc. | | Redis | `redis:7` container | varies | Some ArNS / network state | | Contiguous data store | `${CONTIGUOUS_DATA_PATH}/` | unbounded | Files keyed by content hash; cleanup workers prune | | Response cache hints | `Cache-Control` headers in response | per-route | e.g., `/raw/:id` HIT serves with `max-age=43200` | | Edge cache (optional) | nginx/Varnish/CDN in front | per-zone | Only present if the operator deployed one — gateway doesn't run one itself | `X-Cache: HIT` on a `/raw/` response means the gateway's contiguous data store served bytes without going to peers/Arweave. The edge layer is *separate* — nginx's `X-Cache-Status: HIT` is a different layer. ### ArNS resolution `CompositeArNSResolver` walks resolvers in order: `TrustedGatewayArNSResolver` (asks `TRUSTED_ARNS_GATEWAY_URL`, default `https://__NAME__.turbo-gateway.com`), then `OnDemandArNSResolver` (queries the on-chain `ario-arns` and `ario-ant` programs directly via `SOLANA_RPC_URL`). The base name set is paginated into an in-memory `ArNSNamesCache` at boot and refreshed on a debounce. Resolved IDs may be Arweave TXs (route to data path) or, on the streaming-head branch, IPFS CIDs (route to `/ipfs/` via the Kubo sidecar). Unknown names log `Unable to resolve name against all resolvers` — that's normal user-error traffic, not a service failure. `/` and `//` requests carry `X-ArNS-*` trust headers in the response. Manifest path resolution still uses `StreamingManifestPathResolver`; the "from index" path is not implemented yet (logs warn `not implemented` then falls back to data-side resolution, which works). ### Network identity, observer, and incentives A gateway is a registered participant in the AR.IO network, not just a piece of software. Two distinct Solana identities matter: - **Operator** — `AR_IO_WALLET` (pubkey, base58) plus a signing key as either `SOLANA_KEYPAIR_PATH` (JSON file) or `SOLANA_PRIVATE_KEY` (base58 secret, Phantom-export form). Stakes ARIO tokens (20,000 minimum = 20000000000 mARIO since v3.0.0), receives rewards, signs cranker instructions when `ENABLE_EPOCH_CRANKING=true`. - **Observer** — `OBSERVER_WALLET` (pubkey) plus a signing key as either `OBSERVER_KEYPAIR_PATH` or `OBSERVER_PRIVATE_KEY`. Submits `save_observations` instructions and signs HTTPSig responses. The two roles are deliberately separable so the observer key can sit on the gateway box while the operator/staking key stays in cold storage. Setting both forms (path and PK string) for the same role is rejected at startup. The Solana key designated for the observer role *is* the HTTPSig signing key. The observer's address is already registered in the on-chain Gateway Registry (GAR), so verifiers can confirm a signed response traces back to a staked, registered gateway with a single GAR lookup. No attestation document is uploaded to Arweave on the Solana path — the chain is the source of truth. Four Solana program IDs configure which network the gateway talks to: `ARIO_CORE_PROGRAM_ID`, `ARIO_GAR_PROGRAM_ID`, `ARIO_ARNS_PROGRAM_ID`, `ARIO_ANT_PROGRAM_ID`. Plus `SOLANA_RPC_URL` (premium provider strongly recommended in production; public mainnet-beta throttles hard). `GET /ar-io/info` echoes the resolved `programIds` object, which is the fastest way to confirm a running gateway is pointed at the network you think it is. Observer service knobs: - `RUN_OBSERVER=true` (default) — runs the observer alongside the gateway - `SUBMIT_CONTRACT_INTERACTIONS=true` (default) — submits generated reports on-chain - `OBSERVATIONS_PER_GATEWAY` (default 3) — how many times each gateway is observed per epoch (majority vote) - `OBSERVATION_WINDOW_FRACTION` (default 0.5) — what fraction of the daily epoch window observations are spread across - `REFERENCE_GATEWAY_MIN_CONSECUTIVE_PASSES` / `REFERENCE_GATEWAY_MIN_EPOCH_COUNT` — eligibility thresholds for using a network gateway as a reference **Epoch cadence is daily.** Observation reports drive reward eligibility. Operators may themselves be selected as observers for other gateways. Network onboarding (operator side, only needed once): - `npm install -g @ar.io/sdk` — provides the `ar.io` CLI binary. Use `ar.io get-gateway --address ` to read state, `ar.io join-network ...` to register. - The CLI accepts the operator key as either `--wallet-file ` (JSON) or `--private-key ''` (Phantom export — `bs58.decode` runs in-process). - Every Solana command needs `-t solana` and the four program IDs (`--core-program-id` / `--gar-program-id` / `--arns-program-id` / `--ant-program-id`) plus `--rpc-url`. As of `@ar.io/sdk@^4.0.0-solana.22` the bundled "mainnet" constants are still placeholders, so passing `--mainnet` alone resolves to non-existent addresses — always pass the explicit program IDs the network team publishes for whichever environment you're targeting (mainnet, staging-devnet, local devnet). - Register via the CLI or the gateways.ar.io portal: operator wallet, observer wallet, FQDN, Properties Tx ID, label, stake (`--operator-stake 20000` in whole ARIO — the CLI converts to mARIO internally via `requiredMARIOFromOptions` → `new ARIOToken(...).toMARIO()`). - Gateway must be reachable at the registered FQDN with valid TLS, AND configured for the same `programIds` set you joined with. The usual cause of a "I joined but nothing's working" report is a gateway whose `.env` still pins the old network's program IDs. - Optional delegated staking — operators can accept delegations with a configurable reward share. `--min-delegated-stake` is mARIO directly (no ARIO→mARIO conversion); `--operator-stake` is the only stake-input flag that takes whole ARIO. ### Serving content at the root path Two mutually-relevant env vars decide what `/` serves: - `APEX_TX_ID` — pin a specific Arweave TX as the apex. Cache-Control bounded by `CACHE_APEX_MAX_AGE` with `must-revalidate` so apex rotations propagate. - `APEX_ARNS_NAME` — resolve an ArNS name as the apex. Comma-separated values position-map to `ARNS_ROOT_HOST` entries when there's more than one host. - `ARNS_ROOT_HOST` — the gateway's primary domain(s) for ArNS subdomain resolution. ### Index sharing sidecar (optional) The `index-swarm` sidecar (compose profile `index-swarm`) publishes this gateway's CDB64 root-TX index bands, signed with the observer key, and subscribes to other gateways' bands, installing them under `data/indexes/installed/` where the gateway loads them without a restart. It resolves publishers through the gateway's own `/ar-io/peers`, so it makes no RPC calls. The gateway serves `/ar-io/indexes` (the signed document) and the byte routes, metered like data. The optional torrent engine (profile `index-swarm-torrent`; setting `INDEX_SWARM_ENGINE_AUTH` turns it on) adds a BitTorrent tier: peers first, the HTTP routes as fallback, every subscriber seeding what it installs. The engine runs on its own Docker network and filters private addresses; on a LAN-only swarm set `INDEX_SWARM_ENGINE_BLOCK_PRIVATE=false`. Health: `index_subscription_manifest_age_seconds` (publisher gone quiet), `index_swarm_installed_bands`, `index_subscription_bytes_total{transport}`, `index_swarm_engine_available`. On a publisher, who shares: `index_swarm_tracker_seeding_hosts` (other hosts seeding a band) and `index_swarm_tracker_seeders`. Full guide: `docs/index-swarm.md`. **A fleet behind a load balancer** (docs/index-swarm.md, "Publishing torrents from a fleet"): only the node `/ar-io/indexes*` is pinned to publishes and runs the engine and tracker; its peer port goes to its own public IP, not through the LB (set `INDEX_SWARM_ENGINE_PUBLIC_HOST`), and the tracker is direct or `/announce` proxied with `INDEX_SWARM_TRACKER_TRUSTED_PROXIES`. Rate limiter and x402 need nothing new: byte routes and the WebSeed are metered, the document and `.torrent` are free. Other nodes that answer lookups subscribe over HTTP with `url` pointed at the publishing node, allowlisted in its `RATE_LIMITER_IPS_AND_CIDRS_ALLOWLIST`. **Ports and disk.** The engine's peer port (`INDEX_SWARM_ENGINE_PORT`, 6881 TCP+UDP) and the tracker's (`INDEX_SWARM_TRACKER_PORT`, 6969 TCP) are Docker-published, so ufw neither blocks nor protects them (filter where Docker forwards: the `DOCKER-USER` chain with Docker's default iptables backend; with its nftables backend, which has no `DOCKER-USER`, a chain in your own table on the `forward` hook at a priority before Docker's), and a port already held on the host stops the container starting. A publisher needs its peer port reachable, and the tracker port only when its torrents announce to it directly (`INDEX_SWARM_TRACKERS` naming this node's port): not when `/announce` is proxied through a load balancer, nor when it publishes with no tracker (it warns, and peers then meet only through DHT and the WebSeed). A subscriber needs neither (behind NAT it still downloads and seeds to peers it dials), and `index-swarm-status` reports a closed port as WARN, not FAIL. The engine has UPnP off. `INDEX_SWARM_MAX_DISK_BYTES` (50 GiB from the setup script) is a ceiling, not a reservation: nothing checks free space, so a full disk under `data/indexes` is the operator's to prevent. That directory must be one filesystem (renames, hard links); move it whole with `INDEX_SWARM_DATA_PATH`. Upload defaults to 10 MB/s and 100 GB a UTC day, up to about 3 TB a month. **Build and check bands with `./tools/ar-io-node`** (docs/cli.md, "For scripts and agents"): `index-band-build` (CSV records in, header check against `--gateway-url`, published into `data/indexes/published/root-tx-index`) and `index-band-verify`. JSON on stdout, exit 1 on failure, safe to rerun (an identical band reports `"unchanged": true`). Point `--gateway-url` at a gateway that holds the roots (`https://turbo-gateway.com`); a gateway that must fetch them times out and the band is refused as under-checked, not wrong. Any other command runs as the SDK's `ar.io` CLI. **Bands are built by the `index-export` service** (compose profile `index-export`; docs/index-swarm.md, "Producing bands"), once a day at `INDEX_EXPORT_RUN_AT_UTC` from `INDEX_EXPORT_SOURCES` (this gateway's ClickHouse, else SQLite; peers and a bundler overlay merge). It needs `INDEX_EXPORT_START_HEIGHT` for its first run and `INDEX_EXPORT_HEADER_CHECK_URL` always (a gateway holding the roots; not `core:4000` on a gateway that fetches them). Dry-run before starting: `docker compose --profile index-export run --rm index-export --once --dry-run`. Bring it up by name with `--no-deps`. Its `/healthz` is loop liveness only; outcomes are `index_export_runs_total{result,reason}` and `index_export_last_success_timestamp_seconds{kind="d"}` (alert past 2 days). `couldnt_check` retries itself (15 min doubling to 2 h); `rejected` (a wrong header, conflicts over 1%) never does and needs a look. A stale `data/indexes/export/lock` (5 min untouched) is cleared by the next run; `state.json` there is re-derivable except adoptions. With `parquet-l1` in `INDEX_EXPORT_KINDS` it also builds L1 bands from `core.db`, and adds lookup files (layout `l1-3`) to published L1 bands in place, keeping their ids (`index_export_runs_total{kind="derive"}`). Upgrade a gateway that subscribes to `parquet-l1` before its publisher writes `l1-3` (an older one refuses those bands, naming the layout, and keeps its copies), and a publisher's `index-swarm` with or before its `index-export` (an older sidecar can't describe an `l1-3` band, so it goes unpublished). **Set up and check with the scripts, not by hand:** `./tools/index-swarm-setup --subscribe --torrent --dry-run` shows the `.env` changes (subscription, CDB64 sources, lookup order, engine password); `--restart` applies them, recreating only what needs it with the running gateway's compose files. `./tools/index-swarm-status` is the first thing to run when asked whether index sharing works: read-only, ok/WARN/FAIL per check with the fix, exit 1 on a failure. ### Filters and webhooks (the gateway's contract with sidecars) Three filters compose JSON expressions (`tags`, `attributes`, `or`/`and`/`always`/`never`) — see `docs/filters.md`: - `ANS104_UNBUNDLE_FILTER` — which bundles get unbundled - `ANS104_INDEX_FILTER` — which data items inside matched bundles get persisted - `WEBHOOK_INDEX_FILTER` — which indexed items trigger webhook emission Webhook events fired: `data-cached`, `tx-indexed`, `ans104-data-item-indexed`. `WEBHOOK_TARGET_SERVERS` is a comma-separated list of URLs that receive POSTs. If you change a filter to broaden indexing, expect proportionally higher webhook volume. ## Common operations ### Bring up ```bash # from the ar-io-node repo root docker compose --profile clickhouse up -d ``` The `--profile clickhouse` flag matters: without it, ClickHouse and the auto-import container stay down and the streaming pipeline goes silent. Other profiles (`grafana`, `bundler`, `litestream`, `ao`, `hb`) gate their own services. Wait for ready: ```bash until curl -sf http://localhost:4000/ar-io/healthcheck >/dev/null; do sleep 2; done ``` ### Rebuild from local source (after dev-branch code changes) ```bash docker compose --profile clickhouse build core clickhouse-auto-import docker compose --profile clickhouse up -d --force-recreate core clickhouse-auto-import ``` `--force-recreate` is required — without it, compose sees the same image *name* (`:latest`) and skips recreation even though the image *digest* changed. Confirm with `docker inspect --format '{{.Image}}'` — the SHA must differ from before. ### Disk pressure recovery (in order of safety) ```bash docker system df # see what's hogging docker builder prune -af # build cache (almost always biggest) docker container prune -f # stopped containers docker image prune -af # untagged images sudo journalctl --vacuum-size=500M # if /var/log/journal is large ``` Avoid `docker volume prune -af` without listing first (`docker volume ls -f dangling=true`) — there can be old SQLite/CH data in dangling volumes worth keeping. ### Block / unblock content (gateway admin API) ```bash ADMIN=$(grep '^ADMIN_API_KEY=' .env | cut -d= -f2-) TX= # Block (the gateway has PUT only — there is no DELETE handler today) curl -sf -X PUT "http://localhost:4000/ar-io/admin/block-data" \ -H "Authorization: Bearer $ADMIN" -H 'Content-Type: application/json' \ -d "{\"id\":\"$TX\",\"hash\":\"$TX\",\"source\":\"manual\",\"notes\":\"...\"}" ``` If a moderation sidecar is present (and wired via `WEBHOOK_TARGET_SERVERS`), prefer routing the block through it — depending on the sidecar's config, it may also bust an edge cache after blocking. ### HTTPSig overhead (for self-measurement) Read the in-process histogram for the truest per-request signing cost — no benchmarking needed: ```bash curl -sf http://localhost:4000/ar-io/__gateway_metrics | grep '^httpsig_signing_duration' # mean per sign = sum / count (typically 70-120 µs; under 500 µs at p100) ``` Don't draw conclusions from `/ar-io/info` or `/ar-io/healthcheck` benchmarks — those carry no trust-trigger headers and are never signed. ## Pitfalls (apply to any deployment) 1. **`clickhouse-auto-import` logs `Failed to connect to core port 4000 after 0 ms`** — the script's connection state goes stale when core is recreated mid-cycle, and auto-import sleeps 1 h between cycles, so the bad state lingers. Recreating just the auto-import container clears it; manual `curl http://core:4000/...` from inside the container will succeed even while the script is failing. 2. **Streaming tables empty (`new_blocks` / `new_transactions = 0`) despite `CLICKHOUSE_STREAMING_ENABLED=true`** — running the published `:latest` core image instead of a local build. The streaming columns/tables are populated only by the dev-branch code. Rebuild + force-recreate. 3. **Webhook error flood (`getaddrinfo ENOTFOUND `)** — a configured sidecar isn't running. Bring it up. Don't strip the target unless intentionally retiring the sidecar — those targets are how moderation and provenance subsystems work. 4. **Bundle unbundling looks idle (low matched count)** — `ANS104_UNBUNDLE_FILTER` is intentionally narrow on most operator configs (e.g., one App-Name + one owner). Confirm via `/ar-io/info`. Not a bug. 5. **Gateway has no `DELETE /ar-io/admin/block-data` handler** — only PUT. Sidecars that send DELETE (e.g. an unblock action) will see failures. Document but don't try to "fix" by changing the sidecar — the gateway is the side that needs the route. 6. **`/ar-io/info` and `/ar-io/healthcheck` are unsigned** — they carry no trust-trigger headers, so HTTPSig signing skips them. Use `/raw/` against a cached tx for HTTPSig benchmarks. 7. **`docker logs | tail` can dump megabytes of binary** — the streaming pipeline sometimes logs cert chains as raw byte arrays. Filter with `grep -viE '"buffer":|raw bytes|^[\[0-9, ]+$'`. 8. **Five-edit dance for adding a database method** — SQL statement, worker impl in `StandaloneSqlite`, queue wrapper in main DB class, case in worker message handler, interface in `types.d.ts`. Skip any one and you get silent failures. 9. **Observer restart-loops with `Epoch 0 PDA not found at — has prescribe_epoch run yet?`** — the network you're pointed at hasn't had its first epoch initialized. The observer can't bootstrap without entropy from `epoch[N].prescribed_observers`, and that PDA only exists after a cranker has called `create_epoch`. Confirm with `ar.io get-current-epoch ...` — `"Epoch 0 not found"` means network-side, not gateway-side. Wait for an active cranker (yours or another operator's) to bootstrap epoch 0. 10. **Cranker started cleanly but never logs any `[crank:*]` activity** — `EpochSettings.enabled === false` on the target network. The cranker bails silently at `epoch-cranker.ts:173` (debug log only) so it doesn't burn SOL submitting ix against a paused network. Verify with `ar.io get-epoch-settings ...`. No fix on the gateway side; this is a network-operations state. 11. **ArNS names return 404 even though `ar.io get-arns-record --name ` returns the record fine** — SDK version drift. `ArNSNamesCache.hydrate` paginates the on-chain registry; an SDK pin that's significantly older than the deployed `ario-arns` program may parse paginated responses incorrectly, hydrating "successfully" but with most entries missing. Logs show `Successfully hydrated ArNS names cache` quickly (~few seconds for thousands of records) followed by `Base name not found in ArNS names cache` on lookups. Fix: bump `@ar.io/sdk` in `package.json` to the latest `^4.0.0-solana.*` and rebuild. 12. **`/ar-io/info` shows the OLD network's `programIds` after a migration / image swap** — the `.env` has explicit `CORE_IMAGE_TAG` / `ENVOY_IMAGE_TAG` / `OBSERVER_IMAGE_TAG` pins that shadow the compose defaults. Updating only the compose defaults won't move a gateway whose `.env` still pins old image tags. Either update both, or remove the explicit pins from `.env` and let compose defaults win. 13. **Index bands installed but lookups still go to the network** — `ROOT_TX_LOOKUP_ORDER` must put `cdb` right after `db`, and `CDB64_ROOT_TX_INDEX_SOURCES` must start with `data/indexes/installed/root-tx-index`; both need a gateway restart. Bring them up by service name: the sidecar alone with `docker compose --profile index-swarm up -d index-swarm`, and with the torrent engine `docker compose --profile index-swarm --profile index-swarm-torrent up -d index-swarm index-swarm-engine` (its init runs first by dependency). A bare `up` with a profile starts every default service too. 14. **Every visitor shares one rate-limit bucket, or all get 429/402 together** — the gateway believes `X-Forwarded-For` only from `TRUSTED_PROXIES` (default: loopback and private ranges). Behind a CDN or a load balancer on public addresses, the CDN's own address is then taken as the client. Add its published ranges to `TRUSTED_PROXIES` (keep the private defaults if nginx sits in between) and recreate core. Never set `TRUSTED_PROXIES=none` with the standard compose: every request through Envoy (port 3000) would carry Envoy's address and share one bucket. `none` is only for a core that clients reach directly. Allowlists match the client address only, never other addresses in its headers. 15. **A range request on an index band file returns `200` with the whole file, and the next one returns `402`** - a proxy in front has a `proxy_cache` zone covering `/ar-io/indexes`. When a location has a cache zone, nginx strips the client's `Range` from the upstream request and fetches the whole object so it can store it, which happens before it sees any `Cache-Control`, so no response header can prevent it and `proxy_buffering off` alone does not fix it. Whether the *client* sees the full body depends on whether the response was cacheable: a metering gateway sends `private`, so nothing is stored and nginx returns `200` with everything on every request, while a gateway that does not meter sends `public` and nginx answers the range from cache, appearing to work while still pulling each whole file once and filling the zone with band files of up to 1.35 GB. Use `proxy_cache off` for the prefix (no trailing slash: with one, nginx answers the unslashed form with a `301`, and the document at `/ar-io/indexes` is what subscribers poll, with HTTPSig signing `@path`). Core and envoy both answer `206` correctly, so the default stack is fine and this only bites once an operator adds a cache or CDN. It fails silently, since a `200` with more bytes than asked for looks like success, and it breaks BitTorrent WebSeeds, media seeking and clients that query the Parquet bands in place over HTTP. **Check every server block that serves the gateway**: a TLS listener and an internal cache listener are often in different files, and fixing one leaves the other broken. After `systemctl reload nginx`, old workers finish their requests on the old config, so re-test a few seconds later before concluding a fix failed. The following `402` is metering, not misconfiguration: the accidental whole-file transfer consumed the address's allowance. ## When something breaks: where to look | Symptom | First check | |---|---| | 5xx from gateway | `docker logs --tail 50 ` (filtered), then envoy logs | | Indexer falling behind | SQLite `MAX(height)` per `block_importer` log, peer health, `/data` free space | | ClickHouse stale | `docker compose --profile clickhouse logs --tail 50 clickhouse-auto-import`; verify curl-to-core works from inside the container | | Streaming table stuck at 0 | image SHA — almost always the published `:latest` rather than local build | | Admin API 401/404 | `ADMIN_API_KEY` mismatch with `.env` | | Webhook silent | `WEBHOOK_TARGET_SERVERS` in gateway `.env`; sidecars must be on the same docker network and reachable by service name | | `/data` filling | `du -h -d 1 /data \| sort -hr` — usually the `contiguous` cache | | Observer not submitting `save_observations` | `docker logs ar-io-node-observer-1 --since 10m \| grep -iE 'crank\|epoch'`. Combine with `ar.io get-current-epoch ...` to distinguish gateway-side issues from a paused/uninitialized network | | ArNS resolution returning 404 for names that demonstrably exist | Check the SDK pin in the container: `docker exec ar-io-node-core-1 /nodejs/bin/node -e "console.log(require('@ar.io/sdk/package.json').version)"`. Drift behind the deployed program causes silent partial cache hydration | | `/ar-io/info` shows wrong `programIds` after migration | `.env` likely has explicit `*_IMAGE_TAG` lines pinning the old images. Update them or remove to let compose defaults take over | ## Metrics worth watching Quick-scan list of the prometheus metrics that surface most operator-grade issues. All exposed at `/ar-io/__gateway_metrics` on core's HTTP port. ### Pipeline throughput + saturation | Metric | Surfaces | |---|---| | `bundles_unbundled_total` (rate) | bundle-level throughput | | `data_items_indexed_total{parent_type="transaction"}` (rate) | data-item write rate (SQLite writer capacity) | | `bundles_unbundle_in_flight` | parser workers currently active — if pinned at worker count for minutes, suspect cascade hangs | | `queue_length{queue_name="…"}` | the 4-queue chain — see "ANS-104 unbundling pipeline" | | `data_items_dropped_total{queue_name="…"}` | **silent data loss** at indexer hard cap. Any nonzero rate is a config issue, not a transient | | `bundles_unbundle_skipped_total{reason="…"}` | `high_queue_depth` (soft gate engaged), `queue_full` (BRW pushing past cap), `no_workers` (config error) | ### Per-source health | Metric | Surfaces | |---|---| | `get_data_stream_successes_total{class,source}` | which cascade tier / which peer actually serves bytes | | `get_data_stream_errors_total{class,source}` | tier/peer failure rates. Persistent 100 % errors on a peer mean it's dead but the network process still advertises it | | `get_data_errors_total{…}` | pre-stream errors (connection refused, 4xx, 5xx) | ### Sign of saturation, not failure | Metric | What it means | |---|---| | `nodejs_event_loop_utilization` | 0–1, sampled at scrape. Near 1.0 = main thread pegged | | `nodejs_eventloop_lag_p99_seconds` | tail latency on the event loop. > 0.1 s under load = something blocking on the main thread | | `process_resident_memory_bytes` | overall memory growth; sudden steps usually map to a queue filling | | `httpsig_signing_duration_seconds_*` | per-response signing cost (the histogram is more honest than benchmarking) | ### Performance-interpretation guidance - **Sample windows lie.** A 60–90 s rate can show 0/hr or 60K/hr depending on whether you caught the soft-gate cycle in its open or closed half. Trust averages over many minutes, not single snapshots. - **Saturated workers ≠ broken workers.** `in_flight = N` for the configured worker count with throughput still climbing is the healthy steady state; only worry if the corresponding `_total` counter is *also* flat for many minutes. - **The cascade goes through phases.** When trusted-gateway A goes hot-cache for an hour, throughput spikes; when it rotates out, throughput dips. Aggregate over a day, not a window. ## Adjacent skills (use these, don't reinvent) - **`release`** — handles version bumps, image SHA pinning, tags, GitHub releases. - **`testing`** — picks the right test layer (unit / property / e2e / auto-verify / parquet integration / load test) for a given change. - Deployment-overlay skill (e.g. `vilenarios-gateway`) — adds the local-machine specifics (domains, paths, sidecar config) on top of this generic skill. This skill covers the **operations** layer: what's running, where, and how to keep it that way. For code-level guidance see `CLAUDE.md` in the relevant repo.