--- name: deploy description: > Deploy or run the ScienceDiscovery stack, either as host processes (scripts/start-stack.sh) or with Docker Compose. Use when the user asks to deploy, start, run, or containerize ScienceDiscovery, or asks for the UI URL, SSH port forwarding, or start/stop/restart/log commands. --- # Deploy ScienceDiscovery (local or Docker) Project-local skill for **ScienceDiscovery**. Authoritative facts live in [`README.md`](../../../README.md) → **Quick start**, in [`docs/zh/getting-started/deployment.md`](../../../docs/zh/getting-started/deployment.md) (Chinese) → local mode and Docker deployment, and in [`docs/zh/reference/configuration.md`](../../../docs/zh/reference/configuration.md) → variable tables. Read the matching section before improvising; never invent ports, tokens, flags, or paths. You **assist** the user through a deployment. You do not silently reconfigure their machine. ## Rules 1. **Ask the deployment form first** (Docker Compose vs. host processes). Do not install, build, or start anything before the user answers. 2. **Detect before deploying.** Environment checks are read-only commands only. Report what is missing; do not fix it on your own initiative. 3. **Never run `sudo`** on the user's behalf unless they explicitly ask for that specific command in this conversation. 4. **Never install global software** (`apt`/`dnf`/`brew`/`npm -g`, persistent `sysctl`, systemd units, editing files outside the repository). List the gap and the command *you suggest the user run themselves*. After written consent you may assist, still preferring project-local installs (`uv`/`pnpm`/`corepack`) over system-wide ones. 5. **Environment must pass before deploying.** On any failed check, stop, report, and wait for the user's decision. 6. Repository-local writes are fine (`.env`, `.sciencediscovery-data/`), but say what you are writing before you write it. 7. After a successful start, always report the **URL and the bearer token source**, then the management commands for the mode actually used. ## Step 1 — Ask the form > Docker Compose, or host processes? | | Docker Compose | Host processes | |---|---|---| | Host needs | Docker Engine 24+ and Compose v2 only | Node 22.19+, pnpm, Python 3, uv, bubblewrap | | Entry point | `docker compose up -d` → `scripts/start-stack.sh --mode docker` | `./scripts/start-stack.sh --mode local` | | Runs | Detached, restarts on failure | Foreground terminal (Ctrl-C stops it) | | State | Host bind mount `./data` (no named volume) | `./data` (`SCIENCE_AGENT_DATA_DIR`) | | Published on | `127.0.0.1:4310` by default | `0.0.0.0:4310` by default | Both need Linux x86_64 or aarch64 with unprivileged user namespaces — the bubblewrap sandbox is not optional and Docker Desktop on macOS/Windows is unsupported. ## Step 2 — Detect (read-only) Run from the repository root. Report each check as pass/fail with the value seen. **Common** ```bash uname -s -m # expect: Linux x86_64 or aarch64 ss -ltn | grep -E ':(4310|4311)' || echo "ports free" ``` Sandbox availability is decided by **running the product's probe**, never by reading a sysctl. Use the probe below for host-process mode, and the in-container probe under *Docker mode* for Compose. A sysctl value is a diagnostic to reach for after a probe fails — see Step 3. **Host-process mode** (mirrors `README.md` → Quick start → Requirements) ```bash node --version # v22.19+ pnpm --version # 11.1.2 python3 --version uv --version # 0.9+ bwrap --version # 0.6+; 0.8+ recommended (adds --disable-userns where the environment allows it; otherwise the runner logs a startup warning and omits it) curl --version | head -1 ``` Sandbox preflight — the same probe `scripts/start-stack.sh` and `packages/sandbox-capability` run, harmless and read-only: ```bash bwrap --unshare-all --unshare-user --die-with-parent \ --ro-bind /usr /usr --symlink usr/bin /bin --symlink usr/lib /lib \ --symlink usr/lib64 /lib64 --proc /proc /usr/bin/true && echo "sandbox ok" ``` **Docker mode** ```bash docker version --format '{{.Server.Version}}' # 24+ docker compose version # v2.x docker info >/dev/null && echo "daemon reachable" id -u; id -g # → SCIENCE_AGENT_UID / _GID ``` The container's sandbox can only be probed once the image exists, and the entry point already probes it at startup: a failure is printed to `docker compose logs`. To confirm it positively after `up -d`: ```bash docker compose exec sciencediscovery sh -c ' bwrap --unshare-all --unshare-user --die-with-parent \ --ro-bind /usr /usr --symlink usr/bin /bin --symlink usr/lib /lib \ --symlink usr/lib64 /lib64 --proc /proc /usr/bin/true' && echo "sandbox ok" ``` Keep the outer `sh -c`. With `bwrap` as the exec session's first process the loopback setup fails for reasons unrelated to sandbox capability, which reads as a false negative. ## Step 3 — Report gaps, do not close them yourself Present a short table: check / expected / actual / suggested fix. Suggested fixes are **for the user to run**, quoted as such: - Missing Node/pnpm/uv → suggest user-level installs that need no sudo and touch nothing outside `$HOME`: Node 22 tarball extracted to `~/opt/node22` (add its `bin` to PATH), `corepack enable --install-directory ~/.local/bin && corepack prepare $(node -p "require('./package.json').packageManager") --activate` for the pinned pnpm, and the official uv installer script (installs to `~/.local/bin`). Missing bubblewrap has no user-level path — it is a distro package the user must install. - Sandbox probe **passed** → the environment is fine. Do not read `kernel.apparmor_restrict_unprivileged_userns`, do not report it, and never ask for `sudo sysctl -w ...=0`. That restriction is configured per AppArmor profile, so the value is routinely `1` on Ubuntu 24.04+ hosts where the probe passes; quoting a root command there is wrong advice. - Sandbox probe **failed** → report it and work through the causes in order: the container's `security_opt` (Docker mode needs `seccomp`, `apparmor` and `systempaths` all unconfined); then the host AppArmor configuration, where `/etc/apparmor.d/` can grant `userns create` per program; and only then `sysctl kernel.unprivileged_userns_clone` and `sysctl kernel.apparmor_restrict_unprivileged_userns`. If the last one is the confirmed cause, quote `sudo sysctl -w kernel.apparmor_restrict_unprivileged_userns=0`, say it needs root and does not persist, and stop for the user's decision. - Port already in use → offer `SCIENCE_AGENT_PORT` (host mode) or `SCIENCE_AGENT_PUBLISH_PORT` (Docker) instead of killing the other process. When moving the runner port, also update the explicit `SCIENCE_AGENT_RUNNER_URL` in `.env` — it does not follow `*_PORT` automatically. - npm / PyPI unreachable or very slow → set `SCIENCE_AGENT_NPM_REGISTRY` / `SCIENCE_AGENT_PYPI_INDEX` in `.env` (e.g. the Huawei Cloud mirrors; see the variable table in `docs/zh/reference/configuration.md`). Both are scoped to `start-stack.sh`'s install commands and never touch user or global npm/uv config. If the sandbox probe keeps failing, the API and UI still start and `/health` still reports the runner, but every `run_python` / `run_shell` fails. Say this plainly and let the user choose whether to continue. ## Step 4a — Deploy: host processes ```bash cp .env.example .env # first time only; never overwrite an existing .env ./scripts/start-stack.sh --mode local # install + build + start ./scripts/start-stack.sh --mode local --no-build # start only, after a previous build ``` - Runs in the **foreground**; it starts the runner (127.0.0.1:4311) and the memory-graph sidecar (127.0.0.1:17674) in the background and the API (0.0.0.0:4310) in front. It also starts the bundled Python MCP sidecars. - Backgrounding it (`nohup ./scripts/start-stack.sh --mode local &`, tmux, a supervisor) changes how it must be **stopped** — see *Stopping a backgrounded host stack* below. Ctrl-C no longer applies. - First run provisions `.sciencediscovery-data/envs/gateway` and `.sciencediscovery-data/envs/paper` through `uv` and needs outbound network. - `./scripts/run-local.sh [--no-build]` is the compatibility wrapper; `pnpm start` and `pnpm server` go through it. - Never launch it with `sudo`. For unattended operation, hand the user a supervisor option (tmux, or a systemd **user** unit) — do not install one. ## Step 4b — Deploy: Docker Compose ```bash cp .env.docker.example .env # or merge its keys into an existing .env # set SCIENCE_AGENT_UID / SCIENCE_AGENT_GID to `id -u` / `id -g` when they are not 1000 mkdir -p data # before `up`: Docker creates a missing bind-mount source as root docker compose build # first build is long and needs network docker compose up -d curl -fsS http://127.0.0.1:4310/health ``` - `./data` is bind-mounted to `/app/data` and is the only persisted location; it survives `docker compose down` and rebuilds. - A uid/gid mismatch is the most common first-run failure: the entry point exits with an explicit "not writable" message. Fix it via `SCIENCE_AGENT_UID` / `SCIENCE_AGENT_GID` and recreate the container. The same message appears when the data directory did not exist at `up` time: Docker then created it as root, so remove it and `mkdir -p` it as the user first. - The sign-in URL in `docker compose logs` always names the container port 4310. When `SCIENCE_AGENT_PUBLISH_PORT` differs, tell the user to substitute the published port before opening it. - The first start seeds micromamba into `./data/scientific-envs/bin/` and then creates the starter Python environment in the background (conda-forge, about 2 GB into `./data`). The UI works meanwhile; `runner.scientificEnvs.startersReady` in `/health` turns `true` when it is done. `SCIENTIFIC_ENVS=0` skips it. - Several instances on one host: give each its own Compose project name, port and data directory. The service sets no `container_name`, so the project name alone separates container and network names. ```bash mkdir -p data-b COMPOSE_PROJECT_NAME=sciencediscovery-b SCIENCE_AGENT_PUBLISH_PORT=4320 \ SCIENCE_AGENT_DATA_HOST_DIR=./data-b docker compose up -d ``` Every later `ps` / `logs` / `down` needs the same project name, or it acts on the other instance. Add `SCIENCE_AGENT_IMAGE` when the instances are built from different checkouts. - The service already sets `seccomp=unconfined`, `apparmor=unconfined` and `systempaths=unconfined` for bubblewrap. Do not add capabilities or `privileged: true`. Dropping `systempaths=unconfined` still works, but the runner then falls back to binding the container's `/proc` and warns at startup: sandboxed code sees the container's process list. ## Step 5 — Report the URL - Default: - Sign in with `SCIENCE_AGENT_AUTH_TOKEN`. There is no shipped default: when the variable is unset, the stack generates a token on its first start, prints it at startup, and stores it in `/secrets/auth-token`. Read the value from the user's `.env` or that file — do not print a token into a shared channel. - Host-process mode binds all interfaces by default, so `http://:4310` also works. Docker publishes on `127.0.0.1` unless `SCIENCE_AGENT_PUBLISH_HOST` is changed, and the `Open to sign in` URL in its log always says port 4310 — report the published port instead when it differs. Auth is one bearer token with no TLS: recommend loopback plus SSH forwarding over exposing the port, and tell the user to change the token first if they do expose it. - There is no built-in model. A usable session needs a profile plus credential under **System configuration → Model registry** (`docs/zh/reference/runtime-behavior.md` → 模型). ## Remote host, browser on a laptop When the stack runs on a remote development machine, forward the port instead of publishing it. Run from the **laptop**: ```bash remote_user="alice"; remote_host="science-host.example" ssh -N -L 4310:127.0.0.1:4310 "${remote_user}@${remote_host}" local_port=4310; remote_port=4310 ssh -f -N -o ServerAliveInterval=30 \ -L "${local_port}:127.0.0.1:${remote_port}" "${remote_user}@${remote_host}" ``` Then open `http://127.0.0.1:` locally. `` is the published port: `SCIENCE_AGENT_PORT` in host-process mode, `SCIENCE_AGENT_PUBLISH_PORT` in Docker mode. Pick a different `` if 4310 is taken locally. Stop the tunnel by closing the session, or by killing the backgrounded `ssh -f` process. Forward only the API port. The runner (4311) is loopback-only by design. ## Management commands | Action | Host processes | Docker Compose | |---|---|---| | Start | `./scripts/start-stack.sh --mode local [--no-build]` | `docker compose up -d` | | Stop (foreground) | Ctrl-C in that terminal (also stops the runner) | `docker compose down` (`./data` survives) | | Stop (backgrounded) | signal the instance's process group — see below | same | | Restart | stop as above, then start again | `docker compose restart` | | Rebuild + restart | start once without `--no-build` | `docker compose up -d --build` | | Logs | stdout/stderr of the foreground terminal (or the supervisor's log) | `docker compose logs -f` | | State | `ss -ltn \| grep 4310` | `docker compose ps` (includes the health check) | | Health | `curl -fsS http://127.0.0.1:4310/health` | same, against the published host/port | Component health endpoint in both modes: runner `http://127.0.0.1:4311/health` (reachable from inside the container in Docker mode). The API's own `/health` echoes the runner status, so it is usually the only one worth checking. ## Stopping a backgrounded host stack Ctrl-C only reaches a **foreground** stack. Started with `nohup ... &`, tmux or any supervisor, the script sits in its own process group and is reparented to init, so the terminal's SIGINT never reaches it and its services keep running. Signalling the script alone is also not enough: `start-stack.sh` runs the API in its own foreground and its `trap cleanup EXIT INT TERM` only fires once that API exits. `kill -TERM ` therefore leaves the API, the runner, the memory-graph sidecar and the bundled MCP sidecars running and the ports bound. Stop the whole **process group** instead, and identify it from the port this instance published — never from the script name, because a developer host often runs several stacks on different ports at once: ```bash api_port=4310 # this instance's SCIENCE_AGENT_PORT api_pid=$(ss -ltnp | grep ":${api_port} " | grep -o 'pid=[0-9]*' | head -1 | cut -d= -f2) pgid=$(ps -o pgid= -p "$api_pid" | tr -d ' ') pgrep -g "$pgid" -a # review the group before signalling it kill -TERM -"$pgid" for _ in $(seq 1 30); do ss -ltn | grep -qE ":(${api_port}|4311|17674)\b" || break; sleep 0.5; done ss -ltn | grep -E ":(${api_port}|4311|17674)\b" || echo "ports released" pgrep -g "$pgid" -a || echo "no survivors" ``` List the group with `pgrep -g`, not `ps -g`: `ps` reads a numeric `-g` argument as a session id, so it silently prints nothing and the group looks empty. - Escalate to `kill -KILL -"$pgid"` only for what survives, and only after the 30-second wait: SIGTERM lets the runner tear its sandboxes down. - **Never** `pkill -f start-stack.sh`, `pkill -f 'node.*server.js'` or `killall node`. Those match every stack and every unrelated Node process on the host; the port-derived process group matches exactly one instance. - Neighbour check afterwards: the other stacks' ports must still be listening and their `/health` must still answer. - `ss -ltnp` prints the PID only for sockets you own; for another user's stack ask that user rather than widening the match. - If the instance uses non-default ports, substitute its own `SCIENCE_AGENT_PORT`, `SCIENCE_AGENT_RUNNER_PORT` and `SCIENCE_AGENT_MEMORY_GRAPH_PORT` everywhere above. ## Troubleshooting | Symptom | Cause / next step | |---|---| | `data ... is not writable by uid ...` | Docker uid/gid mismatch → set `SCIENCE_AGENT_UID` / `_GID`, recreate | | `WARNING: bubblewrap cannot create a sandbox` in logs | The probe failed → work Step 3's order: `security_opt`, then host AppArmor profiles, then the sysctl. Never open with the `sysctl` fix | | `.sciencediscovery-data/envs/gateway is missing` | Started with `--no-build` before a build → run once without it (that venv holds the interpreter for the bundled Python MCP servers) | | API up but every run fails | Usually the sandbox warning above, or no model profile configured | | Port already bound (`Bind for 127.0.0.1:4310 failed: port is already allocated` in Docker mode) | Change `SCIENCE_AGENT_PORT` / `SCIENCE_AGENT_PUBLISH_PORT`; substitute the new port in the sign-in URL | | Sign-in URL from `docker compose logs` does not open | It names container port 4310; replace it with `SCIENCE_AGENT_PUBLISH_PORT` | | Second Docker instance exits with the "not writable" message | Its data directory was created by Docker as root → `mkdir -p` it as the user before `up` | | Model on the host is unreachable from the container | `127.0.0.1` inside the container is the container. Add `extra_hosts: ["host.docker.internal:host-gateway"]` to the service in a `docker-compose.override.yml` and use `host.docker.internal:`, or the host's LAN IP | | Model calls need a proxy (Docker mode) | System configuration → Network proxies with a `custom_url` record, or inject `HTTPS_PROXY` through `docker-compose.override.yml` and pick the `environment` proxy type. Compose forwards only the keys of `.env.docker.example` | | Web says "Local service access token rejected" | The pasted value is not this instance's token (a model API key, or another data directory's token) → use `./data/secrets/auth-token` | | Ctrl-C did not stop a host stack, ports still bound | It was backgrounded (`nohup`/tmux/supervisor), so SIGINT never reached it → stop its process group, see *Stopping a backgrounded host stack* | | Killed the script but the API/runner survived | `start-stack.sh` only runs its cleanup trap after its foreground API exits → signal the process group, not the script PID | | `does not support --disable-userns` warning at runner startup | Expected on bwrap < 0.8 (e.g. Ubuntu 22.04's 0.6): nested-userns hardening is skipped, everything else isolates normally. Upgrade bubblewrap for the stronger profile | | `supports --disable-userns but cannot use it here` warning at runner startup | Expected under LXC and container runtimes that mount `/proc/sys` read-only, so bubblewrap cannot write `user.max_user_namespaces`. The option is omitted and executions run; everything else isolates normally. Do not grant `privileged` or `systempaths=unconfined` to silence it | | First `docker compose build` fails on network | The build resolves pnpm and both uv environments; retry with network available | ## Checklist 1. Asked the user which form, and got an answer 2. Ran the read-only checks for that form; reported pass/fail values 3. Reported gaps with suggested user-run commands — no `sudo`, no global install 4. Deployed only after the user confirmed the environment is acceptable 5. Verified `/health`, then reported URL + token source 6. Gave SSH forwarding instructions when the stack is not on the user's own machine 7. Gave the management command table for the mode used, including the correct stop for a backgrounded host stack **Policy**: this skill assists; it does not change host configuration on its own. Repository-local writes need a heads-up, host-level changes need the user's explicit go-ahead, and root-level changes are always executed by the user.