---
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Troubleshooting"
sidebar-title: "Troubleshooting"
description: "Diagnose and resolve common NemoClaw installation, onboarding, and runtime issues."
description-agent: "Lists fixes for common installation, onboarding, and runtime issues. Use when diagnosing a reported NemoClaw error, a failed onboard, or unexpected sandbox behavior."
keywords:
[
"nemoclaw troubleshooting",
"nemoclaw debug sandbox issues",
"openclaw tool calling",
"raw tool call json",
]
content:
type: "reference"
---
{/* markdownlint-disable MD014 */}
This page covers common installation, onboarding, and runtime issues, along with resolution steps.
The diagnostic commands on this page assume `$$nemoclaw` is on your `PATH`
(re-source your shell profile after an nvm- or fnm-managed install) and that
your user can reach the Docker socket — either as a member of the `docker`
group or by running the Docker commands with `sudo`.
If your issue is not listed here, join the [NemoClaw Discord channel](https://discord.gg/XFpfPv9Uvx) to ask questions and get help from the community. You can also [file an issue on GitHub](https://github.com/NVIDIA/NemoClaw/issues/new).
## Installation
### `$$nemoclaw` not found after install
If you use nvm or fnm to manage Node.js, the installer may not update your current shell's PATH. The `$$nemoclaw` binary is installed but the shell session does not know where to find it.
Run `source ~/.bashrc` (or `source ~/.zshrc` for zsh), or open a new terminal window.
When installing from a source checkout with `npm install`, NemoClaw first tries `npm link`. If the global npm prefix is not writable, it writes a managed shim to `~/.local/bin/nemoclaw` instead. Add `~/.local/bin` to your `PATH` if the command is still not found. Source-checkout installs also bootstrap OpenShell when it is missing before running preflight. If a source install still reports that `openshell` is not available, re-run the installer from the repository root and check that `~/.local/bin` is on your `PATH`.
### Installer Reports an Unfinished Installation
If `$$nemoclaw` reports that its compiled CLI is missing or incomplete, an earlier install or upgrade did not finish.
Rerun the installer command that you used to install NemoClaw.
The installer replaces the incomplete CLI, then attempts to recover existing sandboxes before onboarding.
Follow any recovery guidance that the installer prints, then rerun `$$nemoclaw`.
### Installer fails on unsupported platform
The installer checks for a supported OS and architecture before proceeding. If you see an unsupported platform error, verify that you are running on a tested platform listed in the Container Runtimes table in the quickstart guide.
### Node.js version is too old
NemoClaw requires Node.js 22.19 or later. If the installer exits with a Node.js version error, check your current version:
```bash
node --version
```
If the version is below 22.19, install a supported release. If you use nvm, run:
```bash
nvm install 22
nvm use 22
```
Then re-run the installer.
### Installer Reports `No SHA-256 tool available (sha256sum/shasum)`
The installer must verify the nvm installer before it installs or upgrades Node.js. It exits before running the downloaded script when neither `sha256sum` nor `shasum` is available.
On Debian or Ubuntu, install `sha256sum` through `coreutils`:
```bash
sudo apt-get update
sudo apt-get install -y coreutils
```
On another Linux distribution, install its `coreutils` package. On macOS, `/usr/bin/shasum` is normally present; restore it through the operating system if it is missing.
Verify that one supported tool is available, then rerun the installer:
```bash
command -v sha256sum || command -v shasum
```
### Contributor Setup Fails with a JavaScript Heap Out-of-Memory Error
This applies to a source checkout, not to an installed release.
Node.js derives its default old-space limit from host memory. On a host with 8 GB of RAM, that limit is about 2.2 GB. The CLI type check needs more heap than that limit, so `./scripts/dev-setup.sh` stops at the type-check step and Node.js reports `JavaScript heap out of memory`.
Raise the limit, then run setup again:
```bash
export NODE_OPTIONS=--max-old-space-size=5120
./scripts/dev-setup.sh
```
Keep that variable set for later type-check, build, and test commands.
### Image push fails with out-of-memory errors
The sandbox image is approximately 2.4 GB compressed. During image push, the selected container runtime, k3s, and the OpenShell gateway run alongside the export pipeline, which buffers decompressed layers in memory. On machines with less than 8 GB of RAM, this combined usage can trigger the OOM killer.
If you cannot add memory, configure at least 8 GB of swap to work around the issue at the cost of slower performance.
### Docker is not running
These steps apply to the default Docker path. Native rootless Podman does not require Docker; retain `NEMOCLAW_GATEWAY_RUNTIME=podman` and follow [Podman](#podman).
Check the host before onboarding:
```bash
$$nemoclaw host probe
```
The command does not start Docker or apply a repair. A `host.docker.daemon_unreachable` finding means Docker is installed but NemoClaw cannot reach the daemon. For JSON output and exit-code details, refer to [System Readiness](system-readiness).
The installer and onboard wizard require Docker to be running on the default Docker path. If you see a Docker connection error on that path, start the Docker daemon:
```bash
sudo systemctl start docker
```
On macOS with Docker Desktop, open the Docker Desktop application and wait for it to finish starting before retrying.
### Docker permission denied on Linux
On Linux, if the Docker daemon is running but you see "permission denied" errors, your user may not be in the `docker` group. The installer can add your user to the group, but Linux does not activate that membership in the current shell automatically. Add your user and activate the group in the current shell:
The default Docker path needs Docker access. On personal Linux development
machines, adding your user to the `docker` group is the standard way to run
Docker without sudo. Members of the `docker` group can control the daemon with
root-level impact, so grant this access only to trusted local accounts; on
shared or managed systems, use your organization's approved Docker access
path. For background, review Docker's [daemon attack surface
guidance](https://docs.docker.com/engine/security/#docker-daemon-attack-surface).
```bash
sudo usermod -aG docker $USER
newgrp docker
```
Then retry `$$nemoclaw onboard`. If the installer stopped after printing `newgrp docker`, run that command and then re-run the installer:
```bash
newgrp docker
curl -fsSL https://www.nvidia.com/nemoclaw.sh | bash
```
### Installer reports Docker access outside the docker group
On Linux, the installer may report that Docker is reachable even though your user is not in the `docker` group. This means the host grants Docker daemon access through another path, such as a custom `DOCKER_HOST`, socket ACL, or managed runtime policy. NemoClaw can continue when `docker info` works, but the diagnostic explains why a negative Docker-permission test will not reproduce on that host.
Check the Docker access path before relying on the host as a clean permission baseline:
```bash
id -nG
echo "${DOCKER_HOST:-}"
docker info
```
The managed default gateway service accepts `DOCKER_HOST` only as an absolute local `unix://` socket path. It rejects remote endpoints and relative socket paths before service startup.
### Onboarding Reports an Invalid Docker Host
The `invalid_docker_host` advisory means that the selected Docker endpoint is not an absolute local `unix://` socket path that NemoClaw can write to the managed OpenShell gateway service environment.
Either `DOCKER_HOST` or `DOCKER_CONTEXT` can select that endpoint, and the advisory title names the variable that did.
A `DOCKER_CONTEXT` selection reports this instead of falling back to Docker's default socket, because that fallback would report readiness for a daemon you did not select. NemoClaw does not use the standalone gateway fallback when this validation fails. Onboarding prints the advisory identifier in parentheses after each action title in the `Suggested fix` list. The terminal output names `invalid_docker_host` when this validation fails, so you can match the message to this section. Remove the selector to use Docker's default local endpoint. Clear both variables: unsetting only `DOCKER_HOST` leaves a `DOCKER_CONTEXT` selection active, and the next run reports the same advisory.
```bash
unset DOCKER_HOST DOCKER_CONTEXT
$$nemoclaw onboard
```
If your Docker daemon uses another local socket, set an absolute `unix://` path before you retry:
```bash
export DOCKER_HOST=unix:///var/run/docker.sock
$$nemoclaw onboard
```
NemoClaw rejects TCP and SSH endpoints, relative socket paths, values that contain single quotes, and values that contain line breaks. Do not wrap the socket path in single quotes inside the variable value.
### Onboarding Reports a Missing Docker Endpoint Socket
The `docker_endpoint_socket_missing` advisory means that you selected a Docker endpoint with no Unix socket at the path it names.
Either `DOCKER_HOST` or a `DOCKER_CONTEXT` selection can name that endpoint.
A missing path and a regular file or directory at that path are both reported here: nothing can listen on either, so NemoClaw withholds the docker-group and start-Docker remedies instead of suggesting a cause it has ruled out.
The advisory prints the selected endpoint and an `ls -l` command for its exact path. Inspect that path, then either point the selector at a socket that exists or clear it to use Docker's default endpoint:
```bash
ls -l
unset DOCKER_HOST DOCKER_CONTEXT
$$nemoclaw onboard
```
This advisory never applies to Docker's default endpoint. A missing default socket means the daemon is not running, which `start_docker` reports instead.
### Onboarding Warns About the Docker Desktop Credential Store in a Headless Session
The `docker_desktop_credential_store_headless` advisory means that the Docker client config sets `credsStore` to `desktop` (macOS) or `desktop.exe` (WSL2) and the session looks headless, for example an SSH session without a GUI. The Docker Desktop credential helper needs an interactive GUI session. Without one, the helper can fail and block every image pull, even for public images. A common failure message from Docker in this state is `A specified logon session does not exist`.
Onboarding preflight prints this warning before the first image pull and then continues, on both fresh and resumed onboarding. NemoClaw reads the config from `$DOCKER_CONFIG/config.json` when `DOCKER_CONFIG` is set, and from `~/.docker/config.json` otherwise. On WSL, NemoClaw probes the credential helper with a read-only `list` call instead of relying on session markers, because WSLg can set `DISPLAY` in every WSL shell.
NemoClaw reads `DOCKER_HOST`, `DOCKER_CONTEXT`, and `DOCKER_CONFIG` from the environment of the `$$nemoclaw` command.
On WSL, NemoClaw automatically uses a temporary credential-free Docker configuration for generated image builds and local-inference probe images when all of these conditions apply:
- `DOCKER_HOST` is unset and the effective Docker context is `default`.
- The client configuration named above selects the Docker Desktop credential helper.
- The read-only helper probe fails.
- The operation does not require registry credentials.
NemoClaw does not modify your Docker client configuration.
It removes the temporary directory after the Docker operation.
If cleanup fails, the warning prints the exact credential-free directory; wait for Docker to finish using it, then remove that directory.
These cases continue to use the configured Docker client state:
- A nonblank `DOCKER_HOST`.
- A non-default context selected with `DOCKER_CONTEXT` or persisted by `docker context use`.
- A responsive helper.
- A custom Dockerfile.
- An operation that might require private-registry credentials.
Restore the Docker Desktop session or helper access for those operations.
The OpenShell gateway pulls a sandbox image, including a managed sandbox image, through the Docker socket configured for the gateway.
It does not read your Docker client configuration, so the credential helper does not affect that pull.
For a public-image-only retry outside the automatic WSL path, resume onboarding with a temporary isolated configuration:
```bash
(
set -eu
temporary_docker_config="$(mktemp -d)"
trap 'rm -rf -- "$temporary_docker_config"' EXIT
DOCKER_CONFIG="$temporary_docker_config" $$nemoclaw onboard --resume
)
```
Alternatively, temporarily remove the `credsStore` entry from the Docker client config named above, then rerun `$$nemoclaw onboard`. Restore the entry afterward if you use registries that need stored credentials in GUI sessions.
### macOS first-run failures
The two most common first-run failures on macOS are missing developer tools and Docker connection errors.
To avoid these issues, install the prerequisites in the following order before running the NemoClaw installer:
1. Install Xcode Command Line Tools (`xcode-select --install`). These are needed by the installer and Node.js toolchain.
2. Install and start a supported container runtime (Docker Desktop or Colima). Without a running runtime, the installer cannot connect to Docker.
### `docker` is missing after installing Colima
Homebrew Colima does not install the Docker CLI binary. If you install only Colima, `colima start` can succeed while later `docker` commands fail with `command not found`.
Install both packages, start Colima with enough resources for the sandbox image build, and verify Docker before onboarding:
```bash
brew install colima docker
colima start --cpu 4 --memory 8
docker info
```
### Permission errors during installation
The NemoClaw installer does not require `sudo` or root. It installs Node.js via nvm and NemoClaw via npm, both into user-local directories. The installer also handles OpenShell installation automatically using a pinned release.
If you see permission errors during installation, they typically come from Docker, not the NemoClaw installer itself. Docker must be installed and running before you run the installer, and installing Docker may require elevated privileges on Linux.
### npm install fails with permission errors
If `npm install` fails with an `EACCES` permission error, do not run npm with `sudo`. Instead, configure npm to use a directory you own:
```bash
mkdir -p ~/.npm-global
npm config set prefix ~/.npm-global
export PATH=~/.npm-global/bin:$PATH
```
Add the `export` line to your `~/.bashrc` or `~/.zshrc` to make it permanent, then re-run the installer.
### Installer fails on NVIDIA Jetson
The installer auto-detects NVIDIA Jetson devices (Orin and Thor) and applies required host configuration before the normal install flow. If the Jetson setup step fails, verify that you have `sudo` access and that Docker is installed and running.
For JetPack 6 (L4T 36.x), the setup switches iptables to legacy mode and adjusts the Docker daemon configuration. For JetPack 7 (L4T 38.x / Thor), only bridge netfilter and sysctl settings are applied. For JetPack 7 (L4T 39.x), bridge netfilter is loaded only when the host is missing it. Some R39 images already ship with `br_netfilter` configured and are left untouched. On affected R39 hosts, the installer prints `loading br_netfilter (required by k3s inside the OpenShell gateway)`. Without this fix, sandbox pods fail DNS resolution against the in-cluster service and the onboard `Setting up OpenClaw inside sandbox` step times out.
If the L4T version is not recognized, the setup step is skipped and the installer continues normally.
### DNS resolution from inside docker fails (corporate firewall)
Some corporate networks block outbound UDP port 53 to public DNS servers and force all host name resolution through DNS over TLS on TCP port 853. Containers do not inherit the host's DNS-over-TLS configuration, so the sandbox build's `npm ci` step times out trying to resolve `registry.npmjs.org` against `1.1.1.1` or `8.8.8.8`.
NemoClaw's preflight runs a short `docker run --rm busybox nslookup nemoclaw-dns-probe-.invalid` probe before starting the sandbox build. The fresh `.invalid` name should return NXDOMAIN through a working resolver, so cached answers cannot hide blocked DNS egress. When the probe confirms a DNS failure, onboarding stops with platform-specific remediation instead of hanging for ~15 minutes and printing a cryptic `Exit handler never called`.
Use the preflight headline to choose the recovery path:
- If no DNS servers could be reached, Docker could not reach its configured resolver. Follow the platform-specific UDP port 53 and Docker DNS steps below.
- If the DNS server was reachable but rejected the query with `NXDOMAIN` or `REFUSED`, the resolver answered, so the UDP port 53 fix is not relevant. Check the resolver used by Docker, such as dnsmasq, Pi-hole, unbound, or systemd-resolved, and remove any forwarding rule, blocklist entry, or ACL that rejects `registry.npmjs.org`. If needed, configure Docker to use an organization-approved resolver that can resolve public names, restart Docker, and retry onboarding.
For an unreachable resolver, pick the matching platform path below, apply it, then re-run `$$nemoclaw onboard`.
- **Linux with systemd-resolved.** Add a `DNSStubListenerExtra` drop-in pointing at the docker bridge gateway IP (the preflight prints the detected IP), then add the same IP to `/etc/docker/daemon.json` under `dns`. Restart `systemd-resolved` and `docker`.
- **macOS with Colima.** Restart Colima with the corporate DNS address, for example `colima stop && colima start --dns `.
- **macOS with Docker Desktop.** Add the corporate DNS address to `~/.docker/daemon.json` under `dns`, then restart Docker Desktop.
- **Windows or WSL.** Configure DNS in the Docker Desktop settings GUI, or apply the Linux fix above when running native docker inside WSL.
Verify the fix worked:
```bash
docker run --rm busybox nslookup registry.npmjs.org
```
When the lookup returns an answer, retry onboarding.
### Direct DNS lookups fail in a Docker-driver GPU sandbox
Egress covered by OpenShell network policies resolves destinations through the gateway. A direct DNS lookup inside the agent network namespace can fail on Docker-driver GPU hosts such as DGX Spark even when policy-covered inference, messaging, and search work normally. Direct in-sandbox DNS depends on Docker and the host resolver and is not a supported NemoClaw network-policy path, so NemoClaw does not use it as a sandbox health check.
Run a manual lookup only when you are diagnosing a custom tool that performs its own DNS resolution:
```bash
openshell sandbox exec --name -- getent hosts host.openshell.internal
openshell sandbox exec --name -- getent hosts example.com
```
If the host alias resolves but the external name does not, Docker's embedded resolver may be forwarding to an upstream DNS server that the sandbox bridge cannot use. Do not treat this result as evidence that a policy-covered feature is unhealthy. Test the affected inference, messaging, or search request through its normal policy path and inspect denied requests with `openshell term`.
If a custom tool requires direct DNS, configure the Docker daemon to use a resolver that containers can reach. On hosts with VPN or split-DNS software, use an upstream resolver that remains reachable from the Docker bridge, then recreate or rebuild the sandbox. Keep bridge networking enabled so the sandbox retains its normal Docker network isolation.
### Host DNS resolution is blocked before provider validation
NemoClaw also checks that the host process can resolve the provider host before it starts NVIDIA provider validation. A firewall rule that blocks host DNS traffic on port `53` can make later validation fail with `curl: (6) Could not resolve host: integrate.api.nvidia.com` even when container DNS probes look healthy. Current onboarding stops earlier with a host DNS diagnostic and remediation hints.
Verify host DNS outside NemoClaw:
```bash
node -e 'require("node:dns").resolve4("integrate.api.nvidia.com", (err, addrs) => { if (err) { console.error(err); process.exit(1); } console.log(addrs.join(",")); })'
```
Fix the host firewall, VPN, or DNS policy so the host can resolve the provider endpoint, then rerun onboarding. If you intentionally use a non-NVIDIA provider and need to bypass only this preflight, set `NEMOCLAW_SKIP_HOST_DNS_PREFLIGHT=1`.
### Port already in use
The NemoClaw dashboard uses port `18789` by default and the gateway uses port `8080`. If another sandbox already owns the dashboard port, onboarding scans ports `18789` through `18799` and uses the next free port. If all ports in that range are occupied, the error lists the owner for each port and suggests using `--control-ui-port` with a port outside the range.
NemoClaw allocates each Hermes sandbox an OpenAI-compatible API port from `8642` through `8652`, so it rejects every port in that range as a dashboard port for any agent. When all ports in the API range are occupied, the error lists the owner for each port. Destroy a listed Hermes sandbox or stop a listed non-OpenShell listener, then rerun onboarding.
On macOS, the port check also tries a privileged `lsof` probe without prompting for a password so root-owned listeners are detected before the sandbox build starts. For a new sandbox, NemoClaw reserves the selected loopback port through sandbox preparation and the image build. If another listener claims the port before NemoClaw binds the reservation, NemoClaw selects another port before changing sandbox resources. NemoClaw releases the reservation immediately before its receipt-owned ForwardTcp service starts. If forwarding then fails, onboarding removes the new sandbox and tells you to resolve the reported error before retrying.
When a previous onboard, upgrade, or sandbox crash leaves a stale `openclaw-gateway` host process holding the dashboard port, `$$nemoclaw onboard --fresh`, `$$nemoclaw destroy` (when destroying the last sandbox), and `$$nemoclaw uninstall` automatically sweep the dashboard port range and signal `SIGTERM` then `SIGKILL` to recover. The sweep only targets processes owned by the current user whose command line matches `openclaw-gateway` or `openshell forward` markers, and skips dashboard ports owned by other live sandboxes.
If onboarding preflight resolves the complete listener set for a gateway port conflict, the diagnostic lists every listener. Each entry contains the process name and PID, or only the PID when NemoClaw cannot read the process name. The diagnostic identifies listeners that fail ownership verification, but it does not print a reusable process-stop command. If NemoClaw resolves no listener, the diagnostic provides an `lsof` inspection command. Before you stop a listener, confirm that it is not part of a second NemoClaw gateway environment. Release that environment with `NEMOCLAW_GATEWAY_PORT= $$nemoclaw uninstall` instead of stopping its process.
If a non-NemoClaw process is already bound to the dashboard port or the gateway port, identify the conflicting process. Stop it only when its command line names the application you intend to stop, you own the process or administer its service, and the application has no active work:
```bash
sudo lsof -i :18789 -sTCP:LISTEN -P -n
```
Stop it through its service manager when one owns it. Otherwise, repeat the listener check immediately before you signal only the PID from that fresh result. Repeat the check again before any `SIGKILL`. Then retry onboarding.
Alternatively, override the conflicting port instead of stopping the other process. Pass `--control-ui-port` with the desired dashboard port:
```bash
$$nemoclaw onboard --control-ui-port 19000
```
You can also set `CHAT_UI_URL` with the desired port:
```bash
CHAT_UI_URL=http://127.0.0.1:19000 $$nemoclaw onboard
```
Or set the port directly:
```bash
NEMOCLAW_DASHBOARD_PORT=19000 $$nemoclaw onboard
```
For an OpenShell gateway port conflict, set `NEMOCLAW_GATEWAY_PORT` to a free non-privileged port that does not overlap NemoClaw's dashboard, vLLM, Ollama, or Ollama proxy ports:
```bash
NEMOCLAW_GATEWAY_PORT=8990 $$nemoclaw onboard
```
Remote/headless hosts should keep the OpenShell gateway on loopback and bind the dashboard forward instead:
```bash
NEMOCLAW_DASHBOARD_BIND=0.0.0.0 NEMOCLAW_GATEWAY_PORT=8990 $$nemoclaw onboard
```
Use `NEMOCLAW_DASHBOARD_BIND=0.0.0.0` again on later `$$nemoclaw connect` calls. If the sandbox was originally created without remote bind, recreate it with the same onboard command plus `--recreate-sandbox` before connecting remotely.
NemoClaw rejects `NEMOCLAW_GATEWAY_BIND_ADDRESS=0.0.0.0` for Docker-driver gateways while gateway JWT auth is active.
Use `NEMOCLAW_GATEWAY_BIND_ADDRESS=0.0.0.0` only on supported gateway modes and only when other hosts on the network should be able to reach the gateway.
### Older-glibc gateway compatibility container
OpenShell 0.0.116 directly supports Linux hosts with glibc 2.39 or newer. On an older trusted host, `NEMOCLAW_OPENSHELL_GATEWAY_CONTAINER_PATCH=1` explicitly opts into NemoClaw's compatibility container. Leave it unset on supported hosts.
The compatibility container uses host networking and mounts the host Docker socket read-only. A read-only socket mount still permits privileged Docker API operations and can control the host, so do not enable this mode on an untrusted or shared host. The gateway remains loopback-bound, and startup fails closed unless the configured Unix socket answers as a Docker daemon. See [Gateway Compatibility Container](../security/security-controls/gateway-authentication-controls#gateway-compatibility-container) for the container boundary and removal conditions.
Refer to [Environment Variables](commands#environment-variables) for the full list of port overrides.
### Running multiple sandboxes simultaneously
Each sandbox requires its own dashboard port. If you onboard a second sandbox without overriding the port, onboarding uses the next free port in the `18789` to `18799` range. `onboard` checks `openshell forward list` before starting a new forward, so a second onboard cannot silently take over the first sandbox's port.
Assign a distinct port only when you want a specific value:
```bash
$$nemoclaw onboard # first sandbox uses default 18789
$$nemoclaw onboard # second sandbox uses the next free port
$$nemoclaw onboard --control-ui-port 19000 # explicit port override
```
Each sandbox then has its own SSH tunnel and its own dashboard URL:
```text
http://localhost:18789 ← first sandbox
http://localhost:19000 ← second sandbox
```
You can verify which tunnel belongs to which sandbox with:
```bash
openshell forward list
$$nemoclaw list
```
`$$nemoclaw list` prints the recorded dashboard URL for each sandbox. These dashboard ports are separate from the gateway-wide inference route.
### A Gateway Port Stays Bound After Uninstall or Re-Onboard
A gateway port that keeps listening after you uninstall, or that keeps serving after you re-onboard without `NEMOCLAW_GATEWAY_PORT` set, can belong to a second environment.
Onboarding under a non-default `NEMOCLAW_GATEWAY_PORT` registers the sandbox on gateway `nemoclaw-` and stores its registry and state under `~/.nemoclaw/gateways//`. Gateway-lifecycle commands such as `$$nemoclaw onboard`, `$$nemoclaw uninstall`, and the deprecated `$$nemoclaw stop` act on the state root that their own `NEMOCLAW_GATEWAY_PORT` selects, while sandbox-scoped commands that read, exec into, or inspect a sandbox, such as `$$nemoclaw exec` and `$$nemoclaw status`, resolve the named sandbox across every per-port registry root and address its recorded gateway binding. A gateway-lifecycle command without that variable uses port `8080` unless the installer recorded an automatic alternate-port selection; in that case, the CLI restores the recorded port. Clearing an explicitly supplied variable does not move an existing sandbox back to the default port.
List the gateways the host still has:
```bash
openshell gateway list
```
If this happened after the deprecated global `$$nemoclaw stop`, read the command's final status line. Without a resolved sandbox name, the command releases a gateway only when a valid, explicitly set `NEMOCLAW_GATEWAY_PORT` selects it. `Host services stopped; managed gateway not released.` means the command stopped scoped host services but intentionally left the gateway listener running. After `openshell gateway list` confirms the exact target port, rerun the deprecated full stop with that scope only when you intend to release that gateway:
```bash
NEMOCLAW_GATEWAY_PORT=9000 $$nemoclaw stop
```
If NemoClaw reports that release was not confirmed, inspect the remaining listener and stop it only after you verify that it belongs to the selected gateway.
If the gateway name and its port-scoped state remain, treat it as a second environment and select that port for cleanup. If the gateway is absent but the port still listens, cleanup did not stop the listener; follow the process or service remediation printed by uninstall before you retry. If uninstall reported that it kept an `openshell-gateway` process owned by another user running, that process still holds the port. This can happen after uninstall exits successfully because NemoClaw does not treat another user's process as a cleanup failure. Ask that user to stop the process, or onboard under a different `NEMOCLAW_GATEWAY_PORT`.
Remove one environment by selecting its port:
```bash
NEMOCLAW_GATEWAY_PORT=9000 $$nemoclaw uninstall
```
Remove every gateway port in one run:
```bash
$$nemoclaw uninstall --all-gateway-ports
```
After you confirm uninstall and it exits with status `0`, run `openshell gateway list` again. For a NemoClaw-managed gateway without `--keep-openshell`, the gateway name that uninstall removed must be absent. An externally supervised gateway or a run with `--keep-openshell` preserves the gateway process and its resources.
The same scoping applies to `$$nemoclaw stop`. When `stop` reports that no valid gateway binding is registered for a sandbox, the sandbox can be registered under a different gateway port. Rerun `stop` with that `NEMOCLAW_GATEWAY_PORT` value set. If that does not find the sandbox, resolve the missing, invalid, or unreadable registry entry that the command reports.
Refer to [Uninstall NemoClaw](../manage-sandboxes/operate-sandboxes/uninstall-nemoclaw) for the full sweep contract.
### A shared inference route conflicts with another sandbox
Refer to [Use Shared Gateway Routes](../inference/manage-inference/use-shared-gateway-routes) for route time-sharing, provider-global compatibility, and status drift fields.
If `inference set` reports a valid shared-route conflict, align the named sandbox records or remove a sandbox you no longer need. If onboarding or `connect` reports a provider-global identity conflict, align the same-name provider's custom endpoint, API family, and credential environment-variable name across the named sandboxes, or remove a conflicting sandbox you no longer need.
If the error names incomplete legacy custom-route metadata, back up and remove the affected sandbox, then re-onboard it with an explicit custom endpoint and API family. For an OpenAI-compatible route, replace the example endpoint, model, and sandbox name in this recovery sequence:
```bash
$$nemoclaw legacy-sandbox destroy
NEMOCLAW_PROVIDER=custom \
NEMOCLAW_ENDPOINT_URL=https://endpoint.example/v1 \
NEMOCLAW_MODEL=your-model-id \
NEMOCLAW_PREFERRED_API=openai-completions \
$$nemoclaw onboard --name legacy-sandbox
```
If the error names an invalid gateway binding, restore the affected row's known-good `gatewayName` and `gatewayPort` metadata from a trusted backup; otherwise back up and remove the sandbox, then re-onboard it. Do not guess or copy a binding from another sandbox because lifecycle commands use it to select the gateway.
## Onboarding
### Cgroup v2 errors during onboard
Older NemoClaw releases relied on a Docker cgroup workaround on Ubuntu 24.04, DGX Spark, and WSL2. Current OpenShell releases handle that behavior themselves, so NemoClaw no longer requires a Spark-specific setup step.
If onboarding reports that the default Docker provider is missing or unreachable, fix Docker first and retry onboarding:
```bash
$$nemoclaw onboard
```
For explicitly selected native Podman, retain `NEMOCLAW_GATEWAY_RUNTIME=podman` and follow [Podman](#podman). An auto-detected Podman compatibility socket without that selector is rejected instead of silently changing providers.
### Cluster fails with `overlayfs snapshotter cannot be enabled` on Docker 26+
Docker Engine 26 and later default fresh installations to the [containerd image store](https://docs.docker.com/engine/storage/containerd/), which exposes its layers via the `overlayfs` snapshotter rather than the legacy `overlay2` graph driver. The k3s server inside the OpenShell cluster image needs to mount its own overlay filesystem on top, and the kernel rejects nesting two non-trivial overlay mounts. The cluster container then loops with:
```text
"overlayfs" snapshotter cannot be enabled for "/var/lib/rancher/k3s/agent/containerd",
try using "fuse-overlayfs" or "native":
failed to mount overlay: ... err: invalid argument
```
This is a Docker default-driver change, not a NemoClaw or OpenShell regression. The same hardware uses the legacy `overlay2` driver and is unaffected when it runs Docker 25 or earlier, or any Docker version with the containerd image store disabled.
NemoClaw detects the Docker 26+ containerd-snapshotter overlayfs configuration during onboarding and transparently builds a small drop-in replacement for the cluster image on the local Docker engine. The patched image installs `fuse-overlayfs` and selects it as the k3s snapshotter, bypassing the kernel-level nested-overlay limitation. No host configuration changes, sudo, or Docker restart required.
The auto-fix runs once per OpenShell version on the affected host. Subsequent onboarding runs reuse the cached patched image. Hosts without the conflict (`Driver: overlay2` in `docker info`, macOS Docker Desktop, or Linux installations that disable the containerd image store) see no change in behavior.
Override knobs:
- `NEMOCLAW_DISABLE_OVERLAY_FIX=1`: skip the auto-fix and run against the unmodified upstream cluster image. Useful for diagnosis or when you have already applied the manual workaround below.
- `NEMOCLAW_OVERLAY_SNAPSHOTTER=native`: build the patched image with k3s's `native` snapshotter instead of `fuse-overlayfs`. The `native` snapshotter copies image layers instead of overlaying them, so it uses more disk but does not depend on FUSE. Default is `fuse-overlayfs`.
If you prefer to disable the new Docker storage driver instead of running the patched image, edit `/etc/docker/daemon.json`:
```json
{
"storage-driver": "overlay2",
"features": { "containerd-snapshotter": false }
}
```
Then restart Docker (`sudo systemctl restart docker`) and re-run `$$nemoclaw onboard`. This restores the legacy `overlay2` driver host-wide, which kills any other running containers. Prefer the auto-fix unless you need the change for unrelated reasons. Switching storage drivers also rebuilds the entire local image graph: previously-pulled images become unusable and Docker re-pulls them on first reference, so expect a cold cache and additional disk usage right after the restart.
### OpenShell version above maximum
Each NemoClaw release validates against a range of tested OpenShell versions. If the installed OpenShell version exceeds the configured maximum, `$$nemoclaw onboard` exits with an error:
```text
✗ openshell is above the maximum supported by this NemoClaw release.
blueprint.yaml max_openshell_version:
```
Upgrade NemoClaw to a version that supports your OpenShell release, or install a supported OpenShell version from the [OpenShell releases page](https://github.com/NVIDIA/OpenShell/releases).
For fresh installs, NemoClaw passes the blueprint range to `install-openshell.sh` and resolves a compatible published OpenShell release before downloading. If GitHub release metadata is unavailable, the script uses its bundled fallback pin and the post-install gate still enforces the configured range.
### Installer Reports an OpenShell Gateway Version Mismatch
On Linux, an existing OpenShell package can provide a systemd user service that starts a different gateway version from the user-local version that NemoClaw installs. The installer stops before onboarding instead of using the two versions together. The error reports both gateway versions and binary paths.
Do not remove the existing OpenShell package if its gateway manages resources
outside NemoClaw. Package removal can stop that gateway. Align the package
with the version in the installer error, or plan the migration of those
resources first.
If you no longer need the APT-installed OpenShell package, remove it and rerun the installer:
```bash
sudo apt remove openshell
curl -fsSL https://www.nvidia.com/nemoclaw.sh | bash
```
The next installer run must continue past the OpenShell installation step without reporting a version mismatch.
### Docker Driver Gateway Reports an Incompatible Migration
Onboarding can stop when the Docker driver gateway log contains both parts of either error:
```text
migration N was previously applied
is missing in the resolved migrations
```
```text
migration N was previously applied
has been modified
```
The first error means the installed OpenShell migration set does not contain migration N. The second error means that migration set defines migration N with different contents. NemoClaw identifies `/openshell.db` as incompatible with the installed OpenShell migration set for both errors. This failure can happen after an OpenShell downgrade. Installing a NemoClaw release that is older than the installed one performs that downgrade, because each release pins one OpenShell version and the installer reinstalls OpenShell at the pin. You reach that state in one of two ways:
- You select an older release with `NEMOCLAW_INSTALL_TAG` or `NEMOCLAW_INSTALL_REF`.
- You install an older OpenShell yourself.
The default installer stops before replacing a newer installed NemoClaw release whenever the selected ref is `lkg` or `refs/tags/lkg`, whether it came from `NEMOCLAW_INSTALL_TAG` or the higher-priority `NEMOCLAW_INSTALL_REF`. Other explicit refs can replace the installed release; use a version tag or full commit SHA when you need a fixed source.
The diagnosis always prints the database path. When an unused archive path is available, it also prints that archive path beside the selected state directory and the profile-specific onboarding command. When no unused archive path is available, it asks you to keep the gateway stopped and inspect the state directory instead.
The selected state directory contains the gateway database, mutual TLS private
keys, JSON Web Token signing material, and every sandbox and provider
registration on the selected gateway. Moving it makes those registrations and
credentials unavailable to the fresh gateway. Other sandboxes on the selected
gateway can require re-onboarding and credential entry.
When NemoClaw proves that one managed service owns the selected port and state, the recovery stops that service and retries onboarding.
The stop command names the upstream OpenShell package unit, the NemoClaw user service, or the Homebrew formula selected for this host.
If the listener, state namespace, or service identity does not match, NemoClaw prints no stop command and uses the standalone ownership checks.
The onboarding retry resolves ownership again before it can offer a state move.
When no managed service owns the selected port and state, the gateway runtime checks the recorded process and scans current gateway process identities for the selected state directory.
It withholds the move unless that scan establishes that no gateway process uses the state.
When NemoClaw prints a stop-and-retry command, run that command first. Do not move the state unless the retry prints a state-move command.
When NemoClaw prints the state move, run the commands it prints:
1. Create the printed `.incompatible` archive with owner-only access. If that path exists, NemoClaw adds a numeric suffix instead of nesting or replacing an earlier archive.
2. Move the selected state directory into the archive as `gateway-state`.
3. Run the printed onboarding command only after the archive and move succeed.
The archive remains beside the selected state directory and retains the previous gateway records and credentials. Keep it owner-only until onboarding completes and every required sandbox and provider registration is restored. Delete the archive only after you no longer need its gateway records or credentials for recovery.
### Installer Reports That the Systemd User Manager Is Unavailable
On Linux, an OpenShell package can install `/usr/lib/systemd/user/openshell-gateway.service` on a host without a reachable systemd user manager. The service query can then return this diagnostic:
```text
Failed to connect to bus: No medium found
```
The installer accepts only recognized user-manager-unavailable diagnostics for the standalone gateway fallback.
It checks `.wants`, `.requires`, and `.upholds` links in the standard systemd user unit paths.
When `NEMOCLAW_GATEWAY_PORT` selects port `8080` and no activation path exists, the installer keeps the standalone gateway on port `8080`.
When `NEMOCLAW_GATEWAY_PORT` is unset, the installer can continue with one qualified `openshell-gateway.service` activation.
The activation must point to a trusted package unit without drop-ins, and its environment must set `OPENSHELL_SERVER_PORT=8080` exactly once.
The installer selects the first nonreserved and unconfigured port from `8990` through `9005` that has no listener, gateway registration, or port-scoped NemoClaw state.
A supported listener probe must conclusively confirm that the port is unused.
Installation stops if no candidate passes, including when listener inspection is unavailable or inconclusive for every remaining candidate.
It does not modify the existing service or activation link.
After selection and before onboarding runs or is deferred, the installer records the selected value under `~/.nemoclaw/gateways//automatic-gateway-port.pending`.
After installer-driven or direct CLI onboarding succeeds, it promotes that file to `automatic-gateway-port`.
Deferred, failed, or cancelled onboarding keeps the pending identity so a later installer or CLI process restores the same port and reaches its onboarding or cleanup path instead of creating another environment.
Later NemoClaw commands restore either recorded identity when `NEMOCLAW_GATEWAY_PORT` is unset.
You can still use the printed export command when a script or surrounding process must carry the same scope explicitly.
Exporting it makes that port an explicit operator selection, including authority for a no-name gateway stop.
If the installer selects port `8990`, it prints this command:
```bash
export NEMOCLAW_GATEWAY_PORT=8990
```
Later `list`, `status`, `connect`, `stop`, `rebuild`, and `uninstall` commands use the recorded automatic port unless this variable explicitly selects another supported port.
A restored automatic port selects command state but does not count as explicit operator authority for the deprecated no-name `$$nemoclaw stop`; name the sandbox or set `NEMOCLAW_GATEWAY_PORT` explicitly when you intend to release a gateway.
The installer does not replace a supported explicit non-default gateway port.
It uses that port's detached lifecycle without automatic selection.
The installer stops when port `8080` is explicit, the activation evidence is ambiguous, the activation is not an `openshell-gateway.service` symlink to an accepted package-unit path, or the effective service port is not proven to be `8080`.
It also stops when `SYSTEMD_UNIT_PATH` overrides the standard paths.
Restore the systemd user manager, then inspect both possible services:
```bash
systemctl --user status openshell-gateway.service
systemctl --user is-enabled openshell-gateway.service
systemctl --user status nemoclaw-openshell-gateway.service
systemctl --user is-enabled nemoclaw-openshell-gateway.service
```
Resolve an unqualified service through its package or platform owner.
Do not delete an activation link or edit a unit file by hand.
Rerun the installer after the owner confirms the effective port and activation state.
Unknown service query errors remain fatal. The installer also stops for malformed effective metadata, an untrusted unit or executable path, an executable failure, or a gateway version mismatch. Follow the reported condition instead of forcing the standalone fallback.
### Sandbox build fails during OpenClaw plugin install
During sandbox creation, the OpenClaw image setup can install managed plugins for selected features such as web search or diagnostics. If the build reaches `openclaw plugins install` and the npm registry or ClawHub is blocked, NemoClaw classifies that narrow failure and prints a policy hint instead of only generic resume guidance. Brave Search uses an external OpenClaw plugin and can reach this install path. Tavily ships with the pinned OpenClaw runtime, so NemoClaw verifies the bundled extension instead of installing a separate Tavily package. The neutral managed image does not declare `plugins.entries.tavily`; onboarding adds that entry only when you select Tavily Search.
Check that the active policy and host network allow the npm registry and ClawHub endpoints needed by the plugin, or disable the feature that requested the plugin. For example, if the plugin is for web search, disable that feature and resume onboarding:
```bash
NEMOCLAW_WEB_SEARCH_PROVIDER=none $$nemoclaw onboard --resume
```
If you want the feature, fix the network or policy path first, then resume onboarding:
```bash
$$nemoclaw onboard --resume
```
### Web search verification reports a warning or security error
When web search is enabled, onboarding checks the selected agent configuration and sends a real search request through the sandbox egress path. Configuration and egress verification are best effort, so those failed checks print a warning and let onboarding finish. The selected credential's live sandbox isolation check is required. If NemoClaw confirms that the raw Brave or Tavily key is visible, or the sandbox does not return a valid isolation result, it reports a security error. The CLI pauses onboarding and exits with a nonzero status. Recreate that sandbox through the supported onboarding flow before using it:
```bash
$$nemoclaw onboard --recreate-sandbox
```
First confirm that the provider credential and matching policy preset exist.
```bash
$$nemoclaw credentials list
$$nemoclaw policy list
```
Look for `-brave-search` with the `brave` preset or `-tavily-search` with the `tavily` preset. Do not replace an `openshell:resolve:env:` value in the sandbox configuration with a raw API key.
Confirm that OpenClaw reports the provider selected during onboarding.
```bash
$$nemoclaw config get --key tools.web.search --format yaml
```
The provider should be `brave` or `tavily` and `enabled` should be `true`. If the provider is wrong, rerun onboarding with `NEMOCLAW_WEB_SEARCH_PROVIDER=brave` or `tavily` and the matching `BRAVE_API_KEY` or `TAVILY_API_KEY`.
Confirm that the generated Hermes configuration selects the Tavily backend.
```bash
$$nemoclaw exec -- cat /sandbox/.hermes/config.yaml
```
The output should include a `web` mapping with `backend: tavily`. If it does not, rerun onboarding with `NEMOCLAW_WEB_SEARCH_PROVIDER=tavily` and `TAVILY_API_KEY`.
Rerunning onboarding with a different provider recreates the sandbox because the provider configuration and credential attachment are build-time inputs. NemoClaw validates the replacement key before it removes the existing sandbox, then backs up and restores the supported workspace state during recreation. If the configuration is correct but the egress probe fails, keep the matching preset applied and inspect the blocked request with `openshell term` before widening any policy rule.
### Sandbox containers cannot reach the gateway
On native Linux Docker-driver hosts, `$$nemoclaw onboard` verifies the route that sandbox containers use to reach the OpenShell gateway. If a host firewall blocks that path, onboarding exits with output like:
```text
✗ Sandbox containers cannot reach the gateway at host.openshell.internal:8080.
A host firewall may be blocking traffic from the OpenShell Docker bridge.
```
Apply the `ufw` command printed by onboarding, then rerun onboarding. If the message does not include a subnet, derive it from the OpenShell Docker network:
```bash
SUBNET=$(docker network inspect openshell-docker --format '{{(index .IPAM.Config 0).Subnet}}')
sudo ufw allow from "$SUBNET" to any port 8080 proto tcp
$$nemoclaw onboard
```
This reachability check uses a disposable Docker probe and does not create or replace a sandbox. If Docker GPU compatibility recreation fails later, follow [GPU routing or compatibility patch failed](#gpu-routing-or-compatibility-patch-failed). That path can restore the pre-patch sandbox. If its diagnostics report manual cleanup, use only the printed exact-container command. That command targets the failed replacement and preserves the restored sandbox.
### Custom OpenClaw image creates without a gateway or dashboard
`$$nemoclaw onboard --from ` treats the supplied Dockerfile as the complete sandbox image rather than adding it on top of the stock managed runtime. If deployment verification cannot reach the gateway, NemoClaw checks for `/tmp/gateway.log`, `/usr/local/bin/nemoclaw-start`, and `/sandbox/.openclaw/openclaw.json` in the custom sandbox. When all three paths are absent, the CLI reports that the image lacks the NemoClaw-managed OpenClaw runtime and does not suggest repeated dashboard port-forward retries. This failure commonly occurs when the custom Dockerfile starts from `ghcr.io/nvidia/nemoclaw/sandbox-base` alone because that image is an intermediate dependency image.
Rebuild the custom image from the full stock Dockerfile and source context for the same NemoClaw release. For the version-pinned plugin workflow, refer to [Install OpenClaw Plugins](../manage-sandboxes/install-openclaw-plugins).
If the sandbox is unreachable or the managed runtime paths are present, NemoClaw retains the existing generic gateway-log and host OpenShell-log guidance because the base-only failure is not proven.
A custom image without the managed runtime can fail while NemoClaw starts the sandbox container. NemoClaw reports exit code 127 without assigning a cause unless captured logs contain the exact `env` error for missing `nemoclaw-start`. When that error is present, the failure output identifies the missing managed startup command and gives the same rebuild guidance. If NemoClaw saves pre-rollback diagnostics, the reported directory contains the captured container logs. If rollback succeeds, NemoClaw restores and starts the pre-patch sandbox container. It does not print a sandbox deletion command for the restored sandbox. If the failed replacement container remains, NemoClaw prints an exact-container Docker cleanup command. If NemoClaw cannot confirm whether the replacement remains but retains its validated exact ID, it prints the same target-safe command. Without a validated exact ID, it reports cleanup as unknown and prints no deletion command. If rollback fails, sandbox and container state can be uncertain. If NemoClaw reports a diagnostics directory, inspect it. Inspect the diagnostics before removing any container.
### `connect` exits because the gateway is down
`$$nemoclaw connect` checks the OpenShell gateway before it tries dashboard forwarding, SSH, or inference repair. If the gateway is not reachable, the command exits early and prints recovery guidance.
Resume onboarding so NemoClaw recreates or reconnects the managed gateway, then retry:
```bash
$$nemoclaw onboard --resume
$$nemoclaw connect
```
Run `$$nemoclaw status` for a broader gateway health report.
If `nemohermes recover` reports that the Hermes secret-boundary validator is missing, the sandbox image predates the recovery-side validator that re-checks `/sandbox/.hermes/.env`.
Current NemoClaw releases fail closed in this state: recovery reports that the validator is missing, leaves an otherwise healthy gateway untouched, refuses to claim the secret boundary was checked, and instructs you to re-image the sandbox with a current Hermes build.
Re-image the sandbox with a current Hermes build before retrying recovery:
```bash
nemohermes rebuild --yes
nemohermes recover
```
### Sandbox Container Is Unhealthy While the Agent Gateway Process Is Alive
Use NemoClaw commands to distinguish a stopped sandbox from a stopped or unhealthy OpenClaw gateway.
```bash
$$nemoclaw status
$$nemoclaw logs
```
If OpenShell reports the sandbox as `Stopped`, run `$$nemoclaw start` so OpenShell owns the sandbox lifecycle transition.
If the sandbox is available but the built-in native agent gateway is stopped or unhealthy, request its supported native restart:
```bash
$$nemoclaw gateway restart
```
The restart command invokes OpenClaw's native lifecycle command, verifies readiness, and checks or repairs host-side forwards.
NemoClaw does not use an in-sandbox serving watchdog to relaunch the gateway.
`recover` reports a stopped built-in native agent gateway and does not relaunch it.
Use `recover` after the gateway is running when you need to verify supported gateway state or repair host-side forwards:
```bash
$$nemoclaw recover
```
If the restart fails for an ordinary process or health failure, correct the reported cause, then reset the sandbox through OpenShell:
```bash
$$nemoclaw stop
$$nemoclaw start
```
Rebuild only if the sandbox still cannot start.
### Invalid sandbox name
Sandbox names must contain 1 to 19 characters. They must be lowercase, start with a letter, contain only letters, numbers, and single internal hyphens, and end with a letter or number. Consecutive hyphens (`--`) are not allowed. The CLI rejects names that do not match these rules. It prints a `Try: ` recovery line whenever it can derive a valid lowercase, hyphen-separated form from the input, so passing `--name MyAssistant` reports `Try: myassistant` and you can rerun with the suggested slug.
The CLI writes the rejected value as a quoted preview instead of raw input. The preview reads at most the first 80 UTF-16 code units from the input and escapes each code unit outside printable ASCII as `\uXXXX`. Escaping can make the preview longer than 80 output characters. This prevents a rejected name from injecting control sequences into terminal or CI output.
Names that collide with global CLI commands are also rejected.
Reserved names include `onboard`, `list`, `setup`, `start`, `stop`, `status`, `debug`, `uninstall`, `credentials`, and `help`.
Using a reserved name would cause the CLI to route to the global command instead of the sandbox.
If the name does not match these rules or is reserved, the wizard exits with an error. Choose a name such as `my-assistant` or `dev1`.
### Sandbox creation fails on DGX
On DGX machines, sandbox creation can fail if the gateway's DNS has not finished propagating or if a stale port forward from a previous onboard run is still active.
Run `$$nemoclaw onboard` to retry. The wizard cleans up stale port forwards and waits for gateway readiness automatically.
### GPU Setup Fails with a Placeholder GPU Name
On Windows, WSL, and native Linux ARM64 hosts, some systems report a placeholder display adapter name even when no NVIDIA GPU firmware is present. This section also applies when preflight reports no GPU on an ARM64 Linux host whose `nvidia-smi` shows a non-placeholder GPU name. NVIDIA NIM and GPU-backed sandbox setup require a real NVIDIA GPU.
When the primary memory-query probe reports a placeholder-named GPU row on an eligible ARM64 Linux host without firmware-confirmed NVIDIA platform metadata, onboarding runs a bounded CUDA workload through the selected runtime provider. When that probe reports one or more plausible non-placeholder NVIDIA GPU names on WSL ARM64, where Windows paravirtualizes the GPU through `/dev/dxg`, onboarding runs the same workload once for every provider-visible GPU UUID. A non-placeholder name that does not identify an NVIDIA GPU or product family does not start the workload. NemoClaw treats a recognized NVIDIA product model from `/sys/class/dmi/id/product_name` or `/sys/firmware/devicetree/base/model`, or a known Tegra device node, as authoritative platform identity. Docker or Podman may pull the CUDA sample image from `nvcr.io` and keeps the image in the local cache after the container exits. Docker uses `--gpus all`; qualified rootless Podman uses the NVIDIA CDI device. The following manual diagnostics exercise the same fixed workload and per-device capacity query:
```bash
proof_image='nvcr.io/nvidia/k8s/cuda-sample@sha256:7c7540bdf1f942d4fb6db97069fd6c289471b54ac29e3c7fcdf914cf77af7d41'
proof_script='set -eu; snapshot="$(nvidia-smi --query-gpu=uuid,index,name,memory.total,memory.free --format=csv,noheader,nounits)"; device_uuids="$(printf "%s\n" "$snapshot" | cut -d, -f1)"; for device_uuid in $device_uuids; do CUDA_VISIBLE_DEVICES="$device_uuid" /cuda-samples/sample >/dev/null; done; printf "%s\n" "$snapshot" | sed "s/^/NEMOCLAW_GPU_DEVICE=/"'
# Docker
docker run --rm --gpus all --entrypoint /bin/sh "$proof_image" -c "$proof_script"
# Qualified rootless Podman socket from NemoClaw preflight
podman --url 'unix://' run --rm --device nvidia.com/gpu=all --entrypoint /bin/sh "$proof_image" -c "$proof_script"
```
A diagnostic result includes one successful CUDA sample invocation and one `NEMOCLAW_GPU_DEVICE=, , , , ` row for every provider-visible GPU. For Podman, replace `` with the exact socket accepted by [Podman preflight](#podman); do not omit `--url`, because a bare command can select another rootful, remote, or named connection. Manual output does not qualify onboarding. Only the proof that NemoClaw runs through its provider-bound authority validates the row structure, binds every workload to its stable GPU UUID, and supplies the capacity evidence used for model selection.
The run is bounded to 3 minutes by default. Set `NEMOCLAW_WSL_GPU_PROOF_TIMEOUT_MS` to a positive integer number of milliseconds to change that bound; values above 15 minutes are capped at 15 minutes. A passing workload with confirmed container reconciliation lets onboarding treat the detected GPU as eligible for GPU passthrough during that run. On Windows-on-Arm, only the documented N1x WSL path can use this proof for platform qualification; other GPU routes remain unsupported. A failed, interrupted, or timed-out workload leaves the GPU unproven and does not enable GPU passthrough. After every completed capture, NemoClaw immediately reconciles the proof container by its exact name and ownership label. After an interrupted capture, it continues that exact check for up to 15 seconds. If the resource appears, NemoClaw removes only its immutable container ID. If an interrupted capture remains absent through that deadline, or removal fails after any capture, NemoClaw invalidates the proof, reports unresolved cleanup, and names the exact resource to verify or remove before retrying. The names-only unified-memory fallback does not run this workload and rejects denylisted names. WSL uses Docker Desktop by default; a qualification-backed rootless Podman provider can run the same proof for WSL-local Ollama.
When GPU detection rejects the `nvidia-smi` report, preflight prints the failed check under the `Local NIM unavailable — no GPU detected` line, for example an absent `/proc/driver/nvidia` interface or a failed bounded CUDA proof. If NemoClaw rejects the detected GPU name during preflight, select a CPU or remote inference provider, or move the setup to a host with a supported NVIDIA GPU and current drivers.
Jetson/Tegra hosts support sandbox GPU passthrough through the compatibility route. Onboarding detects those hosts separately and propagates eligible host group IDs for selected `/dev/nvmap`, `/dev/nvhost-*`, and `/dev/nvgpu/igpu0/*` nodes plus real `/dev/dri/renderD*` character devices. If that path fails, follow the Jetson/Tegra compatibility guidance below instead of treating a missing `nvidia-smi` result as a placeholder adapter.
### Colima socket not detected (macOS)
Newer Colima versions use the XDG base directory (`~/.config/colima/default/docker.sock`) instead of the legacy path (`~/.colima/default/docker.sock`). Some installations expose a top-level Colima socket at `~/.colima/docker.sock`. NemoClaw checks all three paths. If neither is found, verify that Colima is running:
```bash
colima status
```
### Sandbox build is slow or hangs (under-provisioned container runtime)
Default Colima ships with 2 vCPU and 2 GiB of memory, which is not enough headroom for the BuildKit-driven sandbox image build. On macOS Apple Silicon, the build can stall part-way through with no progress and no error, leaving the wizard waiting indefinitely.
Preflight inspects `docker info` for `NCPU` and `MemTotal` and prints a warning when the runtime falls below 4 vCPU or 8 GiB. In interactive onboarding, the warning prompt defaults to abort, so pressing Enter stops the run before the sandbox build reaches the likely stall point. Type `y` only when you intentionally want to continue on the smaller runtime. Non-interactive onboarding prints the warning and continues. On Colima, raise the resources before re-running onboard:
```bash
colima stop
colima start --cpu 6 --memory 12 --disk 100
```
On Docker Desktop, raise CPU and memory limits in _Settings → Resources_, then apply and restart.
To silence the warning when the host is intentionally small, set `NEMOCLAW_IGNORE_RUNTIME_RESOURCES=1` before running `$$nemoclaw onboard`.
### Managed Sandbox Image Build Requires Local BuildKit
On a local Docker-driver gateway, NemoClaw builds each generated OpenClaw or Hermes sandbox image with host-side BuildKit. The generated Dockerfiles include BuildKit-only file-mode, per-step network, and mount controls. NemoClaw stops before sandbox creation in these cases:
- The local build is disabled.
- The staged build context fails trust validation.
- Docker cannot start the build.
- The build exits with an error.
It does not send that generated Dockerfile to the OpenShell gateway's classic Docker API builder because that builder cannot enforce the same instructions.
Keep the local prebuild enabled, and verify that the host Docker installation provides BuildKit:
```bash
unset NEMOCLAW_SANDBOX_PREBUILD
docker info
docker buildx version
```
Repair Docker access or the Docker Buildx plugin when either Docker command fails. Then rerun the original onboarding or rebuild command. For a resumable onboarding session, run:
```bash
$$nemoclaw onboard --resume
```
A passing recovery completes the local BuildKit build before sandbox creation starts. If NemoClaw rejects the staged build context trust boundary, do not change its permissions or move its Dockerfile. Rerun the command so NemoClaw creates a new private staged context. If the new context is also rejected, preserve the complete error and stop instead of forcing the gateway builder.
This requirement does not change user-supplied `--from` contexts, which continue to use the OpenShell gateway builder. It also preserves the gateway fallback when a generated LangChain Deep Agents Code image does not complete its local prebuild.
### Re-onboard fails because port 18789 is held by SSH
After destroying a sandbox and gateway, the SSH port-forward process for the dashboard can be left running. Re-running onboard then fails preflight with `Port 18789 is not available. Blocked by: ssh`.
Current NemoClaw detects this case and kills the orphaned SSH process automatically before retrying the port check. If you see the error on an older release, identify the SSH process. Use fresh listener output to confirm that your user owns the process and that it still listens on local port `18789`:
```bash
sudo lsof -i :18789 -sTCP:LISTEN -P -n
```
Inspect the process separately with `ps -p -o user=,args=`. Stop it only when the owner is your user, the command line is the stale SSH port forward for local port `18789`, and no active terminal or file-transfer session uses that process. Repeat the listener check immediately before you signal only the PID from that fresh result. Then re-run `$$nemoclaw onboard`.
### Sandbox Keeps Using the Previous Messaging Credential
Rerunning `$$nemoclaw onboard --non-interactive` with a replacement `TELEGRAM_BOT_TOKEN`, `DISCORD_BOT_TOKEN`, `SLACK_BOT_TOKEN`, `SLACK_APP_TOKEN`, `WECHAT_BOT_TOKEN`, or `MSTEAMS_APP_PASSWORD` previously reported success while the sandbox kept using the old credential. When you rerun onboarding, NemoClaw evaluates every active messaging credential binding and compares each supplied credential with its SHA-256 hash in the sandbox registry. When you provide a replacement credential, NemoClaw runs the channel's configured checks before it backs up supported workspace and manifest-declared state, destroys the sandbox, recreates it, and restores the backup. Files outside those state paths are not preserved. If an available pre-recreation check fails, onboarding stops before it backs up supported workspace and manifest-declared state or destroys the existing sandbox. Discord and Microsoft Teams require non-empty replacement input but cannot prove upstream credential validity before recreation, so send a real test message after onboarding and confirm that the recreated sandbox receives it and responds. If the channel state changes during rotation, onboarding stops before it destroys the existing sandbox and asks you to retry with the updated state. If you do not supply a credential, or the supplied credential matches the recorded hash, the credential check does not trigger recreation. If you replace a credential for a channel that you stopped with `channels stop`, onboarding does not trigger recreation because that channel is inactive.
If you suspect a sandbox is still using a stale messaging credential, follow [Rotate a Messaging Credential](../security/credential-rotation#rotate-a-messaging-credential) to export the replacement without putting it in shell history. Then rerun onboarding so the credential check runs:
```bash
$$nemoclaw onboard --name \
--non-interactive --yes --yes-i-accept-third-party-software
```
### Sandbox creation killed by OOM (exit 137)
On systems with 8 GB RAM or less and no swap configured, the sandbox image push can exhaust available memory and get killed by the Linux OOM killer (exit code 137).
NemoClaw automatically detects low memory during onboarding and prompts to create a 4 GB swap file. If this automatic step fails or you are using a custom setup flow, create swap manually before running `$$nemoclaw onboard`:
```bash
sudo dd if=/dev/zero of=/swapfile bs=1M count=4096 status=none
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab
$$nemoclaw onboard
```
### Onboarding Reports a Rejected or Unconfirmed Policy Update
Onboarding can submit several policy mutations in sequence when you deselect policy presets and select others. It submits deselection mutations before selection mutations. Each successful mutation updates and verifies the live OpenShell policy before the next mutation starts. If a later mutation fails, the earlier successful mutations remain in that live policy.
Each mutation uses a temporary `policy.yaml` file in a `nemoclaw-policy-*` directory. When the submission finishes, NemoClaw removes that directory, whether or not the gateway accepted the mutation. If cleanup fails, NemoClaw reports the directory that still holds the policy instead of reporting the submission result.
When NemoClaw reports this directory, do not retry the policy operation. Use Bash to enter and validate the exact path from the error before you remove its contents:
```bash
cleanup_retained_policy() {
local platform_tmp_root retained_policy_dir
if ! platform_tmp_root=$(node -p "require('node:os').tmpdir()"); then
printf 'Validation failed: Node.js could not report the platform temporary directory.\n' >&2
return 1
fi
platform_tmp_root=${platform_tmp_root%/}
IFS= read -r -p 'Enter the temporary policy directory from the error: ' retained_policy_dir
python3 - "$platform_tmp_root" "$retained_policy_dir" <<'PY'
import os
import re
import stat
import sys
class CleanupError(Exception):
pass
def open_retained_policy_directory(platform_tmp_root, reported_path):
if not os.path.isabs(platform_tmp_root):
raise CleanupError("the platform temporary directory is not absolute")
platform_tmp_root_real = os.path.realpath(platform_tmp_root)
if platform_tmp_root_real == os.path.abspath(os.sep):
raise CleanupError("the platform temporary directory cannot be the filesystem root")
reported_basename = os.path.basename(reported_path)
reported_parent = os.path.dirname(reported_path)
if not re.fullmatch(r"nemoclaw-policy-.+", reported_basename):
raise CleanupError("the basename must match nemoclaw-policy-* and include a suffix")
if reported_parent != platform_tmp_root:
raise CleanupError(
f"the path must be a direct child of the actual platform temporary directory {platform_tmp_root!r}"
)
if os.path.realpath(reported_path) != os.path.join(platform_tmp_root_real, reported_basename):
raise CleanupError("the canonical path is outside the actual platform temporary directory")
if (
not hasattr(os, "geteuid")
or not hasattr(os, "O_DIRECTORY")
or not hasattr(os, "O_NOFOLLOW")
or os.stat not in os.supports_dir_fd
or os.unlink not in os.supports_dir_fd
):
raise CleanupError("this platform cannot perform descriptor-relative cleanup")
directory_flags = os.O_RDONLY | os.O_DIRECTORY | getattr(os, "O_CLOEXEC", 0)
root_stat = os.stat(platform_tmp_root_real, follow_symlinks=False)
if not stat.S_ISDIR(root_stat.st_mode):
raise CleanupError("the canonical platform temporary path is not a directory")
root_fd = os.open(platform_tmp_root_real, directory_flags | os.O_NOFOLLOW)
try:
opened_root_stat = os.fstat(root_fd)
if (opened_root_stat.st_dev, opened_root_stat.st_ino) != (root_stat.st_dev, root_stat.st_ino):
raise CleanupError("the platform temporary directory changed while it was being opened")
reported_stat = os.stat(reported_basename, dir_fd=root_fd, follow_symlinks=False)
if not stat.S_ISDIR(reported_stat.st_mode):
raise CleanupError("the path is not a directory or is a symbolic link")
if reported_stat.st_uid != os.geteuid():
raise CleanupError("the current user does not own the directory")
policy_fd = os.open(
reported_basename,
directory_flags | os.O_NOFOLLOW,
dir_fd=root_fd,
)
opened_stat = os.fstat(policy_fd)
if (opened_stat.st_dev, opened_stat.st_ino) != (reported_stat.st_dev, reported_stat.st_ino):
os.close(policy_fd)
raise CleanupError("the directory changed while it was being opened")
return root_fd, policy_fd, reported_basename, opened_stat
except Exception:
os.close(root_fd)
raise
def remove_retained_policy_material(policy_fd):
try:
os.unlink("policy.yaml", dir_fd=policy_fd)
except FileNotFoundError:
pass
remaining = os.listdir(policy_fd)
if remaining:
raise CleanupError(f"the directory contains unexpected entries: {remaining!r}")
def main():
root_fd = None
policy_fd = None
try:
root_fd, policy_fd, basename, opened_stat = open_retained_policy_directory(
sys.argv[1], sys.argv[2]
)
remove_retained_policy_material(policy_fd)
try:
current_stat = os.stat(basename, dir_fd=root_fd, follow_symlinks=False)
except FileNotFoundError as error:
raise CleanupError("the reported path changed during cleanup") from error
if (current_stat.st_dev, current_stat.st_ino) != (opened_stat.st_dev, opened_stat.st_ino):
raise CleanupError("the reported path was replaced during cleanup; the replacement was not removed")
print(
f"Removed retained policy material from {sys.argv[2]!r}. "
"The empty directory intentionally remains."
)
except (CleanupError, OSError) as error:
print(f"Cleanup failed for {sys.argv[2]!r}: {error}.", file=sys.stderr)
raise SystemExit(1) from error
finally:
if policy_fd is not None:
os.close(policy_fd)
if root_fd is not None:
os.close(root_fd)
if __name__ == "__main__":
main()
PY
}
cleanup_retained_policy
```
The procedure opens the validated directory without following a symbolic link and removes `policy.yaml` relative to that open directory. It fails if the directory contains anything else. It intentionally leaves the empty directory because deleting it later by pathname would reintroduce a directory-replacement race. Continue only after the procedure reports that the retained policy material was removed. Do not remove the empty directory by pathname. If validation or cleanup fails, preserve the exact path and complete error message for support. Do not use another removal command on that path.
After temporary-directory cleanup, treat the gateway state as unknown because the cleanup error replaced the submission result. Restore access to the OpenShell gateway, then read the sandbox policy:
```bash
$$nemoclaw policy list
```
If `policy list` reports that it could not query OpenShell, stop because no local policy state can substitute for the live document. If it reports container-runtime recovery guidance, restore that runtime and run `policy list` again. Do not resume or start fresh onboarding until `policy list` reports the live gateway policy.
When NemoClaw reports a rejected or unconfirmed policy mutation, onboarding stops and:
- leaves earlier successful mutations applied and recorded;
- does not record that mutation in the sandbox registry;
- marks the session failed and resumable;
- reports one of the two results below.
After either result, run `policy list` before you resume or start fresh onboarding.
The gateway read the policy and refused it:
```text
OpenShell rejected the policy for sandbox 'my-assistant' (exit ): . The policy was not applied and re-applying it will be rejected again; change the preset selection instead.
```
An OpenShell refusal means this mutation did not change the live policy. Earlier successful mutations in the same onboarding step remain applied and recorded. Use the quoted OpenShell diagnostic to identify what the gateway refused. After you inspect `policy list`, follow [Previous onboarding session failed](#previous-onboarding-session-failed) to start fresh onboarding and choose a different preset selection.
NemoClaw could not confirm the result. This covers a connection that ended before the result arrived, an unreachable gateway, an elapsed deadline, a rejected credential, and a refusal the gateway reported with a status NemoClaw does not recognize as final. NemoClaw reports this whenever the gateway did not return an explicit refusal, because only an explicit refusal proves the policy was not applied:
```text
Could not confirm the policy update for sandbox 'my-assistant': . The gateway may or may not have applied it; read the current policy back before retrying.
```
The gateway state is unknown, so read the sandbox policy with `policy list` before you retry. `policy list` derives applied state from OpenShell and has no local policy record to reconcile.
The same connection problem that made the result unconfirmed can also stop `policy list` from reaching the gateway. When that happens, `policy list` reports that OpenShell is unavailable and does not claim a local policy fallback. If the container runtime is down, it prints that runtime's recovery guidance instead. Stop while `policy list` cannot read OpenShell because the command cannot show whether the gateway applied the mutation. Restore gateway access and run `policy list` again before you resume or start fresh onboarding.
After `policy list` reads the live OpenShell policy, compare that one authoritative result with the change you requested. If the requested addition or removal is already present, do not submit it again. If it is absent, retry the convenience command after OpenShell connectivity is stable. The [failed-session recovery steps](#previous-onboarding-session-failed) provide the resume and fresh onboarding commands for non-policy onboarding state.
### Previous onboarding session failed
If a previous `$$nemoclaw onboard` attempt fails partway through (for example, a provider or inference-setup step reporting an error), NemoClaw records the failure in `~/.nemoclaw/onboard-session.json`.
The resume and fresh-onboarding commands below apply to ordinary onboarding failures.
For an interrupted rebuild, do not manually delete `onboard-session.json` or `.onboard-rebuild-.json` files from the selected state directory.
Rebuild recovery can be retained separately while another sandbox owns the active session.
Follow [Continue an Interrupted Replacement](../manage-sandboxes/operate-sandboxes/recover-and-rebuild-sandboxes#continue-an-interrupted-replacement) to rerun the affected sandbox's rebuild with the same settings.
When you re-run the installer, it detects the failed session and does not silently retry it. Silent retry would loop on the same failure if your original choice, such as an unreachable provider, was the cause.
- In an interactive terminal, the installer prompts whether to resume the failed session or start fresh. Press `R` (or Enter) to retry the same session, or `f` to discard it and make fresh choices.
- In non-interactive mode (piped `curl | bash` with `NEMOCLAW_NON_INTERACTIVE=1`, CI, scripts), the installer refuses and exits with a non-zero status so a scripted re-run cannot loop. You must opt in to one of two paths explicitly:
Start over with new choices to discard the recorded session and provider/model selection.
```bash
curl -fsSL https://www.nvidia.com/nemoclaw.sh | bash -s -- --fresh
```
```bash
curl -fsSL https://www.nvidia.com/nemoclaw.sh | NEMOCLAW_AGENT=hermes bash -s -- --fresh
```
```bash
curl -fsSL https://www.nvidia.com/nemoclaw.sh | NEMOCLAW_AGENT=langchain-deepagents-code bash -s -- --fresh
```
Or use environment variables instead. Set them on the `bash` side of the pipe because only the right-hand process inherits them.
```bash
curl -fsSL https://www.nvidia.com/nemoclaw.sh | NEMOCLAW_FRESH=1 bash
```
```bash
curl -fsSL https://www.nvidia.com/nemoclaw.sh | NEMOCLAW_AGENT=hermes NEMOCLAW_FRESH=1 bash
```
```bash
curl -fsSL https://www.nvidia.com/nemoclaw.sh | NEMOCLAW_AGENT=langchain-deepagents-code NEMOCLAW_FRESH=1 bash
```
Retry the same session.
This is only useful if the original failure was transient, for example a network blip or a stopped Docker daemon, and not a wrong provider choice:
```bash
$$nemoclaw onboard --resume
```
```
OpenClaw resume does not repeat completed non-secret sandbox, web search, messaging, or resource choices. Resume also reuses registered web search and messaging credentials when the same onboarding session recorded their successful OpenShell registration and OpenShell still reports the exact expected name, type, and credential keys. If the session lacks that registration receipt, the provider is missing, or its binding does not match, interactive resume requests the credential again; non-interactive resume preserves the completed choice, reports the required environment variable, and exits so you can export it before retrying `$$nemoclaw onboard --resume`.
### Kubernetes namespace not ready
If onboarding fails with `Kubernetes namespace not ready`, a previous failed or interrupted setup may have left stale OpenShell or NemoClaw state behind. Clean up the failed installation before re-running the installer:
```bash
$$nemoclaw uninstall --yes
curl -fsSL https://www.nvidia.com/nemoclaw.sh | bash
```
The normal uninstall path keeps user data under `~/.nemoclaw/`, including sandbox registry metadata, backups, and saved credentials unless you explicitly remove them. If `$$nemoclaw uninstall` reports that the local uninstall script is missing, follow the CLI's security boundary: download the versioned NVIDIA/NemoClaw tag URL that it prints, inspect the script locally, run that local copy, and then retry the installer.
```bash
curl -fsSLo uninstall.sh
less uninstall.sh
bash uninstall.sh --yes
curl -fsSL https://www.nvidia.com/nemoclaw.sh | bash
```
### Onboarding fails because the inference route served no model catalog
Onboarding verifies the sandbox inference route before it reports the deployment healthy.
When `https://inference.local/v1/models` keeps returning `HTTP 404`, the route answered but served no model catalog, so nothing validated your selected model against the provider.
Onboarding retries for the usual startup budget, then marks the `inference` check failed and exits nonzero instead of reporting a verified deployment.
Re-check the route and the recorded provider and model:
```bash
$$nemoclaw status
```
`$$nemoclaw status` applies the same HTTP 404 rule, so onboarding and status agree about the model catalog.
Confirm the provider and model are configured for this sandbox and that the endpoint serves `/v1/models`.
This result is resumable, so correct the endpoint and continue the recorded session:
```bash
$$nemoclaw onboard --resume
```
## Runtime
### OpenShell gateway and OpenClaw gateway startup order
NemoClaw uses two gateway layers for OpenClaw sandboxes:
- The OpenShell gateway runs on the host side and owns sandbox lifecycle, provider routes, port forwards, and `openshell sandbox list` / `status` queries.
- The OpenClaw gateway runs inside the sandbox container and serves the OpenClaw dashboard, agent API, and sub-agent WebSocket traffic.
Start and recover them in this order: container runtime, OpenShell gateway, sandbox container, then the in-sandbox OpenClaw gateway.
Do not start the OpenClaw gateway by hand before the OpenShell gateway is
healthy. NemoClaw cannot select, inspect, or reconnect the sandbox until
OpenShell can see the owning gateway.
If the host rebooted or the OpenShell gateway is down, first run:
```bash
$$nemoclaw status
```
The status command selects or starts the sandbox's recorded OpenShell gateway when possible, then checks whether OpenShell can still see the sandbox.
If OpenShell reports the registered sandbox missing, status leaves any provider container untouched and reports the missing state.
If you stopped the sandbox intentionally, run `$$nemoclaw start`.
After the sandbox is visible again, use `$$nemoclaw recover` to verify supported state for a running in-sandbox OpenClaw gateway and repair host forwards.
If the built-in native agent gateway is stopped, `recover` reports it and does not relaunch it.
Use `$$nemoclaw gateway restart` when you need a supported in-sandbox gateway restart or runtime configuration reload.
### Reconnect after a host reboot
After a host reboot, the container runtime, OpenShell gateway, and sandbox may not be running. Follow these steps to reconnect.
1. Start the container runtime.
- **Linux:** start Docker if it is not already running (`sudo systemctl start docker`)
- **macOS:** open Docker Desktop or start Colima (`colima start`)
1. Check the managed OpenShell gateway service.
If a custom-port gateway is NemoClaw-managed, skip this service check and continue with the NemoClaw recovery step using the same environment value. Only the default port `8080` uses a NemoClaw-managed service. If the gateway is externally supervised, inspect and restart it through the supervisor declared by `NEMOCLAW_GATEWAY_MANAGEMENT`, regardless of port.
On Apple Silicon macOS with Homebrew, let NemoClaw inspect and restart the official formula service. NemoClaw runs each Homebrew operation inside the checksum-verified temporary trust boundary. Continue to the NemoClaw recovery step below instead of running `brew services` directly.
NemoClaw verifies the staged formula checksum and temporarily trusts only `nvidia/openshell/openshell` around each Homebrew inspection, start, or stop operation. If formula verification fails or Homebrew cannot grant or remove temporary trust, rerun the standard NemoClaw installer and then rerun onboarding:
```bash
curl -fsSL https://www.nvidia.com/nemoclaw.sh | bash
```
The standalone gateway is selected only when Homebrew or both the staged formula and installed keg are absent. When the formula or keg exists, Homebrew remains the lifecycle authority. NemoClaw does not switch to the standalone gateway after a Homebrew inspection, start, or stop failure. Follow the reported repair guidance instead of changing service ownership manually. If the installed service fails inspection, startup, or its health check, NemoClaw prints this log command:
```bash
tail -n 200 "$(brew --prefix)/var/log/openshell/openshell-gateway.out.log" "$(brew --prefix)/var/log/openshell/openshell-gateway.err.log"
```
Rerun the NemoClaw installer to restore the managed service for later onboarding runs.
On Linux package installs, inspect and restart the upstream service.
```bash
systemctl --user status openshell-gateway
systemctl --user restart openshell-gateway
```
If the service fails inspection, startup, or its health check, NemoClaw prints this log command:
```bash
journalctl --user --unit openshell-gateway --no-pager --lines=200
```
On Linux tarball installs, inspect and restart the marked NemoClaw service.
```bash
systemctl --user status nemoclaw-openshell-gateway
systemctl --user restart nemoclaw-openshell-gateway
```
If the service fails inspection, startup, or its health check, NemoClaw prints this log command:
```bash
journalctl --user --unit nemoclaw-openshell-gateway --no-pager --lines=200
```
The tarball unit is under `$XDG_CONFIG_HOME/systemd/user`, or `~/.config/systemd/user` when `XDG_CONFIG_HOME` is not absolute. It starts with your user session; NemoClaw does not enable lingering. On Linux, NemoClaw attempts the standalone fallback when a managed service fails inspection, startup, or its health check. The standalone gateway starts only after NemoClaw verifies exclusive ownership of the gateway port. The fallback does not bypass managed-service trust validation or unsafe environment configuration. These conditions remain hard failures:
- Homebrew formula identity query, metadata, or official-tap validation errors
- Foreign or symlinked systemd units, or an untrusted systemd executable identity
- An invalid `DOCKER_HOST` or a symlinked service environment file
Without Homebrew on macOS, or without a reachable systemd user manager on Linux, continue with the standalone recovery step below.
1. Check sandbox state.
```bash
openshell sandbox list
```
If the sandbox shows `Ready`, skip to step 5.
1. Recover the managed gateway (if needed).
If the sandbox is not listed after the service restart, or `systemctl --user` is unavailable, first ask NemoClaw to reconnect through the recorded sandbox:
```bash
$$nemoclaw status
```
If that cannot restore the gateway registration, resume onboarding to recreate the managed gateway metadata:
```bash
$$nemoclaw onboard --resume
```
Wait a few seconds, then re-check with `openshell sandbox list`. On Docker-driver hosts, NemoClaw also looks for OpenShell-labeled sandbox containers when the gateway is healthy but reports the sandbox as missing. It can start a stopped labeled container, or restore the latest GPU-backup sibling container name and start it.
1. Reconnect.
```bash
$$nemoclaw connect
```
The gateway usually rotates its SSH host keys across a reboot. `connect` detects the resulting identity drift, prunes the stale `openshell-*` entries from `~/.ssh/known_hosts`, and retries automatically. You do not need to edit `known_hosts` by hand or re-run `$$nemoclaw onboard` in this case.
1. Start host auxiliary services (if needed).
If you use the cloudflared tunnel started by `$$nemoclaw tunnel start`, start it again:
```bash
$$nemoclaw tunnel start
```
OpenShell-managed channel messaging is configured during onboarding, not through a separate bridge process from `$$nemoclaw tunnel start`.
Use `$$nemoclaw channels list` to inspect the configured channels and [Choose Messaging Channels](../manage-sandboxes/messaging-channels/choose-messaging-channels) for each channel's support status.
To pause a single bridge without destroying the sandbox, use `$$nemoclaw channels stop `.
If the sandbox remains missing after restarting the gateway, its authoritative OpenShell policy and live workspace are unavailable, so `rebuild --yes` cannot recreate it from NemoClaw registry metadata. Run `$$nemoclaw destroy --yes` to remove the stale local entry, then run `$$nemoclaw onboard`. The missing sandbox's state cannot be recovered unless you already have a separate snapshot; restore that snapshot explicitly after onboarding. For details, refer to [Create and Restore Snapshots](../manage-sandboxes/state-and-backups/create-and-restore-snapshots).
### Gateway Port Stays Bound After Destroying the Last Sandbox
When the default-port gateway runs under the packaged OpenShell gateway service, destroying the final sandbox with `--cleanup-gateway` stops that service before it reaps host gateway processes whose ownership NemoClaw verifies. This does not prove that the port is free: a foreign, ambiguous, or other-user listener may remain. Verify the gateway port after cleanup. The service is stopped, not disabled or removed, and the next onboarding run starts it again. If the systemd user manager is unavailable and the service cannot activate automatically, `destroy` can instead stop the recorded standalone gateway after verifying its exact process identity and port. For other service-stop failures, `destroy` exits non-zero and prints the status command for the service.
If `destroy` prints `Shared NemoClaw gateway left running`, NemoClaw skipped gateway cleanup because `openshell sandbox list` still reported a live sandbox after the post-delete wait, or because that command failed.
The message names those sandboxes or the failed command.
First confirm that the gateway reports no sandboxes:
```bash
openshell sandbox list -g
```
When the list is empty, remove the gateway:
```bash
openshell gateway remove
```
Use the gateway name from the `destroy` message in both commands.
On Apple Silicon macOS with Homebrew, rerun the standard NemoClaw installer to restore the pinned formula and temporary trust contract:
```bash
curl -fsSL https://www.nvidia.com/nemoclaw.sh | bash
```
Then rerun `$$nemoclaw destroy --cleanup-gateway` instead of stopping the Homebrew service directly.
On Linux, stop the service yourself, then rerun `destroy`. Use the service name that matches the install. For package installs:
```bash
systemctl --user stop openshell-gateway
```
For tarball installs:
```bash
systemctl --user stop nemoclaw-openshell-gateway
```
A successful stop is not proof that the port is free. When a managed service fails inspection, startup, or its health check, NemoClaw starts the gateway standalone instead, and the service stays installed but inactive. Stopping an inactive service exits `0` without stopping that standalone gateway. Set `gateway_port` below to the affected gateway port, then confirm the port after the stop:
```bash
gateway_port=PORT
ss -lntp | grep ":${gateway_port}"
```
If a process still holds the port, do not infer ownership from the `openshell-gateway` process name. Rerun the identity-aware cleanup path:
```bash
$$nemoclaw destroy --cleanup-gateway
```
NemoClaw verifies the listener before it reaps host gateway processes. If it reports a foreign or ambiguous owner, use the fresh listener details to identify the owning installation or service and stop it through that owner's lifecycle command. Do not uninstall a selected NemoClaw environment based only on the process name; uninstall removes that environment's gateway state, credentials, and registrations.
### `gateway restart` or `recover` reports `privileged control unavailable`
Built-in OpenClaw and Hermes lifecycle commands require a running sandbox that belongs to the named NemoClaw registry entry. NemoClaw invokes the agent's native lifecycle command through the exact selected OpenShell sandbox and then observes native health and host forwards.
Current built-in images preserve the agent startup wiring and OpenShell credential projection while leaving gateway restart, background work, cron, and child-agent lifecycle with the native agent. Custom images must preserve those attachment and credential contracts without adding a second NemoClaw lifecycle controller.
On a direct-container deployment, first confirm that the sandbox is running:
```bash
$$nemoclaw status
```
For a built-in OpenClaw or Hermes sandbox, `gateway restart` invokes the native
agent lifecycle command inside the exact selected OpenShell sandbox. NemoClaw
then reports native command output and observes gateway health; it does not
authorize, respawn, or quarantine the process.
If the native gateway is stopped, restart it through the agent or restart the
sandbox through OpenShell. If the image itself is incompatible, rebuild it:
```bash
$$nemoclaw rebuild --yes
```
Custom images must preserve the current agent startup wiring and OpenShell
credential projection, but must not add the retired NemoClaw gateway controller.
### Hermes config reports a mutation already in progress
Hermes host config writes and lifecycle seals use the same root-only mutation lock.
A host-side config write binds to the SHA-256 digest of the matching read, installs fresh config inodes, refreshes both config hashes, and restores mutable access before releasing that lock.
If a concurrent command reports `Hermes config mutation is already in progress`, do not remove `/run/nemoclaw/hermes-config-mutation.lock` manually.
Wait for the active config, recovery, or restart command to finish, then retry.
### Hermes startup reports `HERMES_CONFIG_MUTATION_ORPHANED`
This refusal means a root-owned config transaction stopped without a safely recoverable complete state.
NemoClaw does not treat a leftover lock file as permission to continue or adopt partially written config.
Do not delete `/run/nemoclaw/hermes-config-mutation.lock`, the restart state, or the persistent transaction marker manually.
An ordinary in-place `rebuild` can stop when it cannot create a trustworthy backup of the interrupted config state.
If you already have a trusted host-side snapshot, record its selector, destroy the sandbox, re-onboard the same sandbox name from trusted host configuration, and then restore that snapshot:
```bash
$$nemoclaw snapshot list
$$nemoclaw destroy
$$nemoclaw onboard --name --agent hermes
$$nemoclaw snapshot restore
```
Destroying the sealed sandbox permanently discards any state newer than the selected snapshot, so verify that the host-side snapshot exists before confirming destruction. Without a trusted pre-incident snapshot, automatic state-preserving recovery is not available. Recreate the sandbox from host-side onboarding configuration only if you accept losing the inaccessible in-sandbox state.
### Hermes startup reports `HERMES_RESTART_SEAL_ORPHANED`
This refusal means a container recreation discarded the root-only restart metadata under `/run` while the persistent `/sandbox` tree still carries the temporary restart transaction marker, or a non-root entrypoint found a root transaction it cannot safely restore.
NemoClaw cannot safely infer the pre-transaction ownership, modes, or inode flags, so startup fails closed even when `config.yaml` and `.env` still look readable.
Do not repair the paths with manual `chown`, `chmod`, or hash regeneration.
Restore a trusted snapshot into a new sandbox as shown above, or recreate from host-side onboarding configuration if no state-preserving recovery is required.
### Hermes startup refuses an unsafe state directory
Hermes stops startup when `/sandbox/.hermes/sessions`, `/sandbox/.hermes/gateway`, or `/sandbox/.hermes/runtime` is a symbolic link, a file, or cannot be repaired safely.
Startup also stops if the config root or its `hooks`, `image_cache`, or `audio_cache` directories cannot be repaired safely, or if any of those paths changes during descriptor-anchored repair.
The same startup stops when the `logs` directory tree contains a symbolic link, a hard-linked file, another unsafe entry, or changes during descriptor-anchored repair.
It also refuses to repair more than 4,096 retained log entries or a log tree deeper than 64 directories in one startup.
Startup also stops when `/sandbox/.hermes/.hermes_history` is a symbolic link, a nonregular or hard-linked file, or cannot be repaired safely.
`/tmp/nemoclaw-start.log` names the reported state item and reports `Hermes pre-launch layout repair failed`.
When a host-side control command fails, its output can include matching startup diagnostics as `NEMOCLAW_START_LOG=...`.
If the diagnostic reports the maximum log-entry count or repair depth, archive or remove old retained logs from a trusted host-side recovery environment and retry startup.
For symbolic links, hard links, unsafe entry types, or paths that changed during repair, do not run `chown`, `chmod`, or another in-place repair on the reported path.
For those unsafe-path failures, follow [Hermes startup reports HERMES_CONFIG_MUTATION_ORPHANED](#hermes-startup-reports-hermes_config_mutation_orphaned) to restore a trusted host-side snapshot into a recreated sandbox.
If no trusted snapshot exists for an unsafe-path failure, recreate the sandbox from host-side onboarding configuration only if you accept losing its inaccessible state.
### Sandbox is running an outdated agent version
After upgrading NemoClaw, `$$nemoclaw connect` and `$$nemoclaw status` warn if the sandbox is running an older agent version than the current image.
To upgrade the sandbox while preserving workspace state, run:
```bash
$$nemoclaw rebuild
```
The rebuild command backs up state, destroys the old sandbox, recreates it with the current image, and restores state. Create a snapshot before rebuilding if you want an additional safety net:
```bash
$$nemoclaw snapshot create
$$nemoclaw rebuild
```
### Snapshot Sanitization Requires a Verified Python Interpreter
The CLI operations listed below stop when NemoClaw cannot resolve a verified interpreter. The error begins with this text:
```text
python3 is required for snapshot sanitization; install python3 and rerun
```
NemoClaw removes credentials from copied state with an isolated `python3` helper. The operation fails closed when that helper cannot run. These operations can report the error:
- `$$nemoclaw snapshot create`
- `$$nemoclaw rebuild`, including rebuilds started by `$$nemoclaw upgrade-sandboxes`
- `$$nemoclaw backup-all`, including the installer's pre-upgrade backup and an eligible stopped Docker-driver sandbox
A host that already has `python3` can still report this message. NemoClaw does not search `PATH` for this credential-bearing helper. NemoClaw accepts `python3` only at these locations:
- `/usr/bin/python3`
- `/usr/local/bin/python3`
- `/opt/homebrew/bin/python3`
- `/opt/local/bin/python3`
- A `python3` executable beside the canonical Node.js executable
NemoClaw rejects a candidate that fails its ownership, permission, or executable checks. For more information about the interpreter requirement, refer to [Prerequisites](../get-started/prerequisites).
If NemoClaw reports that it removed the incomplete snapshot, install or repair `python3` at an accepted location. Then rerun the complete command.
If cleanup fails, treat the reported directory as retained until you confirm that it is absent.
A retained incomplete snapshot can contain unsanitized credentials. The
directory must remain owner-only. You must not restore, copy, or share the
directory. You must remove the directory only by the exact path that NemoClaw
reports. You must confirm that the exact path no longer exists before you
rerun the complete command.
### Sandbox shows as stopped
When status reports `sandbox_container_stopped`, Docker still has a container for the sandbox, but the container is not running. Use the lightest recovery path first instead of rebuilding immediately.
1. Confirm Docker can still see the labeled container.
```bash
docker ps -a --filter "label=openshell.ai/sandbox-name="
```
1. Run recovery from the host.
```bash
$$nemoclaw recover
```
For a stopped standard sandbox, `recover` asks OpenShell to start that exact sandbox before it waits for readiness. If provider state appears active while OpenShell reports `Stopped`, recovery still submits the start through OpenShell so one lifecycle owner repairs the phase. If OpenShell cannot start the sandbox, recovery reports the failure and does not fall back to a direct provider-container mutation.
```bash
$$nemoclaw status
```
Deep Agents Code does not provide the `recover` command. Use `$$nemoclaw start` to start its existing container.
1. Check status from the host.
```bash
$$nemoclaw status
```
On Docker-driver hosts, a missing OpenShell sandbox remains a missing result even when Docker still has a labeled container. Status does not start, unpause, rename, or replace that container; use the reported OpenShell recovery guidance.
1. Rebuild only if the sandbox cannot be restarted but OpenShell still reports the live sandbox and its current policy:
```bash
$$nemoclaw rebuild --yes
```
Rebuild captures the live sandbox policy and workspace before recreation. If the live sandbox is absent, use the clean-replacement sequence below instead.
### Sandbox is registered locally but missing from the gateway
After a gateway restart, host reboot, or manual OpenShell cleanup, NemoClaw may still have a local registry entry for a sandbox that the live gateway no longer lists. `$$nemoclaw status` and `$$nemoclaw connect` preserve that local registry entry and print recovery guidance instead of deleting it automatically. First restore the gateway and retry status in case the sandbox returns. If it remains absent, run `$$nemoclaw destroy --yes` to remove the stale entry, then run `$$nemoclaw onboard`. Rebuild cannot recreate the missing sandbox because no authoritative OpenShell policy remains. State is recoverable only from a separate snapshot restored after onboarding.
### A command reports that the registry file is not valid JSON
Registry operations that require complete sandbox records, such as `$$nemoclaw list` and `$$nemoclaw onboard`, stop with `Configuration file is present but is not valid JSON`, followed by the path to `sandboxes.json` and the recovery commands.
These operations stop instead of reading the file as an empty registry, so they cannot replace your sandbox records with empty state. Optional messaging health checks omit registry-derived information when they cannot read the registry. It does not rename, move, or rewrite the file.
Follow [Malformed Registry File](host-files-and-state#malformed-registry-file) to keep a copy and remove it.
### Status shows "not running" inside the sandbox
This is expected behavior. When checking status inside an active sandbox, host-side sandbox state and inference configuration are not inspectable. The status command detects the sandbox context and reports "active (inside sandbox)" instead.
Run `openshell sandbox list` on the host to check the underlying sandbox state.
## Deep Agents
### `dcode status` reports a stale inference route
The managed `dcode` runtime reads provider and model settings from `/sandbox/.deepagents/config.toml`. NemoClaw generates that file during onboarding and rebuilds, and Deep Agents provider or model changes require fresh named recreation rather than `inference set`.
Check the host-recorded route first:
```bash
nemo-deepagents status
```
Then connect and compare the in-sandbox identity:
```bash
nemo-deepagents connect
dcode status
```
If the provider or model does not match the host route, rebuild the sandbox to regenerate `config.toml` from recorded metadata:
```bash
nemo-deepagents rebuild
```
To intentionally switch provider or model, recreate the named sandbox with fresh onboarding:
```bash
nemo-deepagents onboard --fresh --name --recreate-sandbox
```
### Trusted route-probe helper is missing
Deep Agents sandboxes created by older NemoClaw images may not contain the image-owned `/usr/local/lib/nemoclaw/dcode-managed-exec` helper. Current `connect`, `status`, and `doctor` fail closed when that helper is missing because NemoClaw cannot run the authoritative `inference.local` route probe safely.
If the output says the trusted Deep Agents Code route-probe helper is missing, rebuild the sandbox with the current image and retry the command:
```bash
nemo-deepagents rebuild
nemo-deepagents status
```
### `dcode` refuses to start because upstream auth state exists
The managed launchers refuse upstream credential state inside `/sandbox/.deepagents/.state/auth.json` and `/sandbox/.deepagents/.state/chatgpt-auth.json`. Those files can contain provider credentials or OAuth state that bypasses NemoClaw's host-owned credential boundary.
Remove the upstream auth state from the sandbox, then start `dcode` again:
```bash
rm -f /sandbox/.deepagents/.state/auth.json /sandbox/.deepagents/.state/chatgpt-auth.json
dcode status
```
Do not put provider credentials in `/sandbox/.deepagents/.env`, project `.env` files, or Deep Agents config files. Register credentials with NemoClaw or OpenShell on the host so the gateway can inject them at egress.
### Managed MCP commands report an older Deep Agents runtime
For the current rebuild and retry workflow, refer to [Agent MCP Capability Is Missing](troubleshoot-mcp-servers#agent-mcp-capability-is-missing).
### Tavily remains blocked after opt-in
Deep Agents does not have a NemoClaw-managed web-search feature. The Tavily flow only opens Python egress for project code or manually configured tools that call Tavily.
Confirm that the target sandbox has the `tavily` preset applied:
```bash
nemo-deepagents policy list
```
If it is missing, apply the preset, register the host-side credential, and rebuild so the provider attaches:
```bash
nemo-deepagents policy add tavily --yes
export TAVILY_API_KEY=tvly-...
nemo-deepagents credentials add tavily-search --type tavily --agent dcode --credential TAVILY_API_KEY
unset TAVILY_API_KEY
nemo-deepagents rebuild
```
If Tavily is still blocked after rebuild, inspect recent policy denials:
```bash
nemo-deepagents logs --tail 50
```
The `tavily` preset is a managed-Python opt-in. It is process-wide for sandbox Python and is not a `dcode`-only boundary.
### Deep Agents read-only path checks fail
The Deep Agents image keeps the managed Python environment under `/opt/venv` read-only and leaves `/sandbox` writable for project and agent state. Writes under `/usr`, `/etc`, or `/opt/venv` should fail, while writes under `/sandbox` and `/tmp` should work.
If a read-only path probe reports that protected paths are writable, rebuild with the current NemoClaw image:
```bash
nemo-deepagents rebuild
```
If startup reports that Landlock enforcement is unavailable, the Deep Agents sandbox fails closed instead of running with reduced filesystem enforcement. Deep Agents uses `compatibility: strict` for its managed filesystem policy, so kernels older than 5.13 or VM-backed Docker runtimes without Landlock support can block sandbox creation. Move the sandbox to a Linux kernel and container runtime that support Landlock, then rerun onboarding or rebuild the sandbox.
### Git clone fails with a certificate verification error
In networks that inspect TLS, OpenShell injects a proxy CA bundle into the sandbox. Current NemoClaw exports that bundle as `GIT_SSL_CAINFO` during sandbox startup and persists it for `$$nemoclaw connect` sessions, so Git can trust the proxy CA. It also forwards standard CA bundle variables for subprocesses, including `GIT_SSL_CAPATH`, `CURL_CA_BUNDLE`, and `REQUESTS_CA_BUNDLE`.
If Git still reports `server certificate verification failed`, reconnect to the sandbox and check that the CA variables are present:
```bash
env | grep -E 'SSL_CERT_FILE|GIT_SSL_CAINFO|CURL_CA_BUNDLE|REQUESTS_CA_BUNDLE' || true
```
`grep` exits non-zero when it finds no matches, so empty output (with the trailing `|| true`) simply means none of these CA variables are set in the current shell.
If they are missing on an older sandbox, upgrade NemoClaw and run:
```bash
$$nemoclaw rebuild
```
### External channel TLS fails behind a corporate MITM proxy (`NET:FAIL`)
On networks where a corporate proxy re-signs external TLS, endpoints such as `api.telegram.org` can fail certificate verification even when network policy allows the connection. Logs show the request opening as `NET:OPEN ... api.telegram.org:443` followed by `NET:FAIL` because the corporate root is not in the OpenShell trust path.
Provide the corporate CA before onboarding, then onboard or rebuild the sandbox.
```bash
export NEMOCLAW_CORPORATE_CA_BUNDLE=/path/to/corporate-ca.pem
$$nemoclaw onboard
```
Refer to [Configure Corporate CA Trust](../security/configure-corporate-ca-trust) for source precedence, host anchor discovery, image and runtime trust, custom Dockerfile requirements, import validation, and `NEMOCLAW_CORPORATE_CA_IMPORT=0`. If the import does not occur, check onboarding output for `baking corporate proxy CA from ...` or a warning that the selected source was skipped.
### A request inside the sandbox fails with `CONNECT tunnel failed, response 403`
Sandbox outbound network access is denied by default and enforced by the OpenShell proxy. When a request targets a host that no applied policy preset allows, the proxy refuses the tunnel and tools surface only the protocol-level error:
```text
fatal: unable to access 'https://example.com/foo/bar/': CONNECT tunnel failed, response 403
curl: (56) CONNECT tunnel failed, response 403
```
This is a network-policy denial, not a tool or certificate problem.
When you run a command through `$$nemoclaw exec -- ...` and it exits non-zero, NemoClaw checks the sandbox audit log for a policy denial recorded after the command started. If it finds one, it appends a short breadcrumb to stderr after the tool's own output, naming the denied `host:port` when it can be extracted safely and showing the commands below:
```text
curl: (56) CONNECT tunnel failed, response 403
$$nemoclaw: recent network policy denial detected for example.com:443 inside sandbox 'oc-fresh'.
The sandbox's egress policy blocked this request; the tool above only saw the proxy's 403.
See the denied flow: $$nemoclaw oc-fresh logs --tail 50
Review applied presets: $$nemoclaw oc-fresh policy list
Allow the host: $$nemoclaw oc-fresh policy add
Silence this hint: export NEMOCLAW_NO_POLICY_HINT=1
```
If the log probe fails, no breadcrumb is added. Unsafe endpoint data is omitted from the breadcrumb, and an invalid sandbox name is shown as ``.
An IPv6 target is named in its RFC 3986 bracketed form:
```text
$$nemoclaw: recent network policy denial detected for [2001:db8::1]:443 inside sandbox 'oc-fresh'.
```
The tool's own stdout/stderr bytes and its exit code are left unchanged. The breadcrumb is printed by the host CLI after the command finishes, and only for a genuine failure with a fresh denial. A command that succeeds, or one that fails for an unrelated reason, prints no breadcrumb. Set `NEMOCLAW_NO_POLICY_HINT` to any non-empty value other than `0` or case-insensitive `false` (for example, `1`, `true`, `TRUE`, `yes`, or `YES`) to suppress it entirely.
The first interactive `$$nemoclaw connect` shell also prints a one-line reminder of this denial signature and the `logs` command below. The reminder is shown once per top-level interactive session, and only when all of these hold: an egress proxy is configured, the shell is interactive with a terminal attached to stderr, and it is a top-level shell (not a nested subshell or pane). Suppress it with `NEMOCLAW_NO_POLICY_HINT=1`. The reminder names the sandbox when NemoClaw receives a valid sandbox name during sandbox creation. If no valid name is available, it shows ``; run `$$nemoclaw list` to see your sandbox names. If the reported sandbox name contains characters that are not valid in a sandbox name (uppercase letters, underscores, control characters, and similar) or exceeds 19 characters, the reminder shows the `` placeholder for safety rather than echoing the untrusted value. The reminder is intentionally proactive: the denial itself is surfaced by the OpenShell proxy, so the `curl`/`git` error text is left unchanged and the reminder points you to the logs instead.
To see which rule denied the request, read the merged logs from the host:
```bash
$$nemoclaw logs --tail 50
```
If the host should be reachable, allow it with a preset or a [custom
preset](../network-policy/customize-network-policy#custom-preset-files):
If the host should be reachable, allow it with a built-in preset or apply a
[reviewed custom preset
file](../network-policy/configure-policies/create-custom-policy-presets) from
the host:
```bash
$$nemoclaw policy add
```
```bash
nemo-deepagents policy add --from-file ./my-preset.yaml --yes
```
Replace `` with a real preset name such as `github`, `pypi`, or `npm`. Run `$$nemoclaw policy add` with no preset to list the available presets.
### An `openclaw` command inside the sandbox fails with `scope upgrade pending approval`
The OpenClaw gateway refuses a command when the device asks for more scopes than are currently approved. The failure names the request id but not the command that clears it:
```text
gateway connect failed: GatewayClientRequestError: scope upgrade pending approval
(requestId: 4899d110-911f-4bc7-ac1a-85d76c7b366f)
```
Record the `requestId` from this native failure.
Open the prepared shell. During setup, `connect` makes a bounded, best-effort
attempt to settle pending requests from the `cli`, `openclaw-cli`, and
`openclaw-control-ui` clients when the requested scopes are limited to pairing,
read, and write. This pass uses the client ID and scopes recorded in the pending
request; it does not authenticate that client metadata. OpenClaw's gateway and
canonical approval operation remain the authorization boundary. If the pass
times out or an approval operation fails, an eligible request can remain
pending:
```bash
$$nemoclaw connect
```
Replace `` with the sandbox name from the failed command.
First rerun the original command unchanged. If it succeeds, no manual approval
is required. If the command still reports a pending approval, such as a request
for `operator.admin`, list the remaining requests in the prepared shell:
```bash
openclaw devices list --json
```
Find the pending entry whose `requestId` exactly matches the native failure.
Confirm that entry's requesting client, device, and requested scopes, then
approve only that same `requestId`:
```bash
openclaw devices approve
```
Then rerun the original command again.
The `exec` command streams the native command output and normally returns its
native exit status. After `exec` fails, NemoClaw does not inspect pending
requests automatically. The `connect` flow attempts
automatic settlement only for allowlisted client IDs requesting pairing, read,
and write scopes. Its helper does not authenticate client metadata; requests
that remain pending are available for manual review of the client, device,
scopes, and matching request ID.
### Onboarding or Sandbox Creation Reports a TLS Certificate Error
Use this recovery when the preserved OpenShell detail for global policy inspection or sandbox creation identifies a gateway certificate-trust failure during onboarding.
The OpenShell gateway certificate may have changed since the CLI last trusted it.
Refresh trust for the registered gateway, then resume onboarding:
```bash
openshell gateway trust -g nemoclaw
$$nemoclaw onboard --resume
```
### `openclaw update` hangs or times out inside the sandbox
This is expected for the current NemoClaw deployment model. NemoClaw installs `openclaw` into the sandbox image at build time, so the CLI is image-pinned rather than updated in place inside a running sandbox.
Do not run `openclaw update` inside the sandbox. Instead:
1. Upgrade to a NemoClaw release that includes the newer `openclaw` version.
2. If you build NemoClaw from source, bump the pinned `openclaw` version in `Dockerfile.base` and rebuild the sandbox base image.
3. Run `$$nemoclaw rebuild` to recreate the sandbox with the updated image. The rebuild command automatically backs up workspace state before destroying the old sandbox and restores it afterward.
### AWS EC2 Instance-Role Credential Discovery Is Unavailable
This is expected in an OpenClaw sandbox. NemoClaw forces `AWS_EC2_METADATA_DISABLED=true` because OpenShell blocks the link-local EC2 Instance Metadata Service endpoint. Existing sandboxes must use a current NemoClaw image before they receive this environment invariant.
Upgrade NemoClaw and rebuild the sandbox:
```bash
$$nemoclaw update --yes
$$nemoclaw rebuild --yes
```
Reconnect and verify the value:
```bash
$$nemoclaw connect
printenv AWS_EC2_METADATA_DISABLED
```
Expected output:
```text
true
```
Only EC2 instance-role discovery is disabled. Static access keys, bearer tokens, shared profiles, SSO and process credentials, web identity, and ECS container credentials remain eligible. NemoClaw's host-local Amazon Bedrock adapter is outside the sandbox credential-discovery boundary and remains available. Do not add `169.254.169.254` to a network policy or override the variable.
### Inference requests time out
Verify that the inference provider endpoint is reachable from the host. Check the active provider and endpoint:
```bash
$$nemoclaw status
```
The main `Inference` line probes `https://inference.local/v1/models` from inside the sandbox and then sends an inference request over the same route, so it reflects the route the agent uses. When that request returns a transient gateway status, `status` repeats both probes up to three total attempts before it reports a failure. It prints the HTTP status, next attempt, and two-second delay to stderr before each retry; the configured timeout envelope for three complete probe pairs and two delays is about 124 seconds. If that line shows `unauthorized`, `unhealthy`, `unreachable`, or `not probed`, inspect the labeled diagnostic lines to identify the failing hop. An `unauthorized` line means the route answered but rejected the request, so refresh the provider credential rather than the route.
For local Ollama and local vLLM, `Inference (ollama backend)` or the corresponding local-backend line reports the host-side service separately. If a local-backend diagnostic fails, start the backend. For Local Ollama, current releases can also print `Inference (auth proxy)` when a proxy token is available. Docker Desktop on Windows Subsystem for Linux (WSL) reaches host loopback directly, so its auth-proxy diagnostic is non-authoritative. Native Docker Engine inside WSL is unqualified; enable Docker Desktop WSL integration, then rerun onboarding.
On non-WSL hosts that front Ollama with the auth proxy, run `$$nemoclaw connect` when the auth-proxy result is `unreachable` to restart the proxy with its persisted token. If the result is `unauthorized`, rerun onboarding to rotate the rejected proxy token and refresh the route. Follow the printed detail for any other auth-proxy failure. For Ollama-backed OpenClaw sandboxes, agent passthrough uses the registered host route to warm an unloaded model after an Ollama daemon restart. If that bounded warm-up fails or times out, NemoClaw reports the result and continues so OpenClaw can emit its canonical backend error.
If the endpoint is correct but requests still fail, check for network policy rules that may block the connection. Then verify the credential and base URL for the provider you selected during onboarding.
If you entered an AWS Bedrock Runtime URL such as `https://bedrock-runtime.us-east-1.amazonaws.com` in the **Other Anthropic-compatible endpoint** flow, NemoClaw auto-detects it and routes sandbox traffic through a host-local adapter. Use the raw Bedrock Runtime host, not an Anthropic `/v1/messages` path, and verify that the model ID or inference profile ID is valid for that region. For auth, export `AWS_BEARER_TOKEN_BEDROCK`, `AWS_PROFILE`, or standard IAM environment credentials before onboarding; if you paste a key at the `COMPATIBLE_ANTHROPIC_API_KEY` prompt, NemoClaw uses it only as the adapter's Bedrock bearer token. Region errors usually mean the pasted endpoint region, `AWS_REGION`, `AWS_DEFAULT_REGION`, or the model/inference profile ID do not match.
The adapter advertises model IDs to readiness checks only after that running adapter completes a successful inference request for them.
Onboarding performs this request before checking sandbox readiness; failed requests do not add a model to the catalog.
This uses the existing Bedrock Runtime inference permissions and does not query the Bedrock control-plane catalog.
For Ollama, vLLM, NIM, and compatible-endpoint inference validation, the default timeout is 180 seconds. The managed NIM startup health wait uses a separate 15-minute (900-second) default and still exits early if the container stops before it becomes healthy. On Docker 29.x or hosts using the containerd image store, managed NIM onboarding resolves and pulls the host-platform image digest when NGC exposes a multi-architecture image index. If you still see NGC repository-format or attestation errors, confirm Docker can run `docker manifest inspect` for the selected image and that you are logged in to `nvcr.io`. If large prompts still cause timeouts, increase it with `NEMOCLAW_LOCAL_INFERENCE_TIMEOUT` before re-running onboard:
```bash
export NEMOCLAW_LOCAL_INFERENCE_TIMEOUT=300
$$nemoclaw onboard
```
For local Ollama and vLLM, onboarding retries the container reachability check and can fall back to the host-side health check when the local backend is healthy. If Ollama times out during a cold model load, NemoClaw retries once with a 300-second probe budget before failing. If all attempts fail, the error includes container reachability diagnostics such as HTTP status and host gateway resolution.
`NEMOCLAW_LOCAL_INFERENCE_TIMEOUT` only covers the inference-server validation probe. The post-create readiness wait has its own budget (`NEMOCLAW_SANDBOX_READY_TIMEOUT`); refer to [Sandbox onboard times out with "did not become ready within Ns"](#sandbox-onboard-times-out-with-did-not-become-ready-within-ns) for the readiness path.
### Sandbox onboard times out with "did not become ready within Ns"
Onboarding ends with:
```text
Sandbox 'my-assistant' was created but did not become ready within 180s.
terminal_resolution: timed_out_retained
Recovery remains blocked while sandbox 'my-assistant' exists.
```
This is a separate budget from `NEMOCLAW_LOCAL_INFERENCE_TIMEOUT`. It covers the readiness wait that follows sandbox creation, including in-sandbox boot, OpenClaw start, and policy load. It does not cover the inference probe.
For a newly created OpenClaw or Hermes sandbox, `Ready` is not the final acceptance signal. Within this same budget, NemoClaw also requires OpenShell to return a durable sandbox ID and accept `openshell sandbox exec --name -- true`. NemoClaw keeps waiting only when OpenShell returns its exact `sandbox is not ready` response. A missing or malformed ID, or another command failure, stops the wait and preserves the sandbox. Inspect the gateway, then run `$$nemoclaw destroy` to check whether authoritative absence permits cleanup. Do not retry same-name onboarding while the retained sandbox exists.
The 180-second default fits typical workstations but can be exceeded when:
- The host is building or uploading the sandbox image for the first time (cold caches, slow link).
- The selected model is large (70B+ parameters or 4-bit/8-bit quantisations that take time to memory-map).
- Onboarding runs on a remote VM where image upload to the gateway streams over the network (for example DGX Station first-run installer).
Raise the budget before re-running onboard:
```bash
export NEMOCLAW_SANDBOX_READY_TIMEOUT=600
$$nemoclaw onboard
```
The variable accepts seconds and applies to the readiness wait only. When the ordinary create deadline expires, NemoClaw preserves the partially created sandbox because OpenShell deletion targets its mutable name. Same-name onboarding remains blocked while that sandbox exists. Run `$$nemoclaw destroy` to check retained recovery. Do not delete the sandbox by mutable name.
The failure path also differs when NemoClaw recreates an OpenShell-managed Docker runtime immediately before this wait. NemoClaw pins the exact OpenShell sandbox ID before recreation. Within the same deadline, NemoClaw requires two consecutive `Ready` observations that each confirm the exact ID and successful command execution. It retries only OpenShell's exact `sandbox is not ready` response. If the deadline expires, the ID changes, or another probe fails, NemoClaw preserves diagnostics and attempts to restore the pre-recreation Docker container. If restoration fails, NemoClaw reports that the sandbox and container state is uncertain. NemoClaw does not start dashboard or other host forwarding, and it does not delete a sandbox by its mutable name. It leaves the sandbox in place for inspection and recovery.
If readiness still fails after the extended budget, inspect the gateway and sandbox status:
```bash
openshell sandbox list
$$nemoclaw status
```
If onboarding instead reports that the sandbox "did not re-register with OpenShell after policy application," the same timeout controls that post-policy command-readiness probe. Raise the budget before retrying, then inspect the same gateway and sandbox status if re-registration still fails.
### Sandbox onboard fails with "entered Error phase before it became ready"
Onboarding ends with:
```text
Sandbox 'my-assistant' entered Error phase before it became ready (waited up to 180s).
```
On a fresh onboard the OpenShell gateway can (re)start its supervisor session and re-register the just-created sandbox. During that window `openshell sandbox list` briefly reports the sandbox in the transient `Error` phase before it flips to `Ready`, as seen on DGX Spark when supervisor restart races the sandbox bootstrap.
NemoClaw polls immediately, starts retrying after 250ms, and backs off to a 2-second cap.
It tolerates 30 consecutive `Error` observations by default so this transient recovers on its own.
Only `Error` that persists through the debounce count is terminal, unless the overall `NEMOCLAW_SANDBOX_READY_TIMEOUT` deadline expires first.
Every terminal observation outside the `Error` phase, including one with no reported phase, fails immediately.
If your host needs more observations for slower re-registration, raise the debounce. Raise `NEMOCLAW_SANDBOX_READY_TIMEOUT` too if the overall deadline is too short. To fail fast on the first `Error` poll, set the debounce to `1`:
```bash
export NEMOCLAW_SANDBOX_READY_ERROR_DEBOUNCE=1
$$nemoclaw onboard
```
If the failure persists after the debounce, the sandbox is stuck. Inspect the retained diagnostics and gateway state:
```bash
openshell sandbox list
$$nemoclaw status
```
### Compatible endpoint fails sandbox validation or later at runtime
Some OpenAI-compatible servers (such as SGLang) expose `/v1/responses` but their streaming mode is incomplete. OpenClaw requires granular streaming events like `response.output_text.delta` that these backends do not emit.
For the compatible-endpoint provider, NemoClaw now defaults to `/v1/chat/completions` and skips the Responses API probe entirely unless you opt in. If you onboarded an older release that selected `/v1/responses`, recreate the named sandbox so the wizard rebuilds the image with chat completions.
Clear any previous Responses API preference before recreating the sandbox.
This command replaces the existing sandbox and interrupts active work.
Finish active work before continuing.
```bash
unset NEMOCLAW_PREFERRED_API
$$nemoclaw onboard --fresh --name --recreate-sandbox
```
Every OpenClaw onboarding run with an OpenAI-compatible endpoint sends a validation request through `inference.local` from inside the sandbox, even when you select no messaging channel.
If that sandbox validation request fails, fix the compatible-endpoint base URL, credentials, model, or network route before continuing onboarding.
Messaging setup is not the root cause.
### Tool calls appear as assistant text
Local model servers must return structured `tool_calls` for OpenClaw to dispatch a tool. When the inference response contains only text that resembles a tool request, the gateway treats it as ordinary assistant text and no tool runs. The TUI can display a response such as:
```json
{ "arguments": { "query": "robotics" }, "name": "memory_search" }
```
This symptom is different from a network or policy block. `$$nemoclaw status`, `$$nemoclaw logs`, and `$$nemoclaw debug --quick` can all look healthy while conversation-level tool dispatch fails.
Ollama can serve local chat and some simple tool surfaces, but agent loops with several tools, long instructions, or multi-turn dispatch need a server that returns structured tool calls consistently.
| Workload | Ollama is usually sufficient | Prefer vLLM with a parser |
| -------------------------------------- | ---------------------------- | ------------------------- |
| Plain chat | Yes | Optional |
| One simple tool with short prompts | Often | Optional |
| Agent loops with several tools | Risky | Yes |
| Long system prompts or sender metadata | Risky | Yes |
| Multi-turn tool dispatch | Risky | Yes |
[Set up vLLM](../inference/local-inference/set-up-vllm) with automatic tool choice and the tool-call parser that matches the model family for persistent agent use. After the parser-aware server is ready, re-run onboarding. Select the **Local vLLM** entry for a server detected on `localhost:${NEMOCLAW_VLLM_PORT:-8000}`; for another address, select **Other OpenAI-compatible endpoint**. On generic hosts, the Local vLLM entry includes an experimental label; on DGX Spark or DGX Station, it does not. On N1x, decline or disable Express, or set `NEMOCLAW_PROVIDER=vllm`, to enter standard onboarding; this route remains unvalidated on N1x.
NemoClaw restores the complete credential-sanitized native OpenClaw configuration across rebuilds, but those edits do not change the OpenShell route, credential binding, or network policy.
Use `$$nemoclaw inference set` when the provider change requires those host-side resources to stay aligned.
Ask the agent to perform an action that requires a tool, then confirm that the TUI does not show a JSON blob as assistant text, the gateway log shows tool dispatch followed by an answer, and `$$nemoclaw status` reports the intended local vLLM or compatible provider. If JSON still appears as text, confirm that vLLM started with automatic tool choice and the parser required by the model family.
### Onboarding rejects an Anthropic-compatible tool call
Validation for an OpenClaw **Other Anthropic-compatible endpoint** selection can fail with either of these diagnostics:
```text
anthropic-streaming-missing-tool-use
anthropic-streaming-missing-tool-use-stop-reason
```
After an ordinary `/v1/messages` request succeeds, NemoClaw sends a streaming request that forces the endpoint to call the `emit_ok` validation tool. The response must include a native Anthropic `tool_use` content block named `emit_ok` and a later `message_delta` with `stop_reason: tool_use`. NemoClaw checks these observations independently: the first diagnostic means the named native block is absent, and the second means the stream does not finish the tool request with the required stop reason.
A text delta containing JSON such as `{"name":"emit_ok","arguments":{"value":"OK"}}` is still assistant text. NemoClaw rejects it instead of treating it as a tool call because OpenClaw cannot dispatch text as a native Anthropic tool request. Fix the endpoint's Anthropic tool parser or chat template so it emits native protocol events, then run onboarding again.
This check applies only to OpenClaw custom Anthropic routes. Hermes and Deep Agents Code keep their existing `/v1/chat/completions` validation and do not run the native Anthropic `emit_ok` probe. For a reasoning-only endpoint, `NEMOCLAW_REASONING=true` skips the streaming sequence and forced tool-call checks; OpenClaw still needs streaming and native tool calls at runtime, so use this only when the selected model cannot complete the onboarding probe.
### Onboarding fails with duplicate Anthropic message_start events
Validation for an OpenClaw **Other Anthropic-compatible endpoint** selection ends with an error like:
```text
Anthropic Messages API (streaming): duplicate message_start
```
For OpenClaw custom Anthropic routes, NemoClaw sends a `stream: true` request to `/v1/messages` and validates the SSE event sequence (exactly one `message_start`, at least one `content_block_delta`, and a `message_stop`). The same request forces the `emit_ok` tool and separately requires a native `tool_use` block plus `stop_reason: tool_use`; JSON-shaped assistant text does not satisfy that tool-call contract. This error means the streaming layer on the endpoint or gateway is malformed even though its non-streaming responses are valid. A working non-streaming response does not imply that streaming works. Some inference gateways proxy plain requests correctly but corrupt the SSE stream, for example by emitting `message_start` twice for one request. OpenClaw uses the streaming path, so without this check the defect would first surface inside the sandbox as a runtime failure.
Hermes and OpenAI-compatible-only agents use the endpoint's `/v1/chat/completions` surface for custom Anthropic selections instead. Current onboarding validates that surface and does not reject those agents because of a malformed native `/v1/messages` stream they will not use. An older Hermes sandbox that still uses native Anthropic Messages can report that no final response was produced; re-run onboarding to select and validate the managed Chat Completions route.
Fix the streaming layer on the endpoint or gateway, or onboard with a different Anthropic-compatible endpoint. The official Anthropic provider does not run this check and is not affected. If an OpenClaw sandbox created by an older release fails at runtime with an empty final response on an Anthropic-compatible endpoint, re-run `$$nemoclaw onboard` so the streaming check can diagnose the endpoint. If the endpoint serves a reasoning-only model, set `NEMOCLAW_REASONING=true` to skip the streaming sequence and forced tool-call checks. Streaming or native tool-call defects then surface at runtime instead of during onboarding.
### `NEMOCLAW_DISABLE_DEVICE_AUTH=1` does not disable pairing
This is expected behavior. OpenClaw 2026.9.1 retired the device-auth bypass. NemoClaw accepts `NEMOCLAW_DISABLE_DEVICE_AUTH` and its provenance input only while managed Dockerfile callers transition, but it does not emit an OpenClaw device-auth bypass key. Changing the input before or after a build does not skip device pairing.
Complete the normal pairing flow for each browser or CLI device. For the security model, refer to [Gateway Authentication Controls](../security/gateway-authentication-controls).
### OpenClaw Config Is Empty After Legacy or Manual Damage
A legacy non-atomic inference write or an interrupted manual edit can leave `/sandbox/.openclaw/openclaw.json` empty.
Current `inference set` writes use an atomic native OpenClaw batch that applies all related values or none.
If NemoClaw cannot confirm the batch, rerun the same `inference set` command instead of using this recovery procedure.
When the file is empty, OpenClaw commands may report that the config is empty instead of showing a raw JSON parse error.
NemoClaw treats `openclaw.json` as ordinary OpenClaw-owned state.
It does not maintain a hash, seal, last-good copy, recovery anchor, baseline, or config guard that can replace or reject the file during startup.
Restore a trusted snapshot that predates the failed write:
```bash
$$nemoclaw snapshot list
$$nemoclaw snapshot restore
```
If no trusted snapshot exists, back up other required state manually, recreate the sandbox through onboarding, and restore only the trusted files.
### `openclaw channels add` or `remove` is blocked inside the sandbox
This is expected.
NemoClaw freezes the messaging channel list into the sandbox image during `$$nemoclaw onboard` or `$$nemoclaw rebuild`.
NemoClaw compiles the selected channel configuration into `NEMOCLAW_MESSAGING_PLAN_B64` for that build.
The build applies the selected agent configuration to `/sandbox/.openclaw/openclaw.json` for OpenClaw or `/sandbox/.hermes/.env` for Hermes, writes reduced runtime metadata to `/usr/local/share/nemoclaw/messaging-runtime-plan.json`, and removes the full build plan from the runtime environment.
Credential bindings remain OpenShell credential placeholders, so raw messaging credentials do not enter the sandbox image or agent configuration.
NemoClaw restores the complete credential-sanitized native OpenClaw configuration across rebuilds, but an in-sandbox config change cannot create the matching OpenShell credential binding or policy.
NemoClaw's sandbox entrypoint intercepts `openclaw channels ` and points to the host-side commands below.
Run the equivalent host-side command instead:
```bash
$$nemoclaw channels list
$$nemoclaw channels add
$$nemoclaw channels remove
```
`channels add` registers credentials with the OpenShell gateway and `channels remove` clears them.
Both offer to rebuild the sandbox so the image reflects the new channel set.
In non-interactive mode (`NEMOCLAW_NON_INTERACTIVE=1`, or any run without a terminal on stdin), the commands stage the change and leave the rebuild to a follow-up `$$nemoclaw rebuild`.
Review [Choose Messaging Channels](../manage-sandboxes/messaging-channels/choose-messaging-channels) for the supported channel IDs, prerequisites, and experimental status before enabling a channel.
WeChat captures its bot token through a host-side QR scan during `$$nemoclaw onboard` or `channels add wechat`. You scan the iLink QR from WeChat on your phone and NemoClaw registers the captured token with the OpenShell gateway.
WhatsApp pairs entirely inside the sandbox. NemoClaw advertises WhatsApp for OpenClaw and Hermes sandboxes after you add the channel on the host. Run `openclaw channels login --channel whatsapp` inside OpenClaw sandboxes, or run `hermes whatsapp` inside Hermes sandboxes.
### Change OpenClaw Configuration Inside the Sandbox
OpenClaw owns `/sandbox/.openclaw/openclaw.json` after first launch.
Connect to the sandbox and use its native configuration commands:
```bash
$$nemoclaw connect
openclaw configure
openclaw config set
openclaw config unset
```
Restart and reconnect preserve the current native configuration.
Rebuild restores the credential-sanitized configuration captured immediately before rebuild.
Snapshot restore reapplies only the credential-sanitized configuration present when that snapshot was created.
The host command `$$nemoclaw config set` returns guidance to use the native OpenClaw command.
Native OpenClaw commands do not create OpenShell credential bindings or widen network policy.
Use NemoClaw credential, inference, and policy commands when a change needs matching host-side state.
Keep provider credentials in OpenShell because values stored directly in `openclaw.json` are readable inside the sandbox and can be removed by snapshot sanitization.
Restore does not change OpenShell policy, credential bindings, or the host-side inference route.
Reconcile those resources after restore when the native configuration expects different values, then restart the gateway when OpenClaw requires it.
### `openclaw doctor --fix` cannot reconcile managed Discord host resources
`openclaw doctor --fix` can update native channel configuration, but it cannot create the matching OpenShell credential binding or network policy.
If a managed Discord, Telegram, or Slack channel is still unhealthy, use the host-side channel command to reconcile those resources.
Do not treat a successful native config repair as proof that the external channel credential and network path are ready.
### Official messaging plugin verification stops the build
OpenClaw messaging image builds verify the installed provenance of official Discord, Slack, WhatsApp, Microsoft Teams and Google Chat plugins.
If the inspection fails, times out or cannot prove a trusted official installation, the build stops before the rebuilt-runtime channel checks.
A missing verified package in the offline build cache also stops installation.
These are build failures; the post-rebuild startup warnings apply only after the image builds successfully.
Report the failure diagnostic, plugin name and NemoClaw version to a maintainer.
NemoClaw manages the package pins and build cache; do not edit them or bypass plugin verification to continue.
After a maintainer resolves the reported build failure, rerun the original onboarding or rebuild command.
### Discord bot logs in, but the channel still does not work
Separate the problem into two parts:
1. Channel selection and provider wiring
Check that onboarding selected Discord and that the sandbox was created with the Discord messaging provider attached. If Discord was skipped during onboarding, rerun onboarding and select Discord again.
1. Native Discord gateway path
Successful login alone does not prove that Discord works end to end. Discord also needs a working gateway connection to `gateway.discord.gg`. If logs show errors such as `getaddrinfo EAI_AGAIN gateway.discord.gg`, repeated reconnect loops, or a `400` response while probing the gateway path, the problem is usually in the native gateway/proxy path rather than in the native channel configuration.
Common signs of a native gateway-path failure:
- REST calls to `discord.com` succeed, but the Discord channel never becomes healthy
- `gateway.discord.gg` fails with DNS resolution errors
- the WebSocket path returns `400` instead of opening a tunnel
- native command deployment fails even though the bot token itself is valid
In that case:
- keep the Discord policy preset applied
- verify the sandbox was created with the Discord provider attached
- inspect gateway logs and blocked requests with `openshell term`
- treat the failure as a native Discord gateway problem, not as a bridge startup problem
### Discord preset validation behind a proxy
The built-in Discord policy preset allows the OpenClaw Node.js runtime and does not allow `curl`. As a result, `curl -s https://discord.com` failing, hanging, or printing no output is not proof that the Discord preset is broken.
Behind the OpenShell proxy, direct DNS-only checks can also be the wrong signal. For example, `dns.resolve("gateway.discord.gg")` can fail even when HTTPS requests routed through the proxy are healthy.
Use the OpenClaw Node.js runtime for the manual REST probe:
```bash
node - <<'NODE'
const https = require("node:https");
https
.get("https://discord.com/api/v10/gateway", (res) => {
console.log(`${res.statusCode} ${res.statusMessage || ""}`.trim());
res.resume();
})
.on("error", (err) => {
console.error(err.message);
process.exitCode = 1;
});
NODE
```
To check Discord CDN egress, use the same Node HTTPS path:
```bash
node - <<'NODE'
const https = require("node:https");
https
.get("https://cdn.discordapp.com/", (res) => {
console.log(`${res.statusCode} ${res.statusMessage || ""}`.trim());
res.resume();
})
.on("error", (err) => {
console.error(err.message);
process.exitCode = 1;
});
NODE
```
Any HTTP status from these probes means the Node.js process reached the endpoint; the exact status can vary by unauthenticated path. A transport error or OpenShell policy denial means the probe failed. If the REST probe works but the Discord channel is still unhealthy, investigate the native gateway path instead of widening the preset. Check the gateway logs and blocked-request output with `openshell term`, and look for `gateway.discord.gg` connection or WebSocket upgrade failures.
### Messaging bridge appears running but no messages arrive
Telegram `getUpdates` allows only one active poller per bot token. Reusing Discord or Slack credentials can create competing gateway or Socket Mode sessions and unreliable message delivery. `$$nemoclaw status` can still report a bridge as running because the gateway process itself is alive.
For Telegram group chats, first check BotFather privacy mode. New Telegram bots default to privacy mode enabled, which prevents group messages from reaching `getUpdates` even when the user mentions the bot. In @BotFather, run `/setprivacy`, choose the bot, and choose **Disable**. Then remove the bot from the affected group and add it back; Telegram applies the privacy-mode change to group delivery only after the bot rejoins.
For Telegram direct messages, make sure the rebuilt sandbox has a DM allowlist. Set `TELEGRAM_ALLOWED_IDS` before rebuild; `TELEGRAM_AUTHORIZED_CHAT_IDS` and `TELEGRAM_CHAT_ID` are accepted as compatibility aliases. Keep the aliases until QA automation and public repro templates have stopped exporting them for at least one full release. Bot API `sendMessage` sends from the bot to a chat, so it only proves outbound Telegram API access. To prove inbound agent routing, send a message from the Telegram client as an allowed user and then watch the gateway log for the agent turn and outbound reply. For a reproducible outbound runtime check, run `NEMOCLAW_RUN_LIVE_E2E=1 npx vitest run --project e2e-live test/e2e/live/messaging-providers.test.ts --silent=false --reporter=default` with `NVIDIA_INFERENCE_API_KEY` set. The check imports the installed OpenClaw Telegram `runtime-api.js`, calls `sendMessageTelegram` through an OpenShell-rewritten credential against a host-side fake Telegram API, and verifies the captured chat, text, token rewrite, and absence of unresolved placeholders. When `TELEGRAM_BOT_TOKEN_REAL` and `TELEGRAM_CHAT_ID_E2E` are also set, the same lane performs an additional real outbound send; it does not prompt for or claim an interactive inbound reply.
To diagnose, open a shell in the sandbox and inspect the gateway log:
```bash
$$nemoclaw connect
tail -f /tmp/gateway.log
```
A repeating line like the following confirms the conflict:
```text
[telegram] getUpdates conflict: 409: Conflict: terminated by other getUpdates request; retrying in 30s.
```
To fix, run `$$nemoclaw destroy` on whichever sandbox should stop polling, or rerun onboarding on it with the channel disabled. NemoClaw checks only the sandboxes in the selected OpenShell gateway's sandbox registry. It cannot detect or prevent Slack credential reuse across independent OpenShell gateways. Run only one active Slack sandbox on each OpenShell gateway. Use distinct Slack bot and app tokens for Slack sandboxes on different OpenShell gateways. Within the selected registry, onboarding, rebuild, and `channels add` abort on a conflict or an incomplete required check, including unavailable credential hashes. Only `channels add --force` can accept the duplicate-consumer or shared-resource risk. Sandboxes created before these checks were added, or managed by independent gateways, may still have a conflict without a NemoClaw warning.
### Landlock filesystem restrictions silently degraded
After sandbox creation, NemoClaw checks whether the host kernel supports Landlock (Linux 5.13+). If the kernel is too old or you are running on macOS (where the Docker VM kernel may lack Landlock), a warning prints:
```text
⚠ Landlock: Docker VM kernel does not support Landlock (requires ≥5.13).
Sandbox filesystem restrictions will silently degrade (best_effort mode).
```
This warning is informational and does not block sandbox creation. The sandbox runs without kernel-level filesystem restrictions, relying on container mount configuration instead. For full filesystem enforcement, run on a Linux kernel 5.13 or later (Ubuntu 22.04 LTS and later include Landlock support).
### Landlock filesystem policy blocks sandbox startup
Deep Agents uses strict Landlock compatibility. If the host kernel, Docker VM, or sandbox filesystem mount cannot enforce the managed read-only policy, OpenShell refuses to start the sandbox instead of silently degrading.
Run Deep Agents on a Linux kernel 5.13 or later with a container runtime that exposes Landlock to the sandbox. After moving to a compatible host or runtime, rerun onboarding or rebuild the sandbox:
```bash
nemo-deepagents rebuild
```
### Sandbox lost after gateway restart
Sandboxes created with OpenShell versions older than 0.0.24 can become unreachable after a gateway restart because SSH secrets were not persisted. Running `$$nemoclaw onboard` automatically upgrades OpenShell to 0.0.24 or later during the preflight check. After the upgrade, recreate the sandbox with `$$nemoclaw onboard`.
### DNS-backed HTTPS endpoint is not supported
NemoClaw rejects an explicit custom endpoint when it resolves a public HTTPS hostname but cannot pin the same peer address across the downstream OpenShell runtime boundary while preserving TLS SNI and host validation. This can appear during a direct blueprint run, custom-endpoint onboarding, or a Hermes host-side `config set` write.
It does not appear during a runtime `$$nemoclaw inference set` switch on an
already-onboarded sandbox; that command routes a DNS-backed HTTPS endpoint
through a local HTTPS Pin Runtime adapter instead of rejecting it. Refer to
[Commands](commands) for details.
Use an HTTPS IP-literal endpoint whose certificate is valid for that address. If your deployment permits non-TLS provider traffic, you can instead use a public HTTP endpoint that NemoClaw can rewrite to a DNS-pinned address. Do not bypass the check with a private or internal address or by editing the persisted sandbox config directly. For the full endpoint rules, refer to [Meet Custom Endpoint Security Requirements](../inference/custom-endpoints/custom-endpoint-security).
### Agent cannot reach external hosts through a proxy
NemoClaw uses a default proxy address of `10.200.0.1:3128` (the OpenShell-injected gateway). If your environment uses a different proxy, set `NEMOCLAW_PROXY_HOST` and `NEMOCLAW_PROXY_PORT` before onboarding:
```bash
export NEMOCLAW_PROXY_HOST=proxy.example.com
export NEMOCLAW_PROXY_PORT=8080
$$nemoclaw onboard
```
These are build-time settings baked into the sandbox image. Changing them after onboarding requires re-running `$$nemoclaw onboard` to rebuild the image.
When `HTTP_PROXY` or `HTTPS_PROXY` is set on the host, NemoClaw adds `localhost`, `127.0.0.1`, `::1`, `0.0.0.0`, the container-host aliases `host.docker.internal` and `host.containers.internal`, and the managed inference hostname `inference.local` to `NO_PROXY` for host-side subprocesses and for the env forwarded into `openshell sandbox create`. This keeps local Ollama health checks, model pulls, and managed inference traffic from being chained through a corporate or desktop proxy at the sandbox-create boundary, while preserving the proxy for external hosts. For the local provider validation probe, NemoClaw removes `HTTP_PROXY`, `HTTPS_PROXY`, and `ALL_PROXY` from the probe process and sets `NO_PROXY=*` instead. A host proxy therefore cannot answer for the local endpoint, including the `host.docker.internal` alias used for Windows-host Ollama. Inside the running sandbox, processes continue to use the OpenShell L7 proxy for `inference.local` so OpenShell's internal routing, DNS, and audit boundaries stay intact.
### Agent cannot reach a host-side HTTP service
When a sandbox needs to call an HTTP service running on the host, use the normal OpenShell network policy path. Expose the service on a host IP address that the OpenShell gateway can reach, create a custom NemoClaw policy preset for that IP and port, and apply it with `$$nemoclaw policy add --from-file`. The sandbox request then flows through the OpenShell proxy while NemoClaw preserves the existing live policy entries.
Do not rely on `host.docker.internal` or `host.openshell.internal` as a general-purpose host-service path. Those names may appear in the sandbox's `/etc/hosts`, but in OpenShell's sandbox network they are not guaranteed to point at a reachable host gateway. Bypassing the proxy with `--noproxy '*'` also bypasses network policy enforcement and audit.
First, make sure the host-side service listens on a non-loopback address. For example, a health endpoint on port `50001` should be reachable from the host IP, not only from `127.0.0.1`:
```bash
curl -s http://10.0.0.5:50001/health
```
Expected output:
```json
{ "status": "ok" }
```
Then create a custom NemoClaw preset for the host-side service. Replace `10.0.0.5`, `50001`, paths, methods, and binaries with the service you want the sandbox to reach:
```yaml
preset:
name: host-memory-api
description: "Host memory API"
network_policies:
host_memory_api:
name: host_memory_api
endpoints:
- host: 10.0.0.5
port: 50001
protocol: rest
enforcement: enforce
rules:
- allow: { method: GET, path: "/health" }
binaries:
- { path: /usr/bin/curl }
```
Preview the private-host trust pins and the exact binary, method, path, port, and access limits:
```bash
$$nemoclaw my-assistant policy add --from-file ./host-memory-api.yaml --trusted-private-host 10.0.0.5 --dry-run
```
After reviewing the generated `allowed_ips`, apply the same bounded request:
```bash
$$nemoclaw my-assistant policy add --from-file ./host-memory-api.yaml --trusted-private-host 10.0.0.5
```
After you apply the policy, retry the request from inside the sandbox without disabling the proxy:
```bash
curl -s http://10.0.0.5:50001/health
```
Expected output:
```json
{ "status": "ok" }
```
If the request is still denied, check the blocked request in `openshell term`. The policy `binaries` list must include the executable path that actually made the request. If the response changes from `policy_denied` to `upstream_unreachable`, the policy matched, but the OpenShell gateway could not reach the host IP and port.
### Agent cannot reach an external host
OpenShell blocks outbound connections to hosts not listed in the network policy. Open the TUI to see blocked requests and approve them:
```bash
openshell term
```
To permanently allow an endpoint, add it to the network policy.
Refer to [Customize the Network
Policy](../network-policy/customize-network-policy) for details.
For Deep Agents, follow [Customize the Network
Policy](../network-policy/customize-network-policy) to choose between built-in
presets, reviewed custom presets, baseline edits, and live-policy replacement.
### Dashboard not reachable after setting a custom port
If you ran `$$nemoclaw onboard` with a custom dashboard port and onboarding completed but the dashboard URL is unreachable, the sandbox was most likely created with an older NemoClaw version that did not pass the dashboard port into the sandbox at startup. The browser may show connection refused or fail to load the page. The gateway inside the sandbox continued listening on the default port 18789 while the SSH tunnel forwarded the custom port, leaving nothing at the other end of the tunnel.
Re-run onboarding on the current NemoClaw release with the desired port. Current versions derive the dashboard port from `CHAT_UI_URL` automatically and inject it into the sandbox:
```bash
CHAT_UI_URL=http://127.0.0.1:19000 $$nemoclaw onboard
```
If you need to run multiple sandboxes at different ports at the same time, refer to [Running multiple sandboxes simultaneously](#running-multiple-sandboxes-simultaneously).
### Control UI config endpoint returns 404 or non-JSON
The Control UI loads its runtime configuration from a gateway endpoint, not from a static `controlui.bootstrap.config.json` file. No `controlui.bootstrap.config.json` path is served, so requesting it returns `HTTP 404 Not Found` with a short plain-text body, and piping that response to `jq` fails with a parse error such as `Invalid numeric literal`.
The supported Control UI config endpoint is `/__openclaw/control-ui-config.json`, served by the OpenClaw gateway on the forwarded dashboard port. It is gated by the gateway auth token:
- An unauthenticated request returns `HTTP 401 Unauthorized` with a JSON body (`{"error":{"message":"Unauthorized","type":"unauthorized"}}`), which is already valid JSON.
- An authenticated request returns `HTTP 200 OK` with the Control UI config as JSON.
Resolve the forwarded dashboard port, then authenticate with the gateway token from `$$nemoclaw gateway-token`:
```bash
openshell forward list # note the dashboard PORT for the sandbox
export DASH_PORT=
TOKEN=$($$nemoclaw gateway-token --quiet)
curl -fsS -H "Authorization: Bearer $TOKEN" \
"http://127.0.0.1:${DASH_PORT}/__openclaw/control-ui-config.json" | jq empty \
&& echo "Control UI config is valid JSON"
```
The token is sensitive; treat it like a password and do not log, share, or commit it. For browser access, use the tokenized URL from `$$nemoclaw dashboard-url` instead of calling the config endpoint directly.
Hermes manages its own dashboard sessions and does not expose an OpenClaw gateway auth token or a `/__openclaw/control-ui-config.json` endpoint. Use `nemohermes status` to see the dashboard and API endpoints for a Hermes sandbox.
### Ollama auth proxy did not start
NemoClaw keeps Ollama bound to `127.0.0.1:11434` and starts a token-gated reverse proxy on `0.0.0.0:11435` so the sandbox can reach Ollama without exposing it to the local network. If the proxy fails to start, onboarding exits before configuring inference.
Check whether the proxy port is occupied by another process:
```bash
sudo lsof -i :11435
```
Stop the conflicting process and re-run `$$nemoclaw onboard`. The wizard cleans up stale proxy processes from previous runs automatically, so most failures resolve by retrying.
If the proxy refuses to start because the backend also listens on a non-loopback interface, use the remediation that matches the reported backend:
- For Ollama, bind the daemon to the reported loopback port with `OLLAMA_HOST=127.0.0.1:`, then restart Ollama and rerun onboarding.
- For an unauthenticated OpenAI-compatible endpoint, bind that endpoint server to a loopback address only on the reported port, then rerun onboarding. Do not apply the Ollama setting to vLLM, llama-server, or another compatible endpoint.
- If recovery cannot identify the backend type, bind the reported service and port to a loopback address only, then rerun onboarding. The neutral diagnostic intentionally does not name Ollama.
In every case, the refusal prevents a backend listener from bypassing the protected route's token check. For an IPv6 endpoint, keep it on its IPv6 loopback address instead of changing it to `127.0.0.1`.
The proxy token is persisted to `~/.nemoclaw/ollama-proxy-token` with `0600` permissions. If the file is missing or unreadable after a host reboot, re-running `$$nemoclaw onboard` regenerates it.
### Ollama auth proxy is unreachable from the sandbox
On native Linux Docker-driver hosts, a host firewall can allow the host proxy check but block sandbox traffic to the Ollama auth proxy. When that happens, onboarding exits before it saves the inference route and prints output like:
```text
✗ Sandbox containers cannot reach the Ollama auth proxy at host.openshell.internal:11435.
A host firewall may be blocking traffic from the OpenShell Docker bridge.
```
Apply the `ufw` command printed by onboarding, then rerun onboarding. If the message does not include a subnet, derive it from the OpenShell Docker network:
```bash
SUBNET=$(docker network inspect openshell-docker --format '{{(index .IPAM.Config 0).Subnet}}')
sudo ufw allow from "$SUBNET" to any port 11435 proto tcp
$$nemoclaw onboard
```
Only an accepted Windows-host Ollama route uses direct Docker Desktop access to port `11434`. WSL-local Ollama uses the auth proxy, including on Docker Desktop WSL. Other hosts without the OpenShell Docker network use different routing models; for proxy-fronted routes, NemoClaw treats an unavailable sandbox-side probe as non-blocking and relies on the regular proxy health check.
### `host.docker.internal` does not reliably reach the host from the sandbox
Configuring an inference provider with a base URL like `http://host.docker.internal:11434/v1` does not reliably reach a host Ollama service from inside the OpenShell sandbox. OpenShell runs sandboxes inside a k3s network, where `host.docker.internal` is not a reliable host-service route. Depending on the platform, it may fail DNS resolution or resolve to an internal gateway/bridge address where the host's port `11434` is not forwarded. The sandbox then sees a DNS failure or `connection refused`:
```bash
getent hosts host.docker.internal
```
Expected output:
```text
172.17.0.1 host.docker.internal host.openshell.internal
```
```bash
no_proxy=host.docker.internal curl -v http://host.docker.internal:11434/api/tags
```
Expected output:
```text
* connect to 172.17.0.1 port 11434 failed: Connection refused
```
For local Ollama on non-WSL hosts, use the auth-proxy URL that NemoClaw's "Local Ollama" onboard option configures automatically:
```text
http://host.openshell.internal:11435/v1
```
`host.openshell.internal` resolves to the same gateway IP, and the [token-gated Ollama auth proxy](#ollama-auth-proxy-did-not-start) binds port `11435` there and forwards requests to `127.0.0.1:11434` on the host. Only an accepted Windows-host Ollama route uses direct Docker Desktop access to port `11434`; WSL-local Ollama uses the authenticated proxy even with Docker Desktop. Native Docker Engine inside WSL is unqualified; enable Docker Desktop WSL integration instead of configuring the proxy URL manually. If you need a different host service exposed to the sandbox, route it through the OpenShell gateway rather than relying on `host.docker.internal`. Refer to issue [#3136](https://github.com/NVIDIA/NemoClaw/issues/3136).
### Local inference health check resolves to IPv6
Local inference health checks now use `127.0.0.1` instead of `localhost`. On systems where `localhost` resolves to `::1` first, older NemoClaw releases could probe the wrong address and report the local backend as unreachable even when it was running. If you see this on a current NemoClaw release, verify that the local backend binds an IPv4 address and not only `::1`.
### Blueprint run failed
View the error output for the failed blueprint run:
```bash
$$nemoclaw logs
```
Use `--follow` to stream logs in real time while debugging.
## DGX Spark
For an end-to-end walkthrough with local inference on DGX Spark, refer to the [NVIDIA Spark playbook](https://build.nvidia.com/spark/nemoclaw).
### Host freezes or logs `NVRM NV_ERR_NO_MEMORY` under local vLLM load
Treat a full host freeze separately from an agent tool-call hang. If the Spark stops responding to SSH and ping, and the journal contains `NVRM NV_ERR_NO_MEMORY` or no software-side crash record, first isolate the local inference server before changing MCP or network policy configuration. For onboarding-time context, refer to [Use an Existing Server](../inference/local-inference/set-up-vllm#use-an-existing-server).
Check whether vLLM is a bring-your-own server or the NemoClaw managed Spark profile:
Inspect the running containers, current memory, and kernel evidence:
```bash
docker ps --format 'table {{.Names}}\t{{.Image}}\t{{.Status}}\t{{.Ports}}'
free -h
journalctl -k --since "24 hours ago" --no-pager | grep -Ei 'NVRM|OOM|out of memory|lockup|watchdog'
```
For a NemoClaw-managed Spark profile, derive the host port from the managed container's fixed `8000/tcp` mapping. Bearer-protected profiles publish two bindings with the same host port, while bearerless profiles publish one all-interface binding.
```bash
VLLM_HOST_PORT="$(
docker port nemoclaw-vllm 8000/tcp |
sed -n 's/^.*:\([0-9][0-9]*\)$/\1/p' |
sort -u
)"
case "$VLLM_HOST_PORT" in
""|*[!0-9]*)
printf '%s\n' 'Could not identify one managed vLLM host port.' >&2
exit 1
;;
esac
if [ "$VLLM_HOST_PORT" -lt 1024 ] || [ "$VLLM_HOST_PORT" -gt 65535 ]; then
printf 'Invalid managed vLLM host port: %s\n' "$VLLM_HOST_PORT" >&2
exit 1
fi
curl -fsS "http://127.0.0.1:${VLLM_HOST_PORT}/health"
```
For an existing vLLM server, set `NEMOCLAW_VLLM_PORT` to its host port in the current shell:
```bash
VLLM_HOST_PORT="${NEMOCLAW_VLLM_PORT:?Set NEMOCLAW_VLLM_PORT to the existing vLLM host port.}"
curl -fsS "http://127.0.0.1:${VLLM_HOST_PORT}/health"
```
Both checks use `/health` because a managed profile can require bearer authentication for `/v1/models`.
For an existing vLLM server, inspect its launch arguments:
```bash
docker inspect --format '{{json .Config.Cmd}}'
```
Large checkpoints without explicit quantization, very long `--max-model-len` values, high `--gpu-memory-utilization`, and multiple concurrent sequences all consume the Spark's shared CPU/GPU memory pool. Before reintroducing agent tools, restart vLLM with a smaller envelope, for example:
```bash
vllm serve \
--max-model-len 32768 \
--gpu-memory-utilization 0.75 \
--max-num-seqs 1 \
--max-num-batched-tokens 4096
```
If the host still logs `NVRM NV_ERR_NO_MEMORY` while loading the model, switch to a smaller or quantized checkpoint. For managed setup, prefer `NEMOCLAW_PROVIDER=install-vllm`, which selects the Spark profile and its registered model-specific serve arguments. After standalone vLLM is stable, re-run onboarding and add MCP servers back one group at a time.
### CoreDNS CrashLoop after onboarding
If CoreDNS in the embedded k3s cluster crashes shortly after setup, it is usually because it resolves against `127.0.0.11`, which does not route inside the gateway container. Run `fix-coredns.sh` to point CoreDNS at the container gateway IP instead, then recreate the sandbox.
### `k3s` cannot find a freshly built image
After building a new sandbox image, `k3s` inside the gateway container sometimes fails to pull it even though the image exists on the host. Remove the gateway registration, then resume onboarding. If a privileged host gateway remains, do not use a host-wide process match. Verify its live owner, exact gateway name and port, command line, PID file, runtime marker, and loaded sandbox namespace immediately before you stop it.
```bash
openshell gateway remove nemoclaw
$$nemoclaw onboard --resume
```
### GPU passthrough on Spark
GPU passthrough is not CI-tested on DGX Spark. It is expected to work when you pass `--gpu` and the NVIDIA Container Toolkit is configured. Verify the toolkit is configured by running `docker run --rm --runtime=nvidia --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi` from the host. If `nvidia-smi` works on the host but onboarding says GPU passthrough was not enabled, install or repair the NVIDIA Container Toolkit, then run `sudo nvidia-ctk runtime configure --runtime=docker && sudo systemctl restart docker`. If a reusable gateway was previously started without GPU passthrough, NemoClaw replaces it automatically only when no other registered sandboxes depend on it, or when `--recreate-sandbox` is recreating the only registered sandbox with the same name. When shared gateway cleanup would be unsafe, follow the targeted destroy or gateway-removal commands printed by onboarding.
### `unresolvable CDI devices nvidia.com/gpu=all` during gateway start
Recent NVIDIA Container Toolkit installs configure the Docker daemon for Container Device Interface (CDI) device injection, which a GPU-enabled gateway start then auto-selects. If no `nvidia.com/gpu` CDI spec has been generated on the host yet, gateway start fails with `Docker responded with status code 500: CDI device injection failed: unresolvable CDI devices nvidia.com/gpu=all`. Outside Station Express, the standard NemoClaw installer detects this gap before onboarding, first tries to enable the NVIDIA CDI refresh systemd units, and can fall back to generating the spec directly with `nvidia-ctk`. Station Express never falls back to direct CDI generation. The generic Ubuntu, Colossus BaseOS, and exact AI Developer Tools paths require the packaged refresh lifecycle to work; if it fails or omits `nvidia.com/gpu=all`, inspect `nvidia-cdi-refresh.service`, repair it, and rerun the printed exact-commit install command. Other factory-runtime profiles stop when the CDI device is missing without enabling or restarting the refresh units. If you run `$$nemoclaw onboard` directly, preflight prints the manual remediation instead. The native Linux fix is the same on Docker hosts whose `docker info` advertises a non-empty `CDISpecDirs`. On WSL with Docker Desktop, Docker may advertise CDI directories even though `--device nvidia.com/gpu=all` is not usable from the WSL distro. For that runtime, NemoClaw skips Linux CDI repair and uses Docker's `--gpus` compatibility path for sandbox GPU access. This compatibility path can be retired once Docker Desktop exposes usable `nvidia.com/gpu` CDI specs inside WSL, or once OpenShell no longer requires host-visible CDI specs for Docker Desktop WSL GPU passthrough.
Enable the refresh units, verify they list `nvidia.com/gpu` entries, then rerun onboarding:
```bash
sudo systemctl enable --now nvidia-cdi-refresh.path nvidia-cdi-refresh.service
nvidia-ctk cdi list
$$nemoclaw onboard
```
For other native Linux installations, if the refresh units are unavailable or do not generate CDI devices, generate the spec directly:
```bash
sudo mkdir -p /etc/cdi
sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml
nvidia-ctk cdi list
```
On WSL with Docker Desktop, confirm Docker Desktop WSL integration is enabled for your distro and run the [canonical Docker proof and capacity query](#gpu-setup-fails-with-a-placeholder-gpu-name) from WSL.
If GPU passthrough is not required on this host, rerun onboarding with `--no-gpu` instead.
### GPU routing or compatibility patch failed
The route depends on the host environment and the operator control. Identify the matching path before applying the recovery guidance.
| Symptom | Route or stage | Recovery |
| ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Native `--gpu` is rejected, host runtime evidence identifies GPU injection failure, or an explicit driver proof fails and host configuration confirms no GPU attachment | Ordinary Linux native attempt | The default native-only route stops. Retry with `NEMOCLAW_DOCKER_GPU_PATCH=fallback` only if you explicitly accept one bounded compatibility retry, or use `=1` to select compatibility before creation. |
| `Cleanup could not be proven safe` | Native-to-compatibility handoff | Run the printed sandbox deletion command, verify both the gateway row and OpenShell-managed Docker containers labeled for that sandbox are absent, then rerun onboarding. |
| The patched container exits or the compatibility attempt fails | Compatibility recreation | Inspect the saved diagnostics and the rollback outcome, then repair the NVIDIA Container Toolkit/CDI configuration. Keep the sandbox when the pre-patch container was restored. Use only a container-specific cleanup command printed after rollback. If no command was printed, inspect the sandbox and its labeled containers before removing anything. Then rerun onboarding. |
| A recreated container inherits only a loopback DNS stub and no usable upstream | Compatibility DNS fallback | Repair the host's `systemd-resolved` upstream configuration, then rerun onboarding. |
For bridge-networked compatibility recreation without an explicit container DNS setting, NemoClaw selects a usable IPv4 upstream from `systemd-resolved` and probes that exact `--dns` path before it stops the original container. If the probe confirms that the resolver is unreachable, recreation stops and leaves the original container in place. An IPv6-only upstream list does not become a compatibility override; NemoClaw preserves Docker's default resolver path instead. Containers with explicit DNS settings or host networking keep their existing DNS path and do not use the fallback probe.
#### Ordinary native Linux bounded fallback
Ordinary Linux GPU onboarding uses native OpenShell GPU injection and stops on failure by default. Unset, `auto`, and `0` all preserve this native-only confinement boundary. `NEMOCLAW_DOCKER_GPU_PATCH=fallback` is the explicit operator authorization for one bounded retry. With that control set, if sandbox creation rejects the native GPU flag before progress, the exact OpenShell-managed container labeled for that sandbox records a host runtime GPU-injection error, or an explicit `nvidia-smi` driver proof fails while that container's immutable host configuration confirms that no GPU was attached, NemoClaw captures redacted diagnostics before one compatibility retry.
For a native flag rejection before creation progress, it proves that the sandbox and its labeled containers are absent without deleting by name.
For the other eligible failures, it requires verified cleanup before retrying.
When the managed runtime requires cleanup tied to its recorded resource identities, NemoClaw stops before the compatibility retry.
Free-form build/list text and sandbox-reported CUDA output never independently authorize the broader retry. Without corroborating host evidence, onboarding fails closed even when `fallback` is set and directs the operator to clean up and explicitly select compatibility with `NEMOCLAW_DOCKER_GPU_PATCH=1` if desired. Before the authorized retry, NemoClaw warns that the legacy GPU compatibility envelope recreates the OpenShell-managed Docker container and may relax container confinement compared with native injection. Specifically, compatibility recreation adds `SYS_PTRACE`, adds `apparmor=unconfined` when the original container has no AppArmor option, and uses a compatibility policy that makes `/proc` writable for the NVIDIA runtime's process-name initialization. These broader settings are why onboarding warns before the swap and retains a native-only opt-out. NemoClaw verifies cleanup with two stable checks (that sandbox is absent from the gateway list and no OpenShell-managed Docker containers labeled for that sandbox remain) before retrying through the compatibility path. Cleanup is polled at most five times, one second apart, and both conditions must pass twice consecutively; otherwise onboarding stops before the retry. These fail-closed safety limits are the internal constants `STABLE_ABSENCE_CHECKS` (2), `MAX_CLEANUP_ATTEMPTS` (5), and `CLEANUP_POLL_INTERVAL_MS` (1,000 ms); they are not configurable through environment variables. The first observation is immediate, so the default bound performs at most four one-second sleeps plus the five gateway/container queries. The bounds are intentionally fixed. Allowing environment input to weaken or extend the cleanup proof would make a security gate deployment-dependent. On a host that cannot prove absence within the bound, onboarding fails closed; select compatibility from the outset with `NEMOCLAW_DOCKER_GPU_PATCH=1` instead of weakening the handoff proof. If deletion or container cleanup cannot be proven safe, onboarding stops before the retry and prints manual cleanup guidance. Image build, upload, TLS, provider, policy, dashboard, and inference failures stay on their existing error paths and do not trigger the GPU compatibility fallback. Set `NEMOCLAW_DOCKER_GPU_PATCH=1` to use only the compatibility path for diagnostics or older host compatibility. Other legacy nonzero values keep that behavior through the `v0.0.x` release line and will be removed in `v0.1.0`.
#### Docker Desktop WSL compatibility route
Automatic GPU onboarding uses the compatibility path directly; it does not make a native attempt first. The path creates the sandbox and then recreates the OpenShell-managed Docker container with NVIDIA GPU flags. `NEMOCLAW_DOCKER_GPU_PATCH=0` is ignored because this runtime requires the compatibility patch for GPU passthrough, and onboarding logs a warning when it is set. To skip GPU passthrough entirely, rerun with `--no-gpu` or set `NEMOCLAW_SANDBOX_GPU=0`.
#### Jetson and Tegra compatibility default
Automatic GPU onboarding uses the compatibility path directly; it does not make a native attempt first. The path recreates the OpenShell-managed Docker container with NVIDIA GPU flags and propagates eligible host group IDs for the supported Jetson GPU device nodes. Use `NEMOCLAW_DOCKER_GPU_PATCH=0` only for troubleshooting because it bypasses that group propagation and CUDA may not initialize.
#### Common compatibility-path recovery
After compatibility recreation starts, onboarding keeps the pre-patch container as a rollback backup until the replacement passes the Ready, direct GPU, and applicable local-inference checks. If a later check fails, onboarding prints failure diagnostics and attempts to restore the pre-patch container before it exits. When rollback succeeds, the pre-patch sandbox remains available. If the failed replacement may remain and NemoClaw retains its validated exact container ID, it prints only an exact-container `docker rm -f` command. When a replacement is not stably running, the failure diagnostic includes its exact runtime ID, inspected state, and a bounded redacted log tail when available. If replacement cleanup cannot be confirmed without a validated exact ID, onboarding reports cleanup as unknown and prints no deletion command. When rollback fails, onboarding reports that the sandbox and container state is uncertain and prints no deletion command. A diagnostic bundle captured before rollback records cleanup as pending and contains no deletion command. Inspect the diagnostics, the sandbox, and its labeled Docker containers before removing anything.
Starting with NemoClaw v0.0.43, the standard installer handles the `/proc//task//comm` permission case during this patch path.
If an older release fails direct GPU proof with that path and `Permission denied`, upgrade NemoClaw and rerun onboarding.
When the failed sandbox remains, recovery stays blocked because OpenShell deletion targets its mutable name. Inspection is diagnostic only and does not authorize deletion. Run `$$nemoclaw destroy` to check for authoritative absence. Fix the NVIDIA Container Toolkit or CDI configuration reported in the diagnostics. Rerun onboarding only after retained recovery completes. If you do not need GPU access inside the sandbox, use `--no-sandbox-gpu` when onboarding can continue.
If sandbox creation fails with `CDI device injection failed: unresolvable CDI devices nvidia.com/gpu=all`, the OpenShell gateway tried `docker create --device nvidia.com/gpu=all` and Docker could not resolve the CDI spec. This injection happens inside the gateway, so `NEMOCLAW_DOCKER_GPU_PATCH=0` does not bypass it. Rerun with `--no-gpu`, or set `NEMOCLAW_SANDBOX_GPU=0` and resume onboarding.
If onboarding reports `OpenShell supervisor did not reconnect to the GPU-enabled container.` even though the diagnostic bundle shows the patched container is running and healthy, the supervisor-reconnect wait is treating a transient Error phase (reported while the OpenShell host re-registers the new container) as fatal. The reconnect wait debounces consecutive Error-phase polls before fast-failing, defaulting to 60 consecutive polls of about 120 seconds in total. Increase the debounce window with `NEMOCLAW_DOCKER_GPU_SUPERVISOR_RECONNECT_ERROR_DEBOUNCE` if your host needs more time to re-register the patched container, for example slow WSL2 + Docker Desktop setups. Set it to an integer above the default of 60, such as `120` (about 240 seconds), and rerun onboarding; the value is clamped to a minimum of `1`. If reconnect still fails after the GPU patch, NemoClaw attempts to restore the pre-patch CPU container before exiting. When rollback succeeds, the output says the pre-patch sandbox was restored. When rollback fails, the error says rollback failed and the pre-patch container was not restored, so inspect Docker state before retrying.
### `pip install` fails with a system-packages error
Recent Ubuntu releases (including DGX Spark's Ubuntu 24.04) mark the system Python install as externally managed, so `pip install` without a virtual environment fails. Use a venv instead. Avoid `--break-system-packages` unless you understand the risk, since it can break host tooling.
```bash
python3 -m venv ~/.venvs/nemoclaw
source ~/.venvs/nemoclaw/bin/activate
pip install ...
```
### Port 3000 conflict with AI Workbench
NVIDIA AI Workbench's Traefik proxy binds ports 3000 and 10000. If you run other services on Spark that expect port 3000, bind them to a different port.
## Windows Subsystem for Linux
For environment setup steps, refer to [Windows Prerequisites](../get-started/additional-setup/windows-preparation).
### `wsl --install --no-distribution` returns Forbidden (403)
Check your network connectivity. If you are behind a VPN, try reconnecting or switching to a different network. If your network or Windows image blocks the online WSL installer, install WSL manually with Microsoft's [offline install guidance](https://learn.microsoft.com/en-us/windows/wsl/install#offline-install). Download the latest official WSL `.msi` package from the [Microsoft WSL releases page](https://github.com/microsoft/WSL/releases/latest), choose the matching `.x64.msi` or `.arm64.msi`, install it, reboot if Windows requests it, then rerun `wsl --status`.
### Bootstrap says "Windows Subsystem for Linux is not fully installed"
The bootstrap script checks `wsl --status` before it installs or opens Ubuntu. If Windows reports that the WSL runtime is not installed, the script attempts `wsl --install --no-distribution` automatically. If the repair command succeeds and WSL reports that the changes require a reboot, reboot and let the bootstrap resume after sign-in. If the repair command succeeds but WSL still cannot be verified and the output does not request a reboot, follow the printed repair guidance instead of rebooting by default. If the repair command fails, follow the printed repair steps. If the repair command returns `Forbidden (403)` or remains blocked, install WSL manually with Microsoft's [offline install guidance](https://learn.microsoft.com/en-us/windows/wsl/install#offline-install). Download the latest official WSL `.msi` package from the [Microsoft WSL releases page](https://github.com/microsoft/WSL/releases/latest), choose the matching `.x64.msi` or `.arm64.msi`, install it, reboot if Windows requests it, then rerun the bootstrap script. Use the same repair flow if the bootstrap says "Windows Subsystem for Linux could not be verified" and reports a nonzero `wsl --status` exit code.
### Bootstrap says "Windows reports that WSL 2 cannot start yet"
The bootstrap script attempts `wsl --install --no-distribution` automatically when `wsl --status` reports that WSL 2 cannot start. If the repair command succeeds, reboot when prompted and let the bootstrap resume after sign-in. If the message persists after repair and reboot, enable virtualization in firmware and confirm that the Virtual Machine Platform optional component is enabled. Manual WSL installation only helps when the WSL runtime is missing or the online installer is blocked.
### `wsl -d Ubuntu` says "There is no distribution with the supplied name"
The Ubuntu package was installed with `--no-launch` but never registered, or Windows finished the install command before the distribution appeared in `wsl -l`. When this happens during the NemoClaw bootstrap, the script prints a sanitized `WSL install output` block. PowerShell transcript headers, footers, temporary transcript paths, and status-file paths are redacted before display so you can paste the useful WSL output into a bug report with less local machine metadata.
If the sanitized output says a reboot is required, reboot and rerun the bootstrap. If it does not request a reboot, register the distro manually or reinstall without `--no-launch`:
```bash
wsl --unregister Ubuntu
wsl --install -d Ubuntu
```
### Bootstrap says a Docker executable "is not signed by a trusted publisher"
The bootstrap script runs elevated, so before it launches `Docker Desktop.exe` or uses `docker.exe`, it checks the resolved executable's Authenticode signature and refuses to run one that is not validly signed by Docker. The script accepts `Docker Inc` as the certificate subject common name. If Docker changes the signer identity, the script refuses the executable until maintainers verify the signer on an official Docker download and update the allowlist. For a current-user installation, the administrator child completes the system changes and returns to the original non-elevated PowerShell process before the script starts Docker Desktop or uses its CLI. If you started the script from an elevated PowerShell window, rerun it from a normal PowerShell window so it can use the current-user installation without administrator privileges. Reinstall Docker Desktop from [docker.com](https://www.docker.com/products/docker-desktop/) or `winget install --id Docker.DockerDesktop`, then rerun the bootstrap script. If reinstalling does not clear the warning, treat the existing executable as untrusted and do not run it manually either.
The script continues after this warning instead of stopping, so the [`docker info` fails inside WSL](#docker-info-fails-inside-wsl) symptom below can appear a few minutes later even though the real cause is the untrusted executable, not WSL integration.
### `docker info` fails inside WSL
Confirm that Docker Desktop is running and that WSL integration is enabled for Ubuntu (Settings > Resources > WSL integration). Then restart WSL:
```bash
wsl --shutdown
wsl -d Ubuntu
docker info
```
### Windows-host Ollama is installed but not shown during onboarding
When NemoClaw runs inside WSL, it checks the Windows-host Ollama HTTP endpoint, the Windows `ollama.exe` process, the listener address, and HTTP `Host` validation.
It reuses the Windows route only when the listener is loopback-only, Docker Desktop can reach it, and an untrusted `Host` value receives `403`.
The route also requires `DOCKER_HOST` to be unset and Docker to use its local `default` context.
If Ollama is installed but does not pass those checks, the wizard should still offer a restart action and the WSL-local install.
If the Windows-host option does not appear, confirm that PowerShell interop is enabled in WSL and that Windows can locate Ollama:
```bash
powershell.exe -NoProfile -Command "Get-Process ollama -ErrorAction SilentlyContinue"
```
If the process is missing, start Ollama from Windows and rerun onboarding.
If the process exists but the endpoint is unprotected or unreachable, use the restart action when the wizard offers it.
The action sets `OLLAMA_HOST=127.0.0.1:11434` and verifies that an untrusted `Host` value is rejected.
If Docker Desktop cannot reach that protected route, choose WSL-local Ollama.
If Docker uses another context or `DOCKER_HOST` selects an endpoint, switch to the local `default` context and unset `DOCKER_HOST`, or choose WSL-local Ollama.
Do not set `OLLAMA_HOST=0.0.0.0:11434`.
### Ollama inference fails or hangs in WSL
Ollama configures context length based on your hardware.
On some GPUs (for example RTX 3500), the default context length is not sufficient for OpenClaw. During onboarding, NemoClaw raises loaded-model context lengths below `16384` to `16384` when `NEMOCLAW_CONTEXT_WINDOW` is unset. Set the variable manually when you need a different value or when you run Ollama outside the managed onboarding path. Force a larger context length:
```bash
pkill -f 'ollama serve'
OLLAMA_CONTEXT_LENGTH=16384 ollama serve
```
Hermes requires at least `64000` tokens. During onboarding, NemoClaw verifies the loaded model's actual `context_length` through Ollama's `/api/ps` endpoint. Resumed onboarding and sandbox rebuilds warm the exact recorded Ollama model and repeat this check before reusing its route. Resume stops when that recorded model is missing, Ollama is unreachable, model warm-up fails, or the runtime context cannot be verified. If you set `NEMOCLAW_CONTEXT_WINDOW` above `64000`, the loaded model must provide at least that larger value. If the runtime value is below the requirement, NemoClaw queries Ollama's `/api/show` endpoint for the selected model's native context window. If the model's native context window is below the requirement, onboarding stops and tells you to select a model that meets the reported requirement. `OLLAMA_CONTEXT_LENGTH` cannot raise a model above its native context window. If the model can meet the requirement, or NemoClaw cannot read its native context window, onboarding instead shows the required `OLLAMA_CONTEXT_LENGTH` value for restarting the host daemon. A missing or malformed runtime value also produces the daemon restart guidance. `NEMOCLAW_CONTEXT_WINDOW` controls Hermes prompt budgeting; it does not raise the model's native context window or the Ollama daemon's runtime context, and it does not bypass this check. Use the daemon restart command only when onboarding shows a required `OLLAMA_CONTEXT_LENGTH` value. The example uses the Hermes minimum; replace `64000` with the larger value from onboarding when applicable.
```bash
pkill -f 'ollama serve'
OLLAMA_CONTEXT_LENGTH=64000 ollama serve
```
Verify that Ollama inference works:
```bash
echo "Hello" | ollama run
```
Replace `` with the model you selected during onboarding (for example `qwen3.5:4b`).
If `ollama serve` fails with `Error: listen tcp 127.0.0.1:11434: bind: address already in use`, check whether Ollama is configured for automatic startup:
```bash
sudo systemctl status ollama
```
If it is active, stop it first, then start with the custom context length:
```bash
sudo systemctl stop ollama
OLLAMA_CONTEXT_LENGTH=16384 ollama serve
```
```bash
sudo systemctl stop ollama
OLLAMA_CONTEXT_LENGTH=64000 ollama serve
```
For additional troubleshooting, refer to the [Windows Setup](../get-started/additional-setup/windows-preparation) page.
For first-time OpenClaw setup, refer to the
[Quickstart](../get-started/quickstart).
For first-time Hermes setup, refer to [Quickstart with
Hermes](../get-started/quickstart).
## Podman
NemoClaw supports a native managed rootless Podman runtime provider.
For native rootless Podman on Linux, select the provider explicitly:
```bash
NEMOCLAW_GATEWAY_RUNTIME=podman $$nemoclaw onboard
```
Native preflight fails closed unless the current-user Podman socket and service, rootless engine, cgroups v2 hierarchy, bridge networking, DNS, host architecture, exact managed-image contract, and `lsof` listener enumeration are all qualified.
It accepts Podman's standard rootless layout: a `0660` current-user socket in a `0755` `podman` directory below a private current-user runtime directory, such as `/run/user/` at `0700`.
It rejects a world-writable socket, a group-writable socket without a private current-user directory in its path, and a path component writable by another user or group.
If `lsof` is unavailable, install it through the host's package-management policy and retry; NemoClaw does not install it.
Do not redirect it with `DOCKER_HOST`, `DOCKER_CONTEXT`, `CONTAINER_HOST`, or a named Podman connection; restore the current user's reported socket authority and retry with the explicit selector.
Native Podman supports stock managed-image onboarding only. Use the Docker provider for an explicit `--from ` or a read-only host mount.
## Hermes
The issues below are common problems you may encounter when running Hermes through `nemohermes`. For setup, refer to [Quickstart with Hermes](../../hermes/get-started/quickstart).
### Discord preset validation behind a proxy
The Hermes Discord policy allows its virtual-environment Python interpreter and the concrete system paths to which that interpreter resolves. It does not grant Node.js or `curl` Discord egress. A Node.js or `curl` failure is therefore not proof that the Discord preset is broken.
Use the Hermes virtual-environment Python runtime inside the affected sandbox for the manual REST probe:
```bash
nemohermes exec -- /opt/hermes/.venv/bin/python -c 'import urllib.error, urllib.request; u="https://discord.com/api/v10/gateway";
try: print(urllib.request.urlopen(u, timeout=20).status)
except urllib.error.HTTPError as error: print(error.code)'
```
To check Discord CDN egress, use the same sandbox-qualified Python runtime:
```bash
nemohermes exec -- /opt/hermes/.venv/bin/python -c 'import urllib.error, urllib.request; u="https://cdn.discordapp.com/";
try: print(urllib.request.urlopen(u, timeout=20).status)
except urllib.error.HTTPError as error: print(error.code)'
```
Any HTTP status means the Python process reached the endpoint; the exact status can vary by unauthenticated path. A transport error or OpenShell policy denial means the probe failed. If the REST probe works but the Discord channel is still unhealthy, inspect the gateway logs and blocked-request output with `openshell term`. Look for `gateway.discord.gg` connection or WebSocket upgrade failures instead of widening the preset.
### Hermes restart reports a config integrity failure
Hermes configuration is mutable.
A current Hermes supervisor does not refuse startup, restart, recovery, or automatic respawn only because `config.yaml` or `.env` changed.
It validates the secret boundary and records one stable snapshot before a replacement gateway consumes the config.
The native `mcp_servers` map in that snapshot is authoritative, including direct edits.
An unsafe config or metadata path reports `unsafe config path`.
A missing, malformed, mismatched, or raced hash or restart metadata file reports `config hash mismatch`.
Do not edit either hash file manually.
If the error reports `secret-boundary refusal`, inspect `/sandbox/.hermes/.env` for raw secret-shaped values.
Replace them through the supported credential flow so the file contains `openshell:resolve:env:` placeholders, then run `nemohermes recover`.
If the error reports `unsafe config path`, do not follow or edit the unsafe path in place. Rebuild from registered configuration to restore the managed layout.
If it reports `config hash mismatch`, do not repair the hash or restart metadata by hand. Rebuild from registered configuration:
```bash
nemohermes rebuild --yes
```
The Hermes entrypoint supervisor remains responsible for the gateway, dashboard, internal API relay, dashboard relay, and gateway log stream throughout recovery.
In the OpenShell-managed topology, that nonroot supervisor repairs failed auxiliaries continuously and recovers a gateway after four consecutive failed listener or HTTP health checks.
A fifth status-75 service-managed restart request within 60 seconds starts a cooldown.
The supervisor waits until the rolling window permits another restart, then relaunches the gateway automatically.
The host repairs only the host-side OpenShell forwards after the supervised processes pass health checks.
### Repeated Hermes Service-Managed Restarts
Valid Hermes config changes alone do not stop relaunch.
The supervisor adopts a stable, secret-boundary-safe config snapshot before launching a replacement gateway. If the startup log reports `HERMES_RUNTIME_PREPARATION_FAILED stage=`, the named environment, config, dependency, provider, or messaging input still failed after five bounded preparation attempts. Correct that input through its supported host-side configuration or credential flow, then restart the sandbox:
```bash
nemohermes logs --tail 50
nemohermes stop
nemohermes start
```
During a restart-rate cooldown, wait for the supervisor to relaunch the gateway automatically.
If the requests recur, inspect the gateway logs and correct the action or configuration that requests each restart.
Other unexpected gateway exit statuses fail immediately and do not consume this restart budget.
An unsafe startup layout also stops automatic relaunch. If layout repair reports the maximum log-entry count or repair depth, archive or remove old retained logs from a trusted host-side recovery environment, then stop and start the sandbox. For symbolic links, hard links, unsafe entry types, or paths that changed during repair, do not repair the path in place. Restore a trusted snapshot into a recreated sandbox or recreate from host-side onboarding configuration.
If an integrity mismatch is reported or the sandbox still cannot start, rebuild it from registered configuration:
```bash
nemohermes rebuild --yes
```
After recovery, verify that Hermes starts and the gateway health check passes.
### Port 8642 in a browser shows a blank page or `Cannot GET /`
`nemohermes onboard` forwards the sandbox's API port, which is `8642` when no other sandbox or host listener already holds it. Hermes serves an OpenAI-compatible API at that port, not a chat dashboard. A browser visit to `http://127.0.0.1:8642/` (or any non-API path) returns nothing renderable.
Confirm the agent is healthy with the API health endpoint instead:
```bash
curl -sf http://127.0.0.1:8642/health
```
Expected output:
```json
{ "status": "ok", "platform": "hermes-agent" }
```
Point an OpenAI-compatible client at `http://127.0.0.1:8642/v1` for chat completions. For terminal use, run `nemohermes launch `.
### Onboarding Reports Hermes Is Not Ready With an Unreachable API Port
Deployment verification probes the OpenAI-compatible API inside the sandbox and through its host-side API port forward. When the API answers inside the sandbox but its host-side API port forward is unreachable, verification reports the API port forward as failed. Onboarding then prints `Hermes is not ready` and exits with a nonzero status:
```text
✗ api: port forward not working (connection refused)
The OpenAI-compatible API on port 8642 is not reachable from the host. Run: nemoclaw recover
```
Use the forward recovery below only when the in-sandbox `gateway` check passed. The output omits passing checks, so confirm that it contains no `gateway` failure. If the output contains a `gateway` failure, follow that diagnostic first.
When the `gateway` check passed and only the `api` check failed, the API remains reachable inside the sandbox. Restore the receipt-owned host forward through NemoClaw:
```bash
nemoclaw recover
```
Then probe the allocated API port from the diagnostic:
```bash
curl -sS -o /dev/null -w '%{http_code}\n' --max-time 3 http://127.0.0.1:/health
```
An HTTP status of `200` or `401` means the managed host forward is reachable. Any other status, `000`, or connection failure requires fresh diagnostics. Run `$$nemoclaw status`, then follow the reported gateway or ForwardTcp ownership guidance. Do not create a separate `openshell forward start` process for a NemoClaw-managed port.
### `docker port` shows no mapping for 8642 even though forwarding works
OpenShell port forwards are host-side relays managed by the OpenShell gateway process, not Docker `-p` publish mappings on the sandbox container. `docker port openshell-hermes-` reflects only Docker-published ports, so it returns nothing for OpenShell-managed forwards even when the host bind is live.
Use OpenShell's own view as the supported acceptance signal:
```bash
openshell forward list # shows the host bind for each forwarded port
curl -sf http://127.0.0.1:8642/health # confirms the relayed endpoint answers
```
If `openshell forward list` does not show the sandbox's API port, run `nemohermes connect --probe-only` (or `nemohermes recover`) to ask the recovery path to re-establish every manifest-declared agent forward port that has gone missing. Recovery targets each sandbox's own ports. A second Hermes sandbox on the same host receives the next free API port, so check which sandbox owns each row before assuming a missing `8642` row belongs to the sandbox you are debugging. A Hermes sandbox onboarded before the API port became per-sandbox carries no allocated port and keeps `8642`. A Hermes sandbox created after that change receives its own port during onboarding, so a second sandbox needs no further action. To move a sandbox that predates the change onto its own port, set `NEMOCLAW_HERMES_API_PORT=` and rerun onboarding with `--recreate-sandbox`. Set `` to a free port from `8642` through `8652`. Onboarding rejects a value outside that range. It also rejects an in-range value that another sandbox or host listener already holds. A recreate keeps the sandbox's registry entry, so `--recreate-sandbox` without the variable keeps the recorded port.
### Install reports `Could not restore the Hermes forward`
The installer reads the sandbox's recorded API port from the sandbox registry before it restores the API forward. It exits with that message when any of these conditions is true:
- The `node` binary is missing.
- The registry file is missing, does not parse, or records no entry for the sandbox.
- The recorded port is not an integer from `8642` through `8652`.
A sandbox registered before the API port became per-sandbox records no port, and the installer uses `8642` for it. Install `node`, or rerun `nemohermes onboard` to register the sandbox again. Then run `nemohermes recover` to re-establish the forward.
### `nemohermes` reports `Sandbox 'X' already exists as OpenClaw`
Each sandbox name maps to exactly one agent type. If a sandbox named `X` was created with the default OpenClaw agent, a later `nemohermes onboard` for the same name exits with:
```text
Sandbox 'X' already exists as OpenClaw.
nemohermes is onboarding Hermes for this sandbox name.
Side-by-side agents are supported, but each sandbox name has one agent type.
```
Pick a distinct sandbox name (the Hermes default is `hermes`; a common pattern is `my-hermes`) so Hermes and OpenClaw sandboxes can coexist on the same host. To convert an existing sandbox to Hermes instead, destroy and re-onboard:
```bash
$$nemoclaw destroy
NEMOCLAW_AGENT=hermes nemohermes onboard
```
### `nemohermes: command not found` immediately after install
`nemohermes` is a thin shim installed alongside `$$nemoclaw` that pre-selects the Hermes agent. The installer drops the shim in the same directory as `$$nemoclaw`; if `$$nemoclaw` is on `PATH` but `nemohermes` is not, the shim symlink was skipped.
Verify the install:
```bash
command -v nemoclaw
command -v nemohermes
```
If only `$$nemoclaw` resolves, re-run the installer with `NEMOCLAW_AGENT=hermes` set so the shim is published:
```bash
export NEMOCLAW_AGENT=hermes
curl -fsSL https://www.nvidia.com/nemoclaw.sh | bash
```
Equivalently, every `nemohermes ` invocation is `NEMOCLAW_AGENT=hermes nemoclaw `.
### Choosing between OAuth and API key for the Hermes Provider
The Hermes Provider supports two authentication paths during onboarding. Pick OAuth when you have a Nous Portal account and an interactive terminal; pick API key when you have a long-lived `NOUS_API_KEY` and want a non-interactive flow.
Set the method explicitly so the wizard skips the prompt:
```bash
# OAuth (default; interactive)
export NEMOCLAW_HERMES_AUTH_METHOD=oauth
nemohermes onboard
# API key (non-interactive)
export NEMOCLAW_HERMES_AUTH_METHOD=api-key
export NOUS_API_KEY=nous_...
nemohermes onboard --non-interactive
```
`NEMOCLAW_HERMES_AUTH_METHOD` accepts `oauth`, `nous-portal-oauth`, `api-key`, and `nous-api-key`. The `NEMOCLAW_HERMES_AUTH` and `NEMOCLAW_NOUS_AUTH_METHOD` variables are back-compatible aliases.
If OAuth is selected and onboarding cannot open the host's default browser (a headless host or SSH session), the device-code prompt still prints the verification URL and user code to the terminal. Copy them to a browser on any other machine to complete the flow.
### API client returns `401 Unauthorized` against port 8642
Hermes uses bearer-token header authentication for client requests, not an OpenClaw-style URL fragment. A request without an `Authorization: Bearer ` header (or with an OpenClaw `#token=` fragment appended to the URL) is rejected with `401`.
Configure your OpenAI-compatible client to pass the Hermes API key in the `Authorization` header. Stored credentials (including `NOUS_API_KEY` and `OPENAI_API_KEY`) are listed by:
```bash
nemohermes credentials list
```
Reset a specific provider's credentials with `nemohermes credentials reset ` and re-onboard if the stored value is wrong.
### Brave Search is unsupported under Hermes
Hermes does not have a NemoClaw Brave Search backend. Adding the `brave` preset to a Hermes sandbox opens Brave's endpoints but does not configure Hermes to use the credential. Use Tavily Search through NemoClaw onboarding instead.
```bash
NEMOCLAW_WEB_SEARCH_PROVIDER=tavily \
TAVILY_API_KEY= \
nemohermes onboard --recreate-sandbox
```
NemoClaw writes `web.backend: tavily`, applies the `tavily` policy preset, and configures request-body credential rewriting for Hermes. If the same onboarding run selected the Nous-managed web gateway, Tavily replaces `nous-web` while selected Nous image, audio, browser, and code tools remain enabled.
### Tavily authentication fails after registration or rebuild
A Hermes Tavily provider needs both the `tavily-hermes-v1` profile and the current credential reference issued by OpenShell.
Older images can save an unversioned `TAVILY_API_KEY` reference in `/sandbox/.hermes/.env`.
Hermes loads that file over its process environment, replacing the current revision-scoped reference even after the correct provider is registered.
A copied reference from an older provider can fail for the same reason.
Current Hermes images omit `TAVILY_API_KEY` from generated gateway and dashboard dotenv files and retain the process-injected reference.
Follow [Repair a Hermes Tavily Provider](../security/credential-rotation#repair-a-hermes-tavily-provider) to preserve supported state and rebuild with the current image.
Keep an already-correct provider rather than resetting it solely to repair the old dotenv file.
Do not save raw keys in the sandbox or broaden the provider's binary allowlist.
The onboarding probe uses Hermes Python and the issued reference in a JSON `api_key` request.
It refuses stale dotenv overrides and raw credentials, and reports successful egress only for HTTP `200` with search results.
Configuration or network failures remain advisory; a detected raw credential blocks handoff.
Confirm a real Hermes search after rebuilding.
### Langfuse works with raw keys but not OpenShell placeholders
Do not replace an OpenShell placeholder with a raw Langfuse key in `~/.hermes/.env`.
That workaround bypasses the sandbox credential boundary.
OpenShell 0.0.116 resolves static credentials only at endpoints declared by their provider profile.
Older provider-backed Langfuse setups could rely on destination-independent placeholder rewriting, so they can stop working after a NemoClaw upgrade or sandbox rebuild even though the Hermes placeholder validator still accepts the value.
For Langfuse Cloud, register the two keys through the checked-in endpoint profile and rebuild the sandbox:
```bash
export LANGFUSE_PUBLIC_KEY=pk-lf-...
export LANGFUSE_SECRET_KEY=sk-lf-...
nemohermes credentials add -langfuse \
--type langfuse-hermes-v1 \
--credential LANGFUSE_PUBLIC_KEY \
--credential LANGFUSE_SECRET_KEY
unset LANGFUSE_PUBLIC_KEY LANGFUSE_SECRET_KEY
nemohermes rebuild
```
If `agent.log` says the keys look like placeholders, the running image does not contain the current Hermes compatibility patch.
Rebuild onto the current managed image before changing credentials.
If Hermes accepts the placeholders but OpenShell reports `credential_unavailable` or an endpoint mismatch, inspect the provider type and endpoint profile.
The provider must expose the exact `LANGFUSE_PUBLIC_KEY` and `LANGFUSE_SECRET_KEY` names and authorize the configured Langfuse host.
For self-hosted Langfuse, import a separate reviewed profile that matches the exact HTTPS origin, including its port, before registering the provider.
Set only the non-secret `HERMES_LANGFUSE_BASE_URL` in `~/.hermes/.env`.
Remove any copied `HERMES_LANGFUSE_PUBLIC_KEY` or `HERMES_LANGFUSE_SECRET_KEY` line because a saved credential-handle placeholder can shadow the current process-injected value after rotation or rebuild.
### Re-onboarding asks every messaging prompt again
`nemohermes onboard --resume` against a Hermes sandbox that was originally onboarded with Telegram, Discord, and Slack credentials re-prompts for each channel's bot token and per-channel settings rather than reusing the stored values. This is tracked in [#3581](https://github.com/NVIDIA/NemoClaw/issues/3581). For unattended re-onboards, export the messaging env vars first so the wizard skips the prompts:
```bash
export TELEGRAM_BOT_TOKEN=...
export DISCORD_BOT_TOKEN=...
export SLACK_BOT_TOKEN=...
export SLACK_APP_TOKEN=...
nemohermes onboard --resume --non-interactive
```