# Troubleshooting Symptom → cause → fix. Everything here was hit and measured, not guessed. --- ## The process just dies. No Python traceback, no CUDA error. **Cause:** the kernel's OOM-killer, triggered by ComfyUI's pinned-memory pool. ComfyUI page-locks up to 90 % of system RAM on Linux (`MAX_PINNED_MEMORY = ram * 0.90` in `comfy/model_management.py`). Pinned pages are unswappable and unreclaimable, so under pressure the kernel has no option except to kill the process. **Confirm it:** ```bash sudo dmesg -T | grep -i oom-kill ``` You are looking for something like `python3 invoked oom-killer` with a large `anon-rss`. **⚠️ `docker inspect` will lie to you here.** A container killed by the kernel OOM-killer reports: ``` ExitCode=0 OOMKilled=false ``` and then quietly restarts. `State.OOMKilled` only reflects the *cgroup* OOM killer, not the system one. Trust `dmesg`, not Docker. **Fix:** add `--disable-pinned-memory` to the ComfyUI launch. Optionally grow swap afterwards as a net (pointless before, since pinned memory cannot be paged out). --- ## It OOMs but VRAM looks fine That is the expected signature, not a contradiction. On our box the 15-second run peaked at **18 884 MiB** VRAM — *lower* than the 5-second run's 23 716 MiB — and still died. More frames means a larger requested decode allocation, which makes ComfyUI evict **more** model weight out of VRAM and into host RAM. Lower VRAM usage can therefore mean *higher* host-RAM pressure. Watching `nvidia-smi` tells you nothing useful about this failure. Watch `free -m` instead. --- ## Where the time actually goes > ⚠️ **Corrected 2026-08-04.** An earlier version of this page claimed both frame counts sampled at > ~11 s/step and that the cost was "almost entirely in decode." That was wrong. Measured over 30+ > runs, sampling dominates and it scales *superlinearly* with frame count. | Clip | Frames | Per-step | Sampling total | Wall clock | |---|---:|---:|---:|---:| | 5 s @ 832x480 | 124 | 11.5 – 12.6 s | 3:49 – 4:08 | ~4.5 min | | 15 s @ 832x480 | 362 | 64.6 – 72.3 s | 21:34 – 23:52 | ~23 min | **Sampling is ~95 % of wall clock on a 362-frame clip** (22:14 of ~23 min). Decode is roughly a minute. 2.9× the frames costs ~5.6× per step — worse than linear, consistent with attention being quadratic in sequence length. So you *can* extrapolate from a short run, but scale by frames **superlinearly**, not by a flat per-step figure. **Consequence for optimisation:** anything that speeds up the DiT (step caching such as `EasyCache` / `LazyCache`, or sequence parallelism across GPUs) targets ~95 % of the runtime. Anything that targets decode is chasing ~5 %. --- ## `VAEDecodeTiled` doesn't change anything It is a literal no-op for H3. In `comfy/ldm/minimax/vae.py`: ```python def decode_tiled(self, z, **kwargs): return self.decode(z) ``` The H3 VAE tiles internally and unconditionally (256 px spatial, 17-frame temporal chunks), and `comfy/sd.py` sets `handles_tiling = True`, so ComfyUI routes tiled requests straight back to the plain decode. The tile parameters are not exposed as node inputs. --- ## Adding a second GPU didn't help Expected. ComfyUI has no tensor-parallelism for diffusion models — one model runs on one device and spills to CPU. `ComfyUI-MultiGPU`'s `CLIPLoaderMultiGPU device=cuda:1` genuinely does relocate the text encoder (measured: idle card went 306 MiB → 15 422 MiB). Host RAM still rose to 29 866 MB, because ComfyUI continues to stage a CPU copy regardless: ``` CLIP/text encoder model load device: cuda:1, offload device: cpu, current: cpu ``` Pinning is a memory *policy*. It is indifferent to how many GPUs you own. Fix the policy first; a second card is useful afterwards for a bigger canvas or for running two clips concurrently, not for this. ⚠️ There is an open MultiGPU issue where pinning a text encoder to a second card produces **entirely black video**. If you use it, verify frame content — see `scripts/verify_output.sh`. --- ## Flags that make this failure *worse* | Flag | Why | |---|---| | `--disable-smart-memory` | Its own help text: forces aggressive offload **to regular RAM**. Worst possible choice here. | | `--high-ram` | Bypasses the pinned-memory budget check entirely. | | `--reserve-vram` / `--vram-headroom` | Less VRAM available → more spilled into host RAM. | | `--cache-lru` | Documented as using more RAM/VRAM. | | `--lowvram` / `--novram` | No-ops under dynamic VRAM ("Doesn't do anything if dynamic vram is enabled"). | --- ## Model files show as empty in the loaders If ComfyUI runs in a container and your `models/` entries are symlinks pointing outside the mounted volume, they dangle silently — the loader dropdown is simply empty. Mount the symlink target path into the container as well. --- ## Container exits immediately with code 126 `/usr/bin/bash: /usr/bin/bash: cannot execute binary file` The base image already declares `ENTRYPOINT ["bash"]`, so passing `bash -c "..."` as the command doubles it and bash tries to execute itself as a script. Pass `-c "..."` alone. --- ## `docker run --gpus "device=2,3"` errors ``` cannot set both Count and DeviceIDs on device request ``` Use `--gpus all` combined with `-e CUDA_VISIBLE_DEVICES=2,3`. --- ## Out of memory on a unified-memory device (DGX Spark) Different cause: the OS page cache competes with the CUDA allocator for the same pool. After moving large files, drop the cache before launching: ```bash sync; echo 3 > /proc/sys/vm/drop_caches ```