--- name: status description: > Diagnose a Lego-RL run that is already in flight (or just finished): which run is alive, how far it has got, and whether its numbers are healthy. Reads the process table, the run log's metric lines and the trials directory, then checks the metrics against this cluster's known failure signatures — R3 pearson collapse, lr=0, grad starvation, env_setup_failed avalanches, val fake-zeros, no-tool-call collapse, and for SAO/critic runs critic starvation and the all-negative-batch collapse — and says which one matches. Read-only, local host only: never kills, never restarts, never edits. Triggers on "how's the run doing", "what step is it on", "is reward going up", "is this run broken", "diagnose the training run". --- # /rl:status — is the live run healthy? Read-only diagnosis of a run in flight. Answers three things in order: **which run · how far · is it sick**. Never kills, restarts, cleans or edits anything — a wrong intervention here costs more than a slow answer. ## Step 1 — Which run? ```bash bash scripts/lib/live_probe.sh train 2>&1 | grep -E '^(WARN|OK) +(job|gpu):' ls -t logs/*.log | head -5 ``` Identify the live run from the `job:trainer` / `job:runner` lines, then map it to its log. **Do not assume the log is under `logs/`.** `scripts/templates/verl/common.env` derives `TRAIN_LOG=${HARBOR_LOG_DIR}/${TRAINER_EXPERIMENT_NAME}.log`, and a real config overrides `HARBOR_LOG_DIR` to a per-experiment directory under the shared trials root; only the template default lands in `/logs`. Get the real path from the runner itself, in this order: ```bash # 1. the runner printed it at startup (works even for a run launched by hand) grep -hoE 'train log: +\S+' logs/launch_*.log *.out 2>/dev/null | tail -3 # 2. or resolve it from the config without launching anything bash scripts/train/train.sh --dry-run 2>&1 | grep -E 'train log|vLLM log|trials' # 3. or find what is actually being written right now find "$(dirname "$HARBOR_TRIALS_DIR")" -mindepth 3 -maxdepth 3 -type d -name logs \ -mmin -30 2>/dev/null | head # or point it at your trials root ``` A run launched by hand as `nohup bash scripts/train/train.sh > foo.out` leaves `foo.out` wherever the launcher's cwd was — usually the repo root, not `logs/`. It holds the launch summary *plus* the same `tee`d stream, so it is a superset of `TRAIN_LOG` and equally good to read; the `train log:` line near its top is the fastest way to recover the canonical path. What it is **not** is a file the dashboard can see, since it is outside any served log dir. Because the repo lives on shared storage, that `.out` is visible from every box while the process is not: on a multi-node run only the ray-head node has the `train.sh` / `tee` / trainer processes. Seeing a growing log with **no matching pid here** means you are on the wrong node — not that the run died. Check `stat -c %Y` on the log before concluding anything from an empty `pgrep`. Multi-node: each node evaluates `EXP_NAME=…$(date …)` separately, so one launch produces `NNODES` exp dirs whose timestamps differ by seconds. Only the ray-head dir holds the trainer log; the others hold just `*_train_gpu_wandb.log`. Diagnose from the head's, but remember trial counts must be summed across **all** sibling dirs. If nothing is alive, say so and offer the **last finished** run instead; make it explicit in the report which of the two you are describing. If several runs are alive, list them and ask which one — do not merge metrics from two runs. ## Step 2 — How far has it got? ```bash LOG= # NOT assumed to be logs/.log grep -oE 'step:[0-9]+ ' "$LOG" | tail -1 # latest step grep -cE ' step:[0-9]+ - training/global_step' "$LOG" ls -t harbor_trials// 2>/dev/null | head -3 tail -40 "$LOG" ``` Report: latest step, wall-clock since launch (`ps -p -o etime=`), average minutes/step, and whether the tail is still moving (compare `stat -c %Y "$LOG"` against now). **A log that has not been written to in >30 min while the process is alive is itself the finding** — that is the deadlock shape, not a slow step. ## Step 3 — Are the numbers healthy? Metrics live on the step lines as `key:value` pairs. Pull the latest step line and read the keys below (these names are exact — they come from the real logs): ```bash grep -E ' step:[0-9]+ - training/global_step' "$LOG" | tail -1 \ | grep -oE '(training/rollout_actor_probs_pearson_corr|actor/(lr|grad_norm|kl_coef|pg_clipfrac)|actor/rollout_corr/(kl|rollout_is_eff_sample_size|rollout_is_ratio_fraction_low)|critic/(rewards/mean|advantages/mean|vf_explained_var|vf_loss|grad_norm|lr)|num_turns/mean|trajectory_filter/[a-z_/]+|response_length/(mean|clip_ratio)):[0-9.e+-]+' ``` Tell a **SAO / critic run** apart first: its config block says `gae` with a `critic` line, and the log carries `critic/vf_explained_var`. On such a run `training/rollout_actor_probs_pearson_corr` and `actor/entropy` are **absent by design** (bypass mode: `old_log_probs == rollout_log_probs`, so pearson would be 1.0 by construction) — their absence is not the R3 signature. Read the `actor/rollout_corr/*` keys instead. | Metric key | Healthy | What a bad value means | |---|---|---| | `training/rollout_actor_probs_pearson_corr` | ≈ 0.999 (≥ 0.99) | **R3 routing replay is misaligned** — training on corrupted logprobs. The single most important gate; a run below this is already wasted. | | `actor/lr` | = the configured lr | `0` → the fully-async + cosine + `total_training_steps=-1` bug; the model is frozen. Runner forces `constant`, so a 0 here means something overrode it. | | `actor/grad_norm` | same order as prior runs (~0.2–0.5) | ~0.03 with very long responses = gradient starvation from token dilution, not a bug to fix mid-run. | | `critic/rewards/mean` | non-zero, trending up | Flat 0 from step 1 = infrastructure, not the model — go to the filter reasons below before touching hyperparameters. | | `num_turns/mean` | tens of turns | Collapsing toward ~1 with reward dropping = the model stopped emitting tool calls and just ends the episode; a real training pathology, not infra. | | `trajectory_filter/reason/env_setup_failed` | ~0 | Non-trivial count = pods cannot start: image unpullable, registry down, or a node missing its insecure-registry trust. | | `trajectory_filter/reason/timeout` | small fraction | A large share means the agent budget is too tight for these tasks, or env exec is stalling. | | `trajectory_filter/invalid_ratio` | < ~0.1 | High = most of the batch is being dropped; the effective batch is far smaller than configured. | | `response_length/clip_ratio` | low | High = responses hitting the window; the tail is being truncated. | | `val-core/…`, `val-aux/num_turns/…` | non-zero at test_freq steps | All-zero val while train reward is fine = the val split's images are unpullable, **not** a model regression. | | `critic/vf_explained_var` (SAO) | leaves <0 within ~20 steps, then 0.2–0.5 | Flat ≤ 0.3 for 50+ steps = the critic never converged; with `critic/grad_norm` far above `CRITIC_GRAD_CLIP` that is **critic starvation** (every update clipped down). Huge negatives on a step whose `critic/returns/min ≈ max` are a degenerate batch, ignore that step. | | `critic/grad_norm` (SAO) | median ~10 on 30B–35B, spikes to 30–70 | Alarm only on **three consecutive steps > 30**; a single spike (even 200+, e.g. an empty batch after sandboxes vanished) is not instability. | | `critic/advantages/mean` (SAO) | ≈ 0 with whitening on | Drifting negative for consecutive steps with whitening off = all-negative batches; the precursor of the think-spam collapse. | | `actor/rollout_corr/rollout_is_eff_sample_size` (SAO) | ≥ 0.99 | Well below = DIS is zeroing many tokens (staleness or backend mismatch); with `rollout_is_ratio_fraction_low` at an exact multiple of 1/batch the DIS mirror is missing and the run trains sequence-TIS. | | `actor/pg_clipfrac` (SAO) | **absent** | Present on a bypass-mode run = the actor is not in `bypass_mode`; DIS never reached the loss. | Also worth a line each when present: `fully_async/processing_time/tp99` (long tail), `fully_async/count/dropped_stale_samples` (staleness pressure), `rollout_corr/kl`. ## Step 4 — Match against known failure signatures Only claim a signature when its **specific** evidence is present. Say "no known signature matched" rather than forcing a match — a wrong diagnosis here sends the user chasing the wrong layer for hours. | Signature | Evidence that must be present | |---|---| | **R3 misalignment** | pearson well below 0.99 on recent steps | | **frozen model** | `actor/lr:0` | | **grad starvation** | `actor/grad_norm` an order below the run's own earlier steps, alongside very long `response_length/mean` | | **env avalanche** | `trajectory_filter/reason/env_setup_failed` climbing across steps; reward down in step | | **val fake-zero** | val metrics 0 while `critic/rewards/mean` is healthy | | **no-tool-call collapse** | `num_turns/mean` falling toward 1 over consecutive steps + reward falling; filter reasons *normal* | | **critic starvation** (SAO) | `critic/vf_explained_var` flat ≤ 0.3 for 50+ steps while `critic/grad_norm` sits well above the configured clip; val flat. Fix on the next run: `CRITIC_GRAD_CLIP=10`, a warm `CRITIC_MODEL_PATH`, `CRITIC_WARMUP=20` | | **all-negative-batch collapse** (SAO) | `critic/advantages/mean` negative on 3+ consecutive steps (whitening off) followed by `num_turns/mean` rising while reward falls — a reward-neutral tool (e.g. `think`) is being relatively reinforced. `GAE_WHITEN_ADVANTAGES=True` on the next run; roll back to before the drift | | **DIS not reaching the actor** (SAO) | `actor/pg_clipfrac` present, or `rollout_is_ratio_fraction_low` landing on exact multiples of 1/batch — the policy-loss mirror in `hydra_args.sh` was removed; the run is not SAO | | **deadlock / stall** | process alive, log mtime old, no new step line; check whether the tail sits in val or in a rollout wait | | **step slowdown** | minutes/step up sharply — compare the node/replica counts in the run's own config block before blaming the tasks | For anything that points off-box (registry, kyverno, node disk, image pulls), **report the symptom and stop**. This skill does not SSH, does not touch the cluster, and must not assert a cluster-side cause it cannot see from here — phrase it as "the symptom points at X; confirm on ", and let the user decide. ## Step 5 — Report Keep metric keys verbatim so they can be grepped. ``` ## harbor status — **<🟢 healthy | 🟡 at risk | 🔴 recommend stopping>** — stage step () · running · ~ min/step · log last written min ago procs runner= trainer= GPUs in use: reward critic/rewards/mean= (last steps: ) grads actor/grad_norm= actor/lr= pearson= critic vf_explained_var= grad_norm= ESS= (SAO runs only) traj num_turns/mean= invalid_ratio= env_setup_failed= timeout= val **Diagnosis** **Recommendations** 1. ``` Never end with an action you already took — this skill takes none. ## Guardrails - read-only: no `kill`, no `ray stop`, no restart, no config edit, no log deletion or rotation (a dangling symlink under a run dir breaks the webui) - local host only: no SSH, no `kubectl` mutation - never merge two runs' metrics into one report - never state a cause you did not read out of the log or the process table - when the evidence is thin, say the evidence is thin