# SAO on Reef This example implements the harness side of [Single-Rollout Asynchronous Optimization](https://arxiv.org/abs/2607.07508). SAO samples one rollout per prompt and grades each on its own. There is no comparison group and no barrier: a scored rollout joins the next optimizer step without waiting for siblings, and the next request is served by the updated weights. The method itself is the `sao` recipe package (`recipes/sao/`). This directory holds the loop around it: three IMOAnswerBench problems as Harbor tasks, a Harbor agent that runs six scored rollouts per problem through Reef, and `run.py`, which runs the tasks in order. The [`sao` recipe page](../../../../docs/user-guide/recipes/sao.rst) documents the recipe's configuration and runtime metrics, and [Evolve your model](../../../../docs/user-guide/evolve-your-model.rst) walks through the training stack this example starts. This README records the example's implementation details, its distance from the paper's protocol, and a completed comparison against GRPO at the paper's model scale. ```text harbor/ three IMOAnswerBench problems as Harbor tasks, run in order by run.py imo-4/ problem_idx 4 (gold: 2^{u-2}) imo-8/ problem_idx 8 (gold: -2023/2024^2) imo-12/ problem_idx 12 (gold: 1/2) task.toml metadata, timeouts, resource limits instruction.md the problem text and the \boxed{} instruction environment/ the Python image the verifier runs in tests/ test.sh runs the verifier grade.py extracts \boxed{}, checks it against the gold answer harness/ agent harness (imports reef_client, not reef) __init__.py lazily exports HarborAgent agent.py HarborAgent: six scored rollouts per problem, one report each grader.py \boxed{} extraction and the strict equivalence rule, shared with stream.py and evaluate.py report.py posts Harbor's verifier reward against the trial's receipts serve.yaml smoke stack config (batch 1): Reef + Ray + Slime/Megatron + SGLang, critic colocated serve-30b.yaml Qwen3-30B-A3B-Thinking-2507 on one 8-GPU node, fed by stream.py serve-30b-multi.yaml the same model with data-parallel training nodes and separate rollout engines, batch 128 run.py the loop, written out: solve, verify, report, task by task run.sh starts the Reef training stack, then runs run.py or SAO_DRIVER stream.py the paper-shaped streaming driver: one rollout per prompt, many prompts in flight export_problems.py writes the problems JSONL that stream.py and evaluate.py read evaluate.py held-out evaluation of a served model with the training grader plot_stream.py training reward and held-out accuracy of one streaming run plot_paper.py held-out accuracy and training dynamics against optimizer step (the Results figure) plot_curve.py held-out accuracy per checkpoint, one panel per benchmark pyproject.toml makes the harness importable results/ held-out accuracy and training curves of the batch-128 comparison ``` ## The harness `HarborAgent` receives one problem per Harbor trial as its `instruction`, looks up the gold answer, and runs `ROLLOUTS` (six) attempts at it. Each attempt is one chat request at `temperature=1.0, top_p=1.0` with a 2048-token generation window. The agent extracts the last `\boxed{}` from the completion and compares it with the gold answer under the strict equivalence rule the Harbor verifier uses: an exact match after whitespace and `$` are stripped, or a numeric evaluation of simple LaTeX (fractions, roots, π) within a relative tolerance of `1e-6`. The reward is binary, 1.0 for a correct answer and 0.0 otherwise. The rule is copied into `harness/grader.py` rather than imported from the task, because the harness is installed on its own into reef-eval's environment; `stream.py` and `evaluate.py` grade with the same module. The last completion is written to `/workspace/answer.txt`, where the Harbor verifier (`harbor/imo-*/tests/grade.py`) scores it independently and records the trial's reward. ## The changes needed for Reef The rollout loop in `HarborAgent.run` contains the two integration points: ```python response, agent_record_id = await asyncio.to_thread(self._ask_reef, instruction) completion = response["choices"][0]["message"]["content"] predicted = extract_answer(completion) score = 1.0 if answers_equal(gold, predicted) else 0.0 self._client.report(SCENARIO, {"score": score, "references": [agent_record_id]}, recipe=RECIPE) ``` 1. Send inference through a scenario-scoped Reef endpoint (`inference_with_record`) and keep the returned `agent_record_id` as the generation receipt. 2. After the local grader computes the reward, report it against that exact receipt. The dispatcher collects `batch_size` reports into one training step; this smoke config sets `batch_size: 1`, so each report is one step. A third piece runs after the trial. Harbor writes `result.json` when the verifier finishes, and a watcher thread posts the verifier's reward as one more report, referencing all six receipts (`harness/report.py`; the report id is derived from the trial id, so a repeated post changes nothing). That report does not train, because its references have already trained and the `sao` recipe does not accept multi-reference samples. It records the trial's verdict against the same receipts. The runtime flow is: ```text reef-eval starts one Harbor trial for the next problem -> the agent sends one OpenAI-compatible chat request through Reef -> the SGLang backend renders the prompt once and calls /generate -> Reef stores the sampled tokens, the loss mask, and the rollout log-probabilities -> the agent extracts \boxed{} and scores it against the gold answer -> the agent reports the score against that rollout's receipt -> SAOProcessor accepts the report and emits one ATIF TrajectoryItem -> the sao training objective hands Slime a batch of one -> the colocated critic computes values; skip-observation GAE builds the advantages -> Slime runs policy_loss with SAO's per-token DIS primitive, after two critic steps -> Megatron performs one optimizer step and synchronizes weights to SGLang -> the next rollout is served by the updated weights ``` ## Included paper problems The three tasks are IMOAnswerBench (`Hwilner/imo-answerbench`) problems `problem_idx` 4, 8, and 12, the slice the paper-scale run below used. They were chosen because the untrained model neither always solves nor always fails them under the strict grader, so the rewards carry a signal. Each `instruction.md` is the problem text followed by the instruction to put the final answer in `\boxed{}`. The verifier applies the same extraction and equivalence rule as the agent to `/workspace/answer.txt`. No LLM judge is involved anywhere in the loop. ## Setup (once) The training stack needs the GPU environment described in [Evolve your model](../../../../docs/user-guide/evolve-your-model.rst): Ray, the Slime driver, and CUDA builds of torch, SGLang, and Megatron. Inside it, from this directory: ```bash pip install -e . hf download Qwen/Qwen2.5-1.5B-Instruct --local-dir ~/models/Qwen2.5-1.5B-Instruct ``` `run.py` also needs `docker`: Harbor runs each task's verifier in its own container. The `docker run` line in Evolve your model does not provide that, so when the stack itself runs inside the reef image, extend it: ```bash docker run --gpus all --network host --ipc host --shm-size 32g -it \ -v ~/models:/root/models \ -v /var/run/docker.sock:/var/run/docker.sock \ -v "$REPO":"$REPO" -w "$REPO" \ reef bash # inside: apt-get update && apt-get install -y docker.io ``` Mount the repo at its host path (`-v "$REPO":"$REPO"`, not `/workspace/Reef`): Harbor's sibling containers bind-mount trial directories by path, and those paths must mean the same thing to the host docker daemon. ## Run ```bash ./run.sh ``` The launcher waits for Reef's health endpoint and stops waiting if Reef exits. Configure bridge startup with `training.ready-timeout` and HTTP startup with `reef.ready-timeout` in `serve.yaml`; startup errors are in `work/reef.log`. Exiting or interrupting the script also stops its Reef process. `run.sh` starts `reef serve -c serve.yaml` with its state under `./work`, waits for `/healthz`, and runs `run.py`. `serve.yaml` describes a two-GPU stack: one Megatron actor with the critic colocated on it, and one SGLang rollout engine, serving `Qwen2.5-1.5B-Instruct`. On the first start Reef loads the Hugging Face weights directly and writes the Megatron checkpoint that later starts load. Ray, Slime, Megatron, and SGLang take minutes to come up; `work/reef.log` has the service log if startup fails. The config omits `services`: Reef assembles the Slime driver and HTTP process, waits for the training bridge, and connects HTTP to the bridge's SGLang workers. The native training adapter captures rollout log-probabilities directly from the engine for SAO; no separate inference process or manual bridge wiring is needed. `training.options` still holds the model and algorithm settings. CLI overrides use the same paths, for example `--training.options.lr 0.000002`. Reef starts and stops the shared Ray runtime automatically; no `ray start` or fixed Ray port is needed. `run.sh` defaults the local cluster's GPU pool to `CUDA_VISIBLE_DEVICES=0,1`; override it at launch to choose different GPUs. To use an existing cluster, set `RAY_ADDRESS`; that cluster's node configuration controls GPU visibility, and Reef leaves it running on exit. Slime allocates the model GPUs; the local driver does not reserve them a second time. `run.py` is the loop, written out. For each task in order, reef-eval's `Lab.run` executes one episode: the agent runs its six rollouts, reporting each one as it is scored, then Harbor's verifier scores the last completion. The ordering is the experiment: task `N+1` is served by the weights task `N` produced. Nothing has to be exported. The service URL, token, scenario, rollout count, and generation window are constants at the top of `harness/agent.py`, and the port and token they use are the ones written in `serve.yaml`. The task list is `TASKS` in `run.py`. ### Reading the release chain and the metrics Each scored rollout publishes one `training` entry to the scenario's version chain: ```bash curl -s -H "Authorization: Bearer reef-local" \ http://127.0.0.1:8900/reef/scenarios/sao-smoke/releases ``` The runtime reports `pg_clipfrac` (the fraction of tokens the DIS mask removed), `train_rollout_logprob_abs_diff` (the mean per-token gap between the engine's and the trainer's log-probabilities, the quantity DIS masks on), `critic/explained_variance` (meaningful only with more than one sample per step), actor and critic `grad_norm`, and the asynchrony telemetry `sao/policy_lag_*`, `sao/queue_age_s_*`, and `sao/effective_token_rate`. Set `observability.wandb.enabled: true` in `serve.yaml` and export `WANDB_API_KEY` to keep them per committed step; `observability.wandb.directory` is where the run files go. ### A larger model Change `inference.model-path`, the GPU counts and parallelism flags, and `--seq-length` and `--rollout-max-response-len` in `serve.yaml`. The objective flags are the paper's reasoning-domain values and do not change with model size. `recipe.config.batch-size` and `training.config.global_batch_size` must stay equal, because each rollout sample is its own data-parallel unit. ## Paper fidelity The integration reproduces: - single-rollout sampling, one rollout per prompt with no comparison group and no slowest-sample barrier, with `batch_size` such rollouts per optimizer step (the recipe default is the paper's 128; this smoke config sets `batch_size: 1`, `--global-batch-size=1`, see below); - a value model colocated with the actor and two critic steps per actor step (`--critic-steps-per-actor=2`), trained at the paper's value learning rate of `5e-6` (`--critic-lr=5e-6`) with the paper's 10-step value warmup (`--num-critic-only-steps=10`: the first ten optimizer steps fit the zero-initialized value head before any policy update); - value targets from Monte-Carlo returns (λ = 1) and policy advantages from the length-adaptive λ with α = 1.5, built by skip-observation GAE in the training backend, so the Reef payload carries no advantages; - the DIS per-token loss with the reasoning-domain mask bounds 0.3 and 5.0, computed against the engine's rollout log-probabilities (`--use-rollout-logprobs`), which is why the deployment selects Reef's token-native SGLang chat backend; - sampling at `temperature=1.0, top_p=1.0`, a constant policy learning rate of `1e-6`, and no entropy bonus. ### Batch size is not group size The paper trains with "a batch size of 128, a group size of 1" (§4.1): one rollout per prompt, 128 prompts per optimizer step. Single-rollout is a statement about the group, not about the step. `run.py` and `serve.yaml` set `batch_size: 1` and `--global-batch-size=1`, which is a different estimator: each optimizer step is one REINFORCE sample with a critic baseline. That is the smallest instance of the paper's loop and a good smoke test, but the gradient of one rollout at the paper's learning rate is mostly noise, and two of its diagnostics are degenerate at that size: `critic/explained_variance` is `1 - Var(R - V) / Var(R)` over the batch, which is identically 0 for one sample, and the per-step reward is a coin flip. The streaming protocol below uses the paper's shape at a budget one node can afford. ### The streaming protocol `stream.py` keeps `SAO_IN_FLIGHT` requests open at all times, each on a problem drawn from a training pool, grades each completion with the verifier's rule and reports it against its receipt; `serve-30b.yaml` takes the optimizer step size from `SAO_BATCH` (the recipe's `batch-size` and the driver's `--global-batch-size`, which must agree). Held-out problem indices are never served, and `evaluate.py` scores the base and the trained weights on them afterwards with fresh samples. ```bash python export_problems.py work/imo_answerbench.jsonl # problem_idx, problem, gold SAO_BATCH=8 SAO_SERVE_YAML=serve-30b.yaml SAO_DRIVER=stream.py \ SAO_PROBLEMS=work/imo_answerbench.jsonl SAO_HOLDOUT=0,1,7,... SAO_POOL=2,5,... \ SAO_IN_FLIGHT=8 SAO_BUDGET=320 ./run.sh python evaluate.py --problems work/imo_answerbench.jsonl --indices 0,1,7,... --runs 8 \ --url http://127.0.0.1:30001/v1/chat/completions --label sao --out work/eval-sao.jsonl python plot_stream.py results/curve.png --records work/records/stream-*.jsonl \ --eval base=work/eval-base.jsonl sao=work/eval-sao.jsonl --batch 8 ``` The host-side scripts need `aiohttp`, `datasets` and `matplotlib`; install them with `pip install -e ".[stream]"` from this directory. The batch-128 SAO runs in Results use `serve-30b-multi.yaml` on an existing Ray cluster (`RAY_ADDRESS`): `SAO_TRAIN_NODES` 8-GPU nodes host the actor, tensor parallel 8 per node and data parallel across nodes with the critic colocated, and `SAO_ROLLOUT_GPUS` GPUs on other nodes host the engines, tensor parallel 4 each. The runs kept `SAO_CKPT_DIR` on shared storage with the retention cap set in the stack file, so Reef's coordinator sees the exports it verifies from whichever node it lands on. `SAO_PROGRESS_FILE` paces the driver on the trainer: write the number of completed optimizer steps into it from a host that sees `SAO_CKPT_DIR`, for example the highest exported step, `ls "$SAO_CKPT_DIR/hf" | sort -n | tail -1`. The problems JSONL has the same columns as `export_problems.py` writes, here filled from the DeepMath pools described in Results. The GRPO(+DIS) control is not part of the recipe: it ran the same topology with a loss family that keeps the DIS primitive and takes Slime's group-relative advantages without a critic, 16 prompts of 8 rollouts per step with each prompt's group reported together; the Results tables record its configuration and numbers. ```bash export RAY_ADDRESS=
:6379 SAO_TRAIN_NODES=4 SAO_ROLLOUT_GPUS=