STARGATE / ROUTING PERFORMANCE07 OCT 2026
Summary / Decision

Pulsar-WaW wins on mixed backend sizes, not on equal ones

We ran WaW and Pulsar-WaW on the five-region development fleet in two shapes with the same 220 GPUs: a heterogeneous fleet of 20 backends with 2-20 GPUs each, and a homogeneous fleet with 11 GPUs per backend. Each fleet run is matched by a simulator run with the same workload, engine model, seed and routing configuration.

On the heterogeneous fleet, Pulsar-WaW delivered 37-38% more goodput than WaW at 450 requests/s across three seeds, and 9-16% more at 350, with higher SLO attainment, cache reuse and lower tail latency in every pair. On the homogeneous fleet Pulsar-WaW goodput was 0.1% lower than WaW's at 450 and 1.3% lower at 550, and WaW had the lower tail latency.

FleetRPSPolicyRunsGoodputSim goodputSLO %Reuse %TTFT p99 s
Heterogeneous350WaW3275.5-281.7272.2-285.994.9-95.383.3-84.49.0-10.1
Heterogeneous350Pulsar-WaW3306.9-319.3323.0-328.097.3-98.690.8-91.85.4-7.7
Heterogeneous450WaW3252.2-266.1245.8-264.985.9-88.373.6-75.215.4-16.3
Heterogeneous450Pulsar-WaW3346.6-366.6370.8-384.095.6-96.287.6-89.78.2-9.8
Homogeneous450WaW1436.9431.2-437.8100.093.81.9
Homogeneous450Pulsar-WaW1436.3420.6-427.5100.092.84.3
Homogeneous550WaW1377.0365.8-399.893.585.611.7
Homogeneous550Pulsar-WaW1372.0302.9-401.491.385.612.6
Bars are the mean of the fleet runs; black whiskers span the three seeds; dashes are the simulator runs for the same seeds. Half-scale engine.
Bars are the mean of the fleet runs; black whiskers span the three seeds; dashes are the simulator runs for the same seeds. Half-scale engine.

No Pylon or MockDynamo container was CPU-throttled in any of these runs (maximum throttled fraction 0.000 across 40 containers per run), so the results reflect routing, not infrastructure limits.

STARGATE / ROUTING PERFORMANCE07 OCT 2026
Study design / What was held equal

Two fleet shapes, one workload, matched seeds

RegionStargatesHeterogeneous GPU workers per backendHomogeneous
usw237, 4, 2, 311, 11, 11, 11
ue1314, 6, 12, 911, 11, 11, 11
ew1319, 8, 10, 1311, 11, 11, 11
an1311, 16, 15, 2011, 11, 11, 11
as2318, 17, 11, 511, 11, 11, 11
ControlSetting
Engine (half scale)Batched mock engine shared by MockDynamo and the simulator: step = 8.0 ms + 0.17 ms per decoding sequence + 0.1 ms per prefill token; 500,000 KV tokens per GPU; 25 sequences per GPU. Twice the step cost and half the KV of the full-scale engine, which moves saturation to rates the development nodes serve without CPU limits
Engine (full scale)4.0 ms + 0.085 ms + 0.05 ms; 1,000,000 KV tokens per GPU (2026-10-06 runs, page 7)
PylonMax engine concurrency = 25 x GPU workers; fallback-window input TPS, maximum weighting; queue-mismatch admission 10,000 ms / 4.0; change-driven stats coalesced to 10 ms; 2-CPU limit
Client10 s TTFT SLO and maximum wait; 30 s timeout; up to 3 attempts per turn
WorkloadGrowing sessions: 4,000-token system prompt, 200-2,000 user and 128-256 output tokens per turn, 5-40 turns, 5 s mean think time; 240 s warmup then 240 s measured
RunsHeterogeneous: 350 and 450 RPS, seeds 1-3, policy order alternated. Homogeneous: 450 and 550 RPS, seed 1. Every MockDC backend restarts before each run and all 15 Stargate replicas must report 20 backends before traffic starts
InstrumentationOne backend per node; CFS throttling counters and each Pylon's input TPS captured at the start and end of the measured window; record files checked line by line against the driver pods
PolicyConfiguration
WaWAffinity ring with 150 virtual nodes per backend, 1 affine backend, 300 ms wait, max_queued 4, utilization comparator
Pulsar-WaWRendezvous weights from maximum input TPS; 300 ms affinity wait; bands widen every 300 ms; max_queued 4; overflow only to free engine slots; 4 s queue bounds
STARGATE / ROUTING PERFORMANCE07 OCT 2026
Heterogeneous fleet / Every run

Pulsar-WaW leads in all six seed pairs

PolicyRPSSeedOfferedGoodput fleet / simSLO fleet / simReuseTTFT p99Failures
WaW3501291275.9 / 284.894.9% / 95.5%83.3%10.1 s2,856
Pulsar-WaW3501324319.3 / 326.498.6% / 98.7%91.8%5.7 s935
WaW3502289275.5 / 285.995.3% / 95.3%83.6%9.0 s2,842
Pulsar-WaW3502318312.4 / 323.098.2% / 98.7%91.7%5.4 s1,128
WaW3503296281.7 / 272.295.2% / 93.2%84.4%9.2 s2,918
Pulsar-WaW3503315306.9 / 328.097.3% / 98.2%90.8%7.7 s1,646
WaW4501301266.1 / 255.088.3% / 85.9%75.2%15.4 s4,828
Pulsar-WaW4501381366.6 / 370.896.2% / 96.3%89.7%8.2 s2,942
WaW4502292253.7 / 264.986.8% / 88.1%74.2%16.1 s5,552
Pulsar-WaW4502362346.6 / 384.095.7% / 97.6%88.3%9.7 s2,964
WaW4503293252.2 / 245.885.9% / 84.9%73.6%16.3 s5,681
Pulsar-WaW4503363347.1 / 381.395.6% / 97.1%87.6%9.8 s3,077

WaW on the fleet stays inside the simulator's seed range at both rates: 252.2-266.1 goodput at 450 against 245.8-264.9 simulated. Pulsar-WaW runs below its simulator prediction, 346.6-366.6 against 370.8-384.0 at 450, but its lowest seed still beats WaW's best by 81 requests/s.

Failures are almost all client timeouts after three attempts. Each run also has a handful of HTTP 502 responses (at most a few dozen in about 80,000 requests) returned before a backend was selected.

STARGATE / ROUTING PERFORMANCE07 OCT 2026
Heterogeneous fleet / Why

WaW overloads the small backends

TTFT p50 of each backend over the measured window against its GPU count, seed 1.
TTFT p50 of each backend over the measured window against its GPU count, seed 1.

WaW's affinity ring gives every backend the same share of sessions whatever its size. At 450 requests/s its 2-4 GPU backends reached up to 5.8 s TTFT p50 and evicted cached sessions, so reuse across the fleet fell to 75.2%. Pulsar-WaW weights backends by measured throughput and kept those backends at or below 3.1 s p50 with 89.7% reuse.

The 2-4 GPU backends hold 9 of the 220 GPUs. In the measured window, WaW's sessions pinned to them stalled: 41 requests succeeded there and 930 attempts hit the 30 s client timeout, so those sessions made almost no progress. Under Pulsar-WaW the same backends completed 1,296 requests with 800 timed-out attempts, and the larger backends carried the rest.

STARGATE / ROUTING PERFORMANCE07 OCT 2026
Homogeneous fleet / Control

On equal backends WaW is as good or slightly better

PolicyRPSSeedOfferedGoodput fleet / simSLO fleet / simReuseTTFT p99Failures
WaW4501437436.9 / 437.8100.0% / 100.0%93.8%1.9 s0
Pulsar-WaW4501436436.3 / 420.6100.0% / 99.3%92.8%4.3 s0
WaW5501403377.0 / 365.893.5% / 92.3%85.6%11.7 s4,634
Pulsar-WaW5501408372.0 / 302.991.3% / 82.3%85.6%12.6 s6,806

With every backend the same size, WaW's equal shares are already the right placement. Pulsar-WaW goodput was 0.1% lower than WaW's at 450 and 1.3% lower at 550, and Pulsar-WaW had the longer tail at both rates (4.3 s against 1.9 s at 450). At 550 Pulsar-WaW also returned 720 overloaded rejections.

Simulator sweep with the full-scale engine, seeds 1 and 2: lines join the seed midpoints, bars show the seed range. The dotted line is goodput equal to nominal rate.
Simulator sweep with the full-scale engine, seeds 1 and 2: lines join the seed midpoints, bars show the seed range. The dotted line is goodput equal to nominal rate.

The full-scale simulator shows the same split: on the heterogeneous fleet Pulsar-WaW holds goodput through 900 requests/s, where WaW has already collapsed, while on the homogeneous fleet WaW keeps up through 1,100 requests/s and Pulsar-WaW collapsed under one of the two seeds.

STARGATE / ROUTING PERFORMANCE07 OCT 2026
Pulsar weights / What the signal measures

Maximum input TPS mostly measures cache hits

Each point is one backend's Pylon max_input_tps divided by its GPU workers. The dotted line is the half-scale engine's prefill limit for uncached tokens per GPU.
Each point is one backend's Pylon max_input_tps divided by its GPU workers. The dotted line is the half-scale engine's prefill limit for uncached tokens per GPU.

Per GPU, the measured maximum varied from 50k to 153k tokens/s across heterogeneous backends (coefficient of variation 0.33) and from 70k to 106k across identical homogeneous backends (0.09). The engine can prefill at most about 10k uncached tokens/s per GPU, so most of the measured throughput is cached input that never reaches prefill.

Total weight still grows with backend size, but sub-linearly: per GPU, the smallest backends reported about three times the throughput of the largest, so the weight favors small backends relative to their capacity. It is still far closer to capacity than WaW's equal shares, which kept the small backends serving at these rates instead of stalling. On equal backends the weight adds noise without information. That noise did not produce uneven placement in these runs: requests per backend varied by a coefficient of 0.08 under Pulsar-WaW and 0.09 under WaW at 450 requests/s. The cause of Pulsar-WaW's longer homogeneous tail is not isolated; its affinity wait and band widening are the next candidates.

A weight based on uncached prefill throughput, or on configured capacity, would separate backend capacity from cache state. That is a follow-up, not something these runs measured.

STARGATE / ROUTING PERFORMANCE07 OCT 2026
Full-scale engine / 2026-10-06

Earlier full-scale runs: same ranking, unresolved collapse

PolicyRPSOfferedGoodput fleet / simSLO fleet / simReuseTTFT p99FailuresPylon peak cores
WaW250245244.8 / 249.1100.0% / 100.0%92.5%2.5 s00.73
Pulsar-WaW250250249.6 / 249.3100.0% / 100.0%94.0%0.6 s00.72
WaW550501498.9 / 531.799.6% / 99.8%85.5%6.4 s1480.87
Pulsar-WaW550520520.2 / 549.8100.0% / 100.0%92.6%2.5 s00.86
WaW900532319.5 / 567.360.1% / 88.3%55.8%19.5 s26,1510.85
Pulsar-WaW900597380.2 / 872.963.7% / 100.0%45.8%12.6 s48,8150.84
PowerOf2250204200.298.3%22.1%10.9 s1850.73

These runs used the full-scale engine, one seed, a 1-CPU Pylon limit and no throttling counters, and backends could share a node after restarts. Pulsar-WaW led at 250 and 550 requests/s. At 900 both policies collapsed fleet-wide after about four minutes while the simulator predicted Pulsar-WaW would hold; Pylon ran close to its CPU limit, so infrastructure could not be ruled out. The half-scale campaign was designed to avoid that limit.

Simulator references on this page come from the current simulator with seed 1. PowerOf2 has no simulator row in this sweep.

STARGATE / ROUTING PERFORMANCE07 OCT 2026
Conclusions / Evidence record

Ship Pulsar-WaW for heterogeneous fleets

What the data supports

Next measurements

Run record

ItemDetail
Half-scale runs16 accepted runs on 2026-10-07; every record file matched its driver pod line by line
Full-scale runs7 accepted runs on 2026-10-06; one earlier WaW 900 run excluded for driver open-file exhaustion
ImagesStargate, Pylon, MockDynamo and the fleet driver built from the reviewed routing stack

Limitations

The mock engine uses estimated step costs and is not calibrated against a real engine; cache reuse assumes a cached session covers its whole prompt. Each homogeneous rate has one fleet seed. Load balancers use an unseeded random generator, so simulator runs are not bit-for-bit reproducible.

Bundle filePurpose
report.pdf / report.htmlPortable report and self-contained browser version
results.json / results.csvRun-level values, per-backend tables, per-minute series, CPU, throttling and Pylon TPS snapshots
graphs/*.svg / *.pngStandalone figures
extract_results.py / build_report.py / render_pdf.pyExtraction, figures and rendering