Pulsar-WaW wins on mixed backend sizes, not on equal ones
We ran WaW and Pulsar-WaW on the five-region development fleet in two shapes with the same 220 GPUs: a heterogeneous fleet of 20 backends with 2-20 GPUs each, and a homogeneous fleet with 11 GPUs per backend. Each fleet run is matched by a simulator run with the same workload, engine model, seed and routing configuration.
On the heterogeneous fleet, Pulsar-WaW delivered 37-38% more goodput than WaW at 450 requests/s across three seeds, and 9-16% more at 350, with higher SLO attainment, cache reuse and lower tail latency in every pair. On the homogeneous fleet Pulsar-WaW goodput was 0.1% lower than WaW's at 450 and 1.3% lower at 550, and WaW had the lower tail latency.
Fleet
RPS
Policy
Runs
Goodput
Sim goodput
SLO %
Reuse %
TTFT p99 s
Heterogeneous
350
WaW
3
275.5-281.7
272.2-285.9
94.9-95.3
83.3-84.4
9.0-10.1
Heterogeneous
350
Pulsar-WaW
3
306.9-319.3
323.0-328.0
97.3-98.6
90.8-91.8
5.4-7.7
Heterogeneous
450
WaW
3
252.2-266.1
245.8-264.9
85.9-88.3
73.6-75.2
15.4-16.3
Heterogeneous
450
Pulsar-WaW
3
346.6-366.6
370.8-384.0
95.6-96.2
87.6-89.7
8.2-9.8
Homogeneous
450
WaW
1
436.9
431.2-437.8
100.0
93.8
1.9
Homogeneous
450
Pulsar-WaW
1
436.3
420.6-427.5
100.0
92.8
4.3
Homogeneous
550
WaW
1
377.0
365.8-399.8
93.5
85.6
11.7
Homogeneous
550
Pulsar-WaW
1
372.0
302.9-401.4
91.3
85.6
12.6
Bars are the mean of the fleet runs; black whiskers span the three seeds; dashes are the simulator runs for the same seeds. Half-scale engine.
No Pylon or MockDynamo container was CPU-throttled in any of these runs (maximum throttled fraction 0.000 across 40 containers per run), so the results reflect routing, not infrastructure limits.
STARGATE / ROUTING PERFORMANCE07 OCT 2026
Study design / What was held equal
Two fleet shapes, one workload, matched seeds
Region
Stargates
Heterogeneous GPU workers per backend
Homogeneous
usw2
3
7, 4, 2, 3
11, 11, 11, 11
ue1
3
14, 6, 12, 9
11, 11, 11, 11
ew1
3
19, 8, 10, 13
11, 11, 11, 11
an1
3
11, 16, 15, 20
11, 11, 11, 11
as2
3
18, 17, 11, 5
11, 11, 11, 11
Control
Setting
Engine (half scale)
Batched mock engine shared by MockDynamo and the simulator: step = 8.0 ms + 0.17 ms per decoding sequence + 0.1 ms per prefill token; 500,000 KV tokens per GPU; 25 sequences per GPU. Twice the step cost and half the KV of the full-scale engine, which moves saturation to rates the development nodes serve without CPU limits
Engine (full scale)
4.0 ms + 0.085 ms + 0.05 ms; 1,000,000 KV tokens per GPU (2026-10-06 runs, page 7)
Pylon
Max engine concurrency = 25 x GPU workers; fallback-window input TPS, maximum weighting; queue-mismatch admission 10,000 ms / 4.0; change-driven stats coalesced to 10 ms; 2-CPU limit
Client
10 s TTFT SLO and maximum wait; 30 s timeout; up to 3 attempts per turn
Workload
Growing sessions: 4,000-token system prompt, 200-2,000 user and 128-256 output tokens per turn, 5-40 turns, 5 s mean think time; 240 s warmup then 240 s measured
Runs
Heterogeneous: 350 and 450 RPS, seeds 1-3, policy order alternated. Homogeneous: 450 and 550 RPS, seed 1. Every MockDC backend restarts before each run and all 15 Stargate replicas must report 20 backends before traffic starts
Instrumentation
One backend per node; CFS throttling counters and each Pylon's input TPS captured at the start and end of the measured window; record files checked line by line against the driver pods
Policy
Configuration
WaW
Affinity ring with 150 virtual nodes per backend, 1 affine backend, 300 ms wait, max_queued 4, utilization comparator
Pulsar-WaW
Rendezvous weights from maximum input TPS; 300 ms affinity wait; bands widen every 300 ms; max_queued 4; overflow only to free engine slots; 4 s queue bounds
STARGATE / ROUTING PERFORMANCE07 OCT 2026
Heterogeneous fleet / Every run
Pulsar-WaW leads in all six seed pairs
Policy
RPS
Seed
Offered
Goodput fleet / sim
SLO fleet / sim
Reuse
TTFT p99
Failures
WaW
350
1
291
275.9 / 284.8
94.9% / 95.5%
83.3%
10.1 s
2,856
Pulsar-WaW
350
1
324
319.3 / 326.4
98.6% / 98.7%
91.8%
5.7 s
935
WaW
350
2
289
275.5 / 285.9
95.3% / 95.3%
83.6%
9.0 s
2,842
Pulsar-WaW
350
2
318
312.4 / 323.0
98.2% / 98.7%
91.7%
5.4 s
1,128
WaW
350
3
296
281.7 / 272.2
95.2% / 93.2%
84.4%
9.2 s
2,918
Pulsar-WaW
350
3
315
306.9 / 328.0
97.3% / 98.2%
90.8%
7.7 s
1,646
WaW
450
1
301
266.1 / 255.0
88.3% / 85.9%
75.2%
15.4 s
4,828
Pulsar-WaW
450
1
381
366.6 / 370.8
96.2% / 96.3%
89.7%
8.2 s
2,942
WaW
450
2
292
253.7 / 264.9
86.8% / 88.1%
74.2%
16.1 s
5,552
Pulsar-WaW
450
2
362
346.6 / 384.0
95.7% / 97.6%
88.3%
9.7 s
2,964
WaW
450
3
293
252.2 / 245.8
85.9% / 84.9%
73.6%
16.3 s
5,681
Pulsar-WaW
450
3
363
347.1 / 381.3
95.6% / 97.1%
87.6%
9.8 s
3,077
WaW on the fleet stays inside the simulator's seed range at both rates: 252.2-266.1 goodput at 450 against 245.8-264.9 simulated. Pulsar-WaW runs below its simulator prediction, 346.6-366.6 against 370.8-384.0 at 450, but its lowest seed still beats WaW's best by 81 requests/s.
Failures are almost all client timeouts after three attempts. Each run also has a handful of HTTP 502 responses (at most a few dozen in about 80,000 requests) returned before a backend was selected.
STARGATE / ROUTING PERFORMANCE07 OCT 2026
Heterogeneous fleet / Why
WaW overloads the small backends
TTFT p50 of each backend over the measured window against its GPU count, seed 1.
WaW's affinity ring gives every backend the same share of sessions whatever its size. At 450 requests/s its 2-4 GPU backends reached up to 5.8 s TTFT p50 and evicted cached sessions, so reuse across the fleet fell to 75.2%. Pulsar-WaW weights backends by measured throughput and kept those backends at or below 3.1 s p50 with 89.7% reuse.
The 2-4 GPU backends hold 9 of the 220 GPUs. In the measured window, WaW's sessions pinned to them stalled: 41 requests succeeded there and 930 attempts hit the 30 s client timeout, so those sessions made almost no progress. Under Pulsar-WaW the same backends completed 1,296 requests with 800 timed-out attempts, and the larger backends carried the rest.
STARGATE / ROUTING PERFORMANCE07 OCT 2026
Homogeneous fleet / Control
On equal backends WaW is as good or slightly better
Policy
RPS
Seed
Offered
Goodput fleet / sim
SLO fleet / sim
Reuse
TTFT p99
Failures
WaW
450
1
437
436.9 / 437.8
100.0% / 100.0%
93.8%
1.9 s
0
Pulsar-WaW
450
1
436
436.3 / 420.6
100.0% / 99.3%
92.8%
4.3 s
0
WaW
550
1
403
377.0 / 365.8
93.5% / 92.3%
85.6%
11.7 s
4,634
Pulsar-WaW
550
1
408
372.0 / 302.9
91.3% / 82.3%
85.6%
12.6 s
6,806
With every backend the same size, WaW's equal shares are already the right placement. Pulsar-WaW goodput was 0.1% lower than WaW's at 450 and 1.3% lower at 550, and Pulsar-WaW had the longer tail at both rates (4.3 s against 1.9 s at 450). At 550 Pulsar-WaW also returned 720 overloaded rejections.
Simulator sweep with the full-scale engine, seeds 1 and 2: lines join the seed midpoints, bars show the seed range. The dotted line is goodput equal to nominal rate.
The full-scale simulator shows the same split: on the heterogeneous fleet Pulsar-WaW holds goodput through 900 requests/s, where WaW has already collapsed, while on the homogeneous fleet WaW keeps up through 1,100 requests/s and Pulsar-WaW collapsed under one of the two seeds.
STARGATE / ROUTING PERFORMANCE07 OCT 2026
Pulsar weights / What the signal measures
Maximum input TPS mostly measures cache hits
Each point is one backend's Pylon max_input_tps divided by its GPU workers. The dotted line is the half-scale engine's prefill limit for uncached tokens per GPU.
Per GPU, the measured maximum varied from 50k to 153k tokens/s across heterogeneous backends (coefficient of variation 0.33) and from 70k to 106k across identical homogeneous backends (0.09). The engine can prefill at most about 10k uncached tokens/s per GPU, so most of the measured throughput is cached input that never reaches prefill.
Total weight still grows with backend size, but sub-linearly: per GPU, the smallest backends reported about three times the throughput of the largest, so the weight favors small backends relative to their capacity. It is still far closer to capacity than WaW's equal shares, which kept the small backends serving at these rates instead of stalling. On equal backends the weight adds noise without information. That noise did not produce uneven placement in these runs: requests per backend varied by a coefficient of 0.08 under Pulsar-WaW and 0.09 under WaW at 450 requests/s. The cause of Pulsar-WaW's longer homogeneous tail is not isolated; its affinity wait and band widening are the next candidates.
A weight based on uncached prefill throughput, or on configured capacity, would separate backend capacity from cache state. That is a follow-up, not something these runs measured.
STARGATE / ROUTING PERFORMANCE07 OCT 2026
Full-scale engine / 2026-10-06
Earlier full-scale runs: same ranking, unresolved collapse
Policy
RPS
Offered
Goodput fleet / sim
SLO fleet / sim
Reuse
TTFT p99
Failures
Pylon peak cores
WaW
250
245
244.8 / 249.1
100.0% / 100.0%
92.5%
2.5 s
0
0.73
Pulsar-WaW
250
250
249.6 / 249.3
100.0% / 100.0%
94.0%
0.6 s
0
0.72
WaW
550
501
498.9 / 531.7
99.6% / 99.8%
85.5%
6.4 s
148
0.87
Pulsar-WaW
550
520
520.2 / 549.8
100.0% / 100.0%
92.6%
2.5 s
0
0.86
WaW
900
532
319.5 / 567.3
60.1% / 88.3%
55.8%
19.5 s
26,151
0.85
Pulsar-WaW
900
597
380.2 / 872.9
63.7% / 100.0%
45.8%
12.6 s
48,815
0.84
PowerOf2
250
204
200.2
98.3%
22.1%
10.9 s
185
0.73
These runs used the full-scale engine, one seed, a 1-CPU Pylon limit and no throttling counters, and backends could share a node after restarts. Pulsar-WaW led at 250 and 550 requests/s. At 900 both policies collapsed fleet-wide after about four minutes while the simulator predicted Pulsar-WaW would hold; Pylon ran close to its CPU limit, so infrastructure could not be ruled out. The half-scale campaign was designed to avoid that limit.
Simulator references on this page come from the current simulator with seed 1. PowerOf2 has no simulator row in this sweep.
STARGATE / ROUTING PERFORMANCE07 OCT 2026
Conclusions / Evidence record
Ship Pulsar-WaW for heterogeneous fleets
What the data supports
On fleets with mixed backend sizes, Pulsar-WaW is clearly better than WaW at and near saturation, in every seed pair and in both the simulator and the fleet.
On equal backends, WaW is as good or slightly better, with a shorter tail. Pulsar-WaW's advantage comes from capacity-aware placement and does not help when every backend is the same.
The simulator ranks the policies the same way as the fleet in every comparison. On the heterogeneous fleet it predicts WaW within its seed range and overpredicts Pulsar-WaW by 1-11%; on the homogeneous fleet it underpredicts Pulsar-WaW, most at 550 requests/s (303 simulated, 372 measured).
Next measurements
Isolate Pulsar-WaW's longer tail on equal backends: vary the affinity wait and band widening interval.
Test a weight based on uncached prefill throughput or configured capacity instead of maximum input TPS.
Repeat the 900 RPS full-scale comparison on larger MockDC nodes.
Repeat the decisive cases on real engines before a production recommendation.
Run record
Item
Detail
Half-scale runs
16 accepted runs on 2026-10-07; every record file matched its driver pod line by line
Full-scale runs
7 accepted runs on 2026-10-06; one earlier WaW 900 run excluded for driver open-file exhaustion
Images
Stargate, Pylon, MockDynamo and the fleet driver built from the reviewed routing stack
Limitations
The mock engine uses estimated step costs and is not calibrated against a real engine; cache reuse assumes a cached session covers its whole prompt. Each homogeneous rate has one fleet seed. Load balancers use an unseeded random generator, so simulator runs are not bit-for-bit reproducible.
Bundle file
Purpose
report.pdf / report.html
Portable report and self-contained browser version