#### Vetoes and Sheds
```promql
# Workers currently excluded, per model
smg_workers_overloaded
# Shed rate by stage
sum by (stage) (
rate(smg_worker_overload_shed_total[5m])
)
```
#### Headroom per Worker
```promql
# Mean KV usage per worker (the token usage signal)
avg by (worker, model) (smg_engine_token_usage >= 0)
# Waiting requests per worker (the queue signal)
sum by (worker, model) (smg_engine_waiting_requests >= 0)
```
### Alerting Example
```yaml
groups:
- name: smg-overload-protection
rules:
- alert: WorkersOverloaded
expr: smg_workers_overloaded > 0
for: 5m
labels:
severity: warning
annotations:
summary: "{{ $value }} worker(s) for {{ $labels.model }} excluded by overload protection"
- alert: OverloadShedding
expr: sum(rate(smg_worker_overload_shed_total[5m])) > 0
for: 2m
labels:
severity: critical
annotations:
summary: "SMG is shedding requests with worker_overload_protection_shed"
```
---
## Production Tuning
### Recommended Configurations
#### :material-shield-check: KV Ceiling Only
The engine-agnostic default: exclude a worker at 90% KV usage.
```bash
smg launch \
--worker-overload-protection
```
**Use when**: Turning protection on for the first time
#### :material-tray-full: KV and Queue Ceilings
Also cap the waiting queue, sized from your own traffic.
```bash
smg launch \
--worker-overload-protection \
--worker-overload-waiting-requests 16
```
**Use when**: Latency-sensitive traffic, where a long queue is a failure even with KV room left
#### :material-server-network: Mixed Fleet
Gateway defaults plus `overload` blocks on the workers that saturate earlier.
```bash
smg launch --worker-overload-protection
# then register small workers with
# "overload": {"waiting_requests": 8}
```
**Use when**: Workers differ in GPU memory or batch size
#### :material-timer-outline: Faster Reaction
Poll more often, so vetoes set and clear sooner and `Retry-After` is shorter.
```bash
smg launch \
--worker-overload-protection \
--load-monitor-interval 5
```
**Use when**: Bursty traffic, and workers can take more frequent load polls
The values above are starting points, not recommendations for your hardware. Size the waiting-requests ceiling from `smg_engine_waiting_requests` at healthy peak load, and set it above that normal peak.
### Tuning Guidelines
| Symptom | Potential Adjustment |
|---------|---------------------|
| Requests shed while engines still have headroom | Raise the thresholds, or give larger workers their own `overload` block |
| Latency climbs but no worker is ever vetoed | Lower `--worker-overload-token-usage`, or add `--worker-overload-waiting-requests` |
| Bursts overshoot the ceiling before the veto lands | Lower `--load-monitor-interval`, or use `least_load` with `--least-load-max-waiting-requests`, which also counts requests dispatched since the last poll |
| Protection never engages on some workers | Check that they report load: `GET /loads` omits workers without a report (for example gRPC TRT-LLM and MLX), and a vLLM gRPC worker started with `--disable-log-stats` reports zeros |
| Cache-aware keeps piling onto a hot worker until it is vetoed | Set `--overload-token-usage-threshold` below `--worker-overload-token-usage`, so cache-aware spreads load before the veto fires |
| gRPC PD requests shed at stage `pd_admission` | Add decode capacity, or raise `--pd-admission-wait-secs` while keeping it under the engine's bootstrap deadline |
Overload protection sheds; it does not queue. Keep [rate limiting](rate-limiting.md) or the [priority scheduler](priority-scheduling.md) in front of it to bound what the gateway accepts, and make sure clients honor `Retry-After`.
---
## What's Next?
### :material-tray-full: Rate Limiting
Gateway-side concurrency limits and queuing, applied before routing.
[Rate Limiting →](rate-limiting.md)
### :material-priority-high: Priority Scheduling
Admit higher-priority traffic first when the gateway is at capacity.
[Priority Scheduling →](priority-scheduling.md)
### :material-scale-balance: Load Balancing
The load-aware policies that read the same load reports.
[Load Balancing →](../routing/load-balancing.md)
### :material-call-split: PD Disaggregation
Prefill/decode routing, where the decode admission window applies.
[PD Disaggregation →](../routing/pd-disaggregation.md)
### :material-api: Admin API
Read the gateway's cached load snapshot with `GET /loads`.
[Get Loads →](../../reference/api/admin.md#get-loads)
### :material-chart-box: Metrics Reference
Overload, PD admission, and engine load metrics.
[Metrics Reference →](../../reference/metrics.md)