specification: API Commons Rate Limits specificationVersion: '0.1' schema: https://raw.githubusercontent.com/api-evangelist/interface-research/main/schema/api-commons.yml#/$defs/RateLimits generated: '2026-08-30' method: searched source: https://agentgateway.dev/docs/standalone/latest/configuration/resiliency/rate-limits/ + .../operations/debug/ provider: AgentGateway providerId: agentgateway created: '2026-05-04' modified: '2026-08-30' description: >- Agentgateway imposes NO rate limits on callers, because there is no hosted agentgateway service to call - the product is a binary you run, and its own admin API is unauthenticated and unthrottled on loopback. Agentgateway is instead a rate LIMITER: the capability section below records the limiting it applies to traffic passing through it, which is what an integrator is actually asking about. THIS FILE REPLACES A SCAFFOLD. The previous version, from a bulk sweep dated 2026-05-04, asserted free/professional/enterprise tiers at 10/100/1000 requests per minute against an api-key scope, and X-RateLimit-* / RateLimit-Policy response headers. None of that was published by agentgateway. It has been removed. limit_count: 0 limits: [] headers: {} headers_note: >- No RateLimit-*, X-RateLimit-* or Retry-After response-header contract is documented for agentgateway's own limiter. This is a genuine gap in the product, and the single most useful runtime signal an agent is missing when it gets a 429 from an agentgateway-fronted route. Recorded as absent rather than assumed. responseCodes: throttled: 429 limiterUnavailable: 500 limiterUnavailableNote: >- With the default failureMode failClosed, an unavailable or erroring remote rate limit service causes agentgateway to DENY the request with 500 Internal Server Error rather than 429 - so a 500 from a rate-limited route is a limiter outage, not a gateway bug. failureMode failOpen allows traffic through instead. capability: note: Rate limiting agentgateway APPLIES to proxied traffic. Configuration, not a limit on you. types: - id: local name: Local rate limits storage: in-memory, per replica shared: false detail: >- Very low overhead. Counters are not shared between replicas nor across restarts, so the docs state plainly they are not appropriate where exact global counts are required, or for long windows such as monthly limits. config_fields: [maxTokens, tokensPerFill, fillInterval, type] - id: remote name: Remote rate limits storage: pluggable external data store shared: true protocol: Envoy Rate Limit Service v3 gRPC protocol_url: https://www.envoyproxy.io/docs/envoy/latest/api-v3/service/ratelimit/v3/rls.proto detail: >- Chosen deliberately so existing Envoy rate limiters can be reused; the Envoy project's reference rate limiter with Redis is the documented example. config_fields: [host, domain, failureMode, descriptors, type] example: https://github.com/agentgateway/agentgateway/tree/main/examples/traffic-ratelimiting-global dimensions: - id: requests name: Request-based default: true detail: Each request consumes 1 unit of capacity. - id: tokens name: Token-based detail: >- Each prompt or completion token consumes 1 unit. Evaluated in two phases. At request time, with `tokenize: false` (the default) the token count is unknown so the request is always allowed unless the limit is 0; with `tokenize: true` agentgateway estimates and rejects if the estimate exceeds the limit. At response time the provider's reported counts are reconciled against the limit - which means a response that blows the budget is still returned, and only SUBSEQUENT requests are limited. caution: >- Since 1.5 the input count charged against a token limit INCLUDES cache-read and cache-creation tokens, normalised across providers. A limit sized against Anthropic or Amazon Bedrock (which exclude cached tokens from their own reported input count) now fills sooner than it used to. descriptors: detail: >- Remote limits key on CEL-valued descriptor entries, so limits can be scoped per organization, per authenticated user, per API key or per any header - e.g. value 'request.headers["x-organization"]'. attaches_to: [route] related_controls: - name: LLM budget and spend limits detail: Per-key dollar or token budgets and per-key model access lists, introduced in 1.5. docs: https://agentgateway.dev/docs/standalone/latest/llm/cost-controls/ - name: Retries, timeouts, fault injection docs: https://agentgateway.dev/docs/standalone/latest/configuration/resiliency/ maintainers: - FN: Kin Lane email: kin@apievangelist.com