specification: API Commons Rate Limits specificationVersion: '0.1' schema: https://raw.githubusercontent.com/api-evangelist/interface-research/main/schema/api-commons.yml#/$defs/RateLimits provider: Microsoft Azure API Management providerId: microsoft-azure-api-management created: '2026-05-04' # Provenance stamped 2026-08-11: this artifact was written by the API Evangelist # bulk sweep dated 2026-05-04, not harvested from the provider. See roadmap#35. method: generated modified: '2026-05-22' reconciled: true tags: - Rate Limiting - API Gateway - API Management - Microsoft Azure description: Azure API Management is itself a rate-limit/quota engine for downstream APIs. Built-in policies (rate-limit, rate-limit-by-key, quota, quota-by-key) let operators throttle by subscription key, IP, or arbitrary expression. Service-level capacity caps depend on the tier (scale units). sources: - https://learn.microsoft.com/en-us/azure/api-management/api-management-sample-flexible-throttling - https://learn.microsoft.com/en-us/azure/api-management/rate-limit-policy - https://learn.microsoft.com/en-us/azure/api-management/api-management-features headers: retryAfter: Retry-After responseCodes: throttled: 429 limits: - name: Consumption tier per-subscription scope: subscription metric: requests_per_minute limit: 'see policy definition; default no service-level cap' notes: Operators define limits via rate-limit / quota policies; consumption tier is metered per call rather than capped at the service level. - name: Tier scale units scope: service metric: varies limit: 'Developer 1, Basic 2, Basic v2 10, Standard 4, Standard v2 10, Premium 12 per region, Premium v2 30' notes: Each scale unit yields documented gateway throughput; configure rate-limit policies on top. - name: Rate-limit policy (per minute) scope: subscription/key/expression metric: requests_per_minute limit: 'configurable per policy' - name: Quota policy (per period) scope: subscription/key/expression metric: requests_per_period limit: 'configurable per policy' policies: - name: rate-limit and rate-limit-by-key description: Apply per-minute (or sub-minute via fixed-period) limits per subscription, key, or arbitrary key expression. Returns 429 with Retry-After when exceeded. - name: quota and quota-by-key description: Apply longer-window (renewable per hour, day, week, month) call or bandwidth quotas per subscription/key. - name: llm-token-limit description: AI-gateway TPM (tokens-per-minute) and token-quota policy for LLM/AI backends. Supports per-subscription, per-IP, or arbitrary counter keys, with optional prompt-token pre-estimation before forwarding to backend. - name: llm-emit-token-metric description: Emit prompt/completion/total token counts as Application Insights custom metrics with configurable dimensions (Client IP, API ID, user ID, etc.) for AI cost attribution. - name: Capacity vs throttling description: Tier scale units define peak capacity; operators must size scale units to match configured policy ceilings. - name: Honor Retry-After description: Clients should honor the Retry-After header returned with 429 responses when the gateway throttles a request. Backend circuit breakers in API Management likewise honor backend Retry-After headers for dynamic recovery. maintainers: - FN: Kin Lane email: kin@apievangelist.com