specification: API Commons Rate Limits specificationVersion: '0.1' schema: https://raw.githubusercontent.com/api-evangelist/interface-research/main/schema/api-commons.yml#/$defs/RateLimits provider: microsoft-azure-cache-for-redis providerId: microsoft-azure-cache-for-redis created: '2026-05-04' generated: '2026-09-17' modified: '2026-09-17' method: searched source: >- https://learn.microsoft.com/en-us/azure/azure-resource-manager/management/request-limits-and-throttling — "Understand how Azure Resource Manager throttles requests" (fetched 2026-09-17, HTTP 200). Replaces a 2026-05-04 bulk-sweep scaffold whose headers and tiers were invented. tags: - Rate Limiting - Quotas - Throttling description: >- Published throttling limits for the Azure Resource Manager control plane that serves management.azure.com — the host every operation in openapi/ is called on. Microsoft.Cache publishes no resource-provider-specific override, so the platform token-bucket limits below are the limits that apply to this API. scope_note: >- These are CONTROL-PLANE limits (creating, scaling, listing caches). They are unrelated to the Redis data plane, where throughput is governed by the cache's tier and size, not by an HTTP quota. algorithm: name: token bucket since: '2024' granularity: per region, per subscription, per service principal, per operation type note: >- The bucket is the maximum number of requests you can have in flight in a second; the refill rate is how quickly tokens return. Global per-subscription limits are 15x the individual service-principal limits and apply across all service principals. Limits may be smaller for free or trial customers. headers: retryAfter: Retry-After remaining: - x-ms-ratelimit-remaining-subscription-reads - x-ms-ratelimit-remaining-subscription-writes - x-ms-ratelimit-remaining-subscription-deletes - x-ms-ratelimit-remaining-tenant-reads - x-ms-ratelimit-remaining-tenant-writes - x-ms-ratelimit-remaining-subscription-resource-requests - x-ms-ratelimit-remaining-subscription-resource-entities-read - x-ms-ratelimit-remaining-tenant-resource-requests - x-ms-ratelimit-remaining-tenant-resource-entities-read limit: null reset: null policy: null note: >- There is no X-RateLimit-Limit / -Reset / RateLimit-Policy header. ARM publishes only the REMAINING counters above (read requests get the read counter, writes the write counter) and Retry-After on exhaustion. The resource-request / resource-entities-read counters appear only when a service overrides the default limit. responseCodes: throttled: 429 retryable_non_throttle: 429 note: >- Microsoft documents that some resource providers also return 429 for a temporary condition (e.g. RetryableErrorDueToAnotherOperation when another operation locks the target). Read error.details to tell a quota exhaustion from a transient lock. limits: - scope: per subscription (per service principal, per region) operation_type: reads burst: 250 limit: 25 window: second unit: requests note: bucket size 250, refill 25/sec - scope: per subscription (per service principal, per region) operation_type: writes burst: 200 limit: 10 window: second unit: requests - scope: per subscription (per service principal, per region) operation_type: deletes burst: 200 limit: 10 window: second unit: requests - scope: per tenant (per region) operation_type: reads burst: 250 limit: 25 window: second unit: requests - scope: per tenant (per region) operation_type: writes burst: 200 limit: 10 window: second unit: requests - scope: per tenant (per region) operation_type: deletes burst: 200 limit: 10 window: second unit: requests - scope: global per subscription (all service principals) operation_type: all limit: null window: second unit: requests note: >- Documented as 15x the individual service-principal limit for each operation type. No absolute number is published, so limit is null rather than guessed. limit_count: 6 recovery: strategy: >- Honour Retry-After before the next request; a request sent before it elapses is not processed and returns a fresh retry value. Read the remaining-counter header matching the verb to stay ahead of exhaustion. Metrics reads via */providers/microsoft.insights/metrics are called out by Microsoft as a common cause of subscription throttling — use the getBatch API instead. observed: probed: false note: >- Limits were read from Microsoft's documentation, not observed on the wire — every management.azure.com call needs an Entra token this pipeline does not hold.