specification: API Commons Rate Limits specificationVersion: '0.1' schema: https://raw.githubusercontent.com/api-evangelist/interface-research/main/schema/api-commons.yml#/$defs/RateLimits provider: Anyscale providerId: anyscale created: '2026-05-08' # Provenance stamped 2026-08-11: this artifact was written by the API Evangelist # bulk sweep dated 2026-05-08, not harvested from the provider. See roadmap#35. method: generated modified: '2026-05-08' reconciled: false tags: - AI - Distributed Computing - Ray - ML Platform - Inference - Rate Limiting - Quotas - Throttling description: >- Anyscale is a control-plane API for managing Ray compute. Throughput limits primarily come from the underlying cloud quotas (per-region instance and GPU quotas in the customer's AWS / GCP account or Anyscale's hosted account). Control-plane API call rates are not publicly documented and are pending reconciliation; service-level rate limits on Ray Serve services are controlled by user code and autoscaling configuration. notes: >- Pending reconciliation against Anyscale's published API rate limits. Practical limits are determined by cloud-provider GPU quotas and per-org instance caps. sources: - https://docs.anyscale.com/ - https://www.anyscale.com/pricing responseCodes: throttled: 429 limits: - name: Control-Plane API scope: organization metric: requests limit: see provider documentation notes: Pending reconciliation. - name: Concurrent Workspaces / Jobs / Services scope: organization metric: concurrent limit: bounded by cloud quotas and org limits notes: Practical concurrency is bounded by AWS / GCP instance and GPU quotas. - name: Cluster Node Counts scope: cluster metric: nodes limit: bounded by autoscaling and cloud quotas notes: Configured per compute config and bounded by cloud GPU quotas. - name: Service Endpoint scope: service metric: requests limit: user-configured notes: Throughput on deployed Ray Serve services is controlled by application autoscaling. policies: - name: Backoff Strategy description: Clients should implement exponential backoff with jitter and honor Retry-After. - name: Cloud Quota Management description: Request AWS / GCP quota increases ahead of large training or inference rollouts. maintainers: - FN: Kin Lane email: kin@apievangelist.com