openapi: 3.2.0 info: title: LiteLLM model management API description: 'Proxy Server to call 100+ LLMs in the OpenAI format. **Customize Swagger Docs** 👉 ```LiteLLM Admin Panel on /ui```. Create, Edit Keys with SSO. Having issues? Try ```Fallback Login``` 💸 ```LiteLLM Model Cost Map```. 🔎 ```LiteLLM Model Hub```. See available models on the proxy. **Docs**' version: 1.102.1 tags: - name: Model Management paths: /models: get: tags: - Model Management summary: Model List description: 'Use `/model/info` - to get detailed model information, example - pricing, mode, etc. This is just for compatibility with openai projects like aider. Query Parameters: - include_metadata: Include additional metadata in the response with fallback information - fallback_type: Type of fallbacks to include ("general", "context_window", "content_policy") Defaults to "general" when include_metadata=true - scope: Optional scope parameter. Currently only accepts "expand". When scope=expand is passed, proxy admins, team admins, and org admins will receive all proxy models as if they are a proxy admin. - healthy_only: When true, hide models whose backing deployments are all marked unhealthy by background health checks. Set `general_settings.model_list_healthy_only: true` to apply this to every caller without the query parameter. Requires `background_health_checks: true` in general_settings, plus either `model_list_healthy_only` or `enable_health_check_routing` to keep deployment health state cached; without health state the listing is returned unfiltered (fail open). Models expanded from wildcard routes (e.g. `openai/*`) are not filtered, and nothing is hidden when `allowed_fails_policy` is configured (cooldown remains the sole exclusion mechanism). Hiding is presentation-only: a hidden model can still be called directly.' operationId: model_list_models_get security: - APIKeyHeader: [] parameters: - name: return_wildcard_routes in: query required: false schema: anyOf: - type: boolean - type: 'null' default: false title: Return Wildcard Routes - name: team_id in: query required: false schema: anyOf: - type: string - type: 'null' title: Team Id - name: include_model_access_groups in: query required: false schema: anyOf: - type: boolean - type: 'null' default: false title: Include Model Access Groups - name: only_model_access_groups in: query required: false schema: anyOf: - type: boolean - type: 'null' default: false title: Only Model Access Groups - name: include_metadata in: query required: false schema: anyOf: - type: boolean - type: 'null' default: false title: Include Metadata - name: fallback_type in: query required: false schema: anyOf: - type: string - type: 'null' title: Fallback Type - name: scope in: query required: false schema: anyOf: - type: string - type: 'null' title: Scope - name: healthy_only in: query required: false schema: anyOf: - type: boolean - type: 'null' default: false title: Healthy Only responses: '200': description: Successful Response content: application/json: schema: {} '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /v1/models: get: tags: - Model Management summary: Model List description: 'Use `/model/info` - to get detailed model information, example - pricing, mode, etc. This is just for compatibility with openai projects like aider. Query Parameters: - include_metadata: Include additional metadata in the response with fallback information - fallback_type: Type of fallbacks to include ("general", "context_window", "content_policy") Defaults to "general" when include_metadata=true - scope: Optional scope parameter. Currently only accepts "expand". When scope=expand is passed, proxy admins, team admins, and org admins will receive all proxy models as if they are a proxy admin. - healthy_only: When true, hide models whose backing deployments are all marked unhealthy by background health checks. Set `general_settings.model_list_healthy_only: true` to apply this to every caller without the query parameter. Requires `background_health_checks: true` in general_settings, plus either `model_list_healthy_only` or `enable_health_check_routing` to keep deployment health state cached; without health state the listing is returned unfiltered (fail open). Models expanded from wildcard routes (e.g. `openai/*`) are not filtered, and nothing is hidden when `allowed_fails_policy` is configured (cooldown remains the sole exclusion mechanism). Hiding is presentation-only: a hidden model can still be called directly.' operationId: model_list_v1_models_get security: - APIKeyHeader: [] parameters: - name: return_wildcard_routes in: query required: false schema: anyOf: - type: boolean - type: 'null' default: false title: Return Wildcard Routes - name: team_id in: query required: false schema: anyOf: - type: string - type: 'null' title: Team Id - name: include_model_access_groups in: query required: false schema: anyOf: - type: boolean - type: 'null' default: false title: Include Model Access Groups - name: only_model_access_groups in: query required: false schema: anyOf: - type: boolean - type: 'null' default: false title: Only Model Access Groups - name: include_metadata in: query required: false schema: anyOf: - type: boolean - type: 'null' default: false title: Include Metadata - name: fallback_type in: query required: false schema: anyOf: - type: string - type: 'null' title: Fallback Type - name: scope in: query required: false schema: anyOf: - type: string - type: 'null' title: Scope - name: healthy_only in: query required: false schema: anyOf: - type: boolean - type: 'null' default: false title: Healthy Only responses: '200': description: Successful Response content: application/json: schema: {} '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /models/{model_id}: get: tags: - Model Management summary: Model Info description: 'Retrieve information about a specific model accessible to your API key. Returns model details only if the model is available to your API key/team. Returns 404 if the model doesn''t exist or is not accessible. Follows OpenAI API specification for individual model retrieval. https://platform.openai.com/docs/api-reference/models/retrieve Query parameters mirror `/v1/models` so the same caller context (team scoping, health filtering, paused deployments) drives both endpoints; the listing''s public id must resolve to the same internal deployment here.' operationId: model_info_models__model_id__get security: - APIKeyHeader: [] parameters: - name: model_id in: path required: true schema: type: string title: Model Id - name: team_id in: query required: false schema: anyOf: - type: string - type: 'null' title: Team Id - name: healthy_only in: query required: false schema: anyOf: - type: boolean - type: 'null' default: false title: Healthy Only responses: '200': description: Successful Response content: application/json: schema: {} '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /v1/models/{model_id}: get: tags: - Model Management summary: Model Info description: 'Retrieve information about a specific model accessible to your API key. Returns model details only if the model is available to your API key/team. Returns 404 if the model doesn''t exist or is not accessible. Follows OpenAI API specification for individual model retrieval. https://platform.openai.com/docs/api-reference/models/retrieve Query parameters mirror `/v1/models` so the same caller context (team scoping, health filtering, paused deployments) drives both endpoints; the listing''s public id must resolve to the same internal deployment here.' operationId: model_info_v1_models__model_id__get security: - APIKeyHeader: [] parameters: - name: model_id in: path required: true schema: type: string title: Model Id - name: team_id in: query required: false schema: anyOf: - type: string - type: 'null' title: Team Id - name: healthy_only in: query required: false schema: anyOf: - type: boolean - type: 'null' default: false title: Healthy Only responses: '200': description: Successful Response content: application/json: schema: {} '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /v2/model/info: get: tags: - Model Management summary: Model Info V2 description: 'Paginated model metadata for proxy deployments (pricing, provider, team access). Returns configured router deployments with enriched `model_info` (costs, provider, context window, etc.). Sensitive fields such as API keys and api_base are omitted. Query parameters: model: Filter to a single public `model_name`. user_models_only: When true, only return models created by the calling user. include_team_models: When true, populate `access_via_team_ids` and `direct_access` on each model and filter to deployments the caller can use. page / size: Pagination controls (defaults: page=1, size=50). search: Case-insensitive partial match on model name or team public name. modelId: Return a single deployment by LiteLLM model id. teamId: Filter to models with direct access or team membership for this team id. sortBy / sortOrder: Sort by model_name, created_at, updated_at, costs, or status. access_group: Only return deployments in this model access group. wildcard_only: Only return deployments whose `model_name` contains `*`. Example request: ``` curl -X GET ''http://localhost:4000/v2/model/info?include_team_models=true&page=1&size=50'' \ --header ''Authorization: Bearer sk-1234'' ``` Example response: ```json { "data": [ { "model_name": "gpt-4", "litellm_params": {"model": "openai/gpt-4.1"}, "model_info": { "id": "abc123", "litellm_provider": "openai", "access_via_team_ids": ["team-1"], "direct_access": true } } ], "total_count": 1, "current_page": 1, "total_pages": 1, "size": 50 } ```' operationId: model_info_v2_v2_model_info_get security: - APIKeyHeader: [] parameters: - name: model in: query required: false schema: anyOf: - type: string - type: 'null' description: Specify the model name (optional) title: Model description: Specify the model name (optional) - name: user_models_only in: query required: false schema: anyOf: - type: boolean - type: 'null' description: Only return models added by this user default: false title: User Models Only description: Only return models added by this user - name: include_team_models in: query required: false schema: anyOf: - type: boolean - type: 'null' description: Return all models across all teams user is in. default: false title: Include Team Models description: Return all models across all teams user is in. - name: debug in: query required: false schema: anyOf: - type: boolean - type: 'null' default: false title: Debug - name: page in: query required: false schema: type: integer minimum: 1 description: Page number default: 1 title: Page description: Page number - name: size in: query required: false schema: type: integer minimum: 1 description: Page size default: 50 title: Size description: Page size - name: search in: query required: false schema: anyOf: - type: string - type: 'null' description: Search model names (case-insensitive partial match) title: Search description: Search model names (case-insensitive partial match) - name: modelId in: query required: false schema: anyOf: - type: string - type: 'null' description: Search for a specific model by its unique ID title: Modelid description: Search for a specific model by its unique ID - name: teamId in: query required: false schema: anyOf: - type: string - type: 'null' description: Filter models by team ID. Returns models with direct_access=True or teamId in access_via_team_ids title: Teamid description: Filter models by team ID. Returns models with direct_access=True or teamId in access_via_team_ids - name: sortBy in: query required: false schema: anyOf: - type: string - type: 'null' description: 'Field to sort by. Options: model_name, created_at, updated_at, costs, status' title: Sortby description: 'Field to sort by. Options: model_name, created_at, updated_at, costs, status' - name: sortOrder in: query required: false schema: anyOf: - type: string - type: 'null' description: 'Sort order. Options: asc, desc' default: asc title: Sortorder description: 'Sort order. Options: asc, desc' - name: exclude_auto_routers in: query required: false schema: anyOf: - type: boolean - type: 'null' description: Omit auto-router deployments (litellm model prefixed `auto_router/`). They select among deployments rather than being deployments themselves, so a caller rendering a deployment list can leave them out. Defaults to false, so existing callers are unaffected default: false title: Exclude Auto Routers description: Omit auto-router deployments (litellm model prefixed `auto_router/`). They select among deployments rather than being deployments themselves, so a caller rendering a deployment list can leave them out. Defaults to false, so existing callers are unaffected - name: access_group in: query required: false schema: anyOf: - type: string - type: 'null' description: Only return deployments whose `model_info.access_groups` contains this access group title: Access Group description: Only return deployments whose `model_info.access_groups` contains this access group - name: wildcard_only in: query required: false schema: anyOf: - type: boolean - type: 'null' description: Only return wildcard deployments, i.e. those whose `model_name` contains `*` default: false title: Wildcard Only description: Only return wildcard deployments, i.e. those whose `model_name` contains `*` responses: '200': description: Successful Response content: application/json: schema: {} '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /v1/model/info: get: tags: - Model Management summary: Model Info V1 description: 'Provides more info about each model in /models, including config.yaml descriptions (except api key and api base) Parameters: litellm_model_id: Optional[str] = None (this is the value of `x-litellm-model-id` returned in response headers) - When litellm_model_id is passed, it will return the info for that specific model - When litellm_model_id is not passed, it will return the info for all models - include_team_models: When true, filter to deployments the caller can use (same as /v2/model/info). - teamId: Filter to models accessible by the given team. - healthy_only: When true, hide models whose backing deployments are all marked unhealthy by background health checks, matching `/v1/models?healthy_only=true`. Set `general_settings.model_list_healthy_only: true` to apply this to every caller without the query parameter. Requires `background_health_checks: true`, plus either `model_list_healthy_only` or `enable_health_check_routing` to keep deployment health state cached; without health state the listing is returned unfiltered (fail open). Ignored when `litellm_model_id` is passed, since that is a direct lookup of one deployment rather than a listing. Hiding is presentation-only: a hidden model can still be called directly. Each model in the list response includes `model_info.access_via_team_ids` and `model_info.direct_access` when the proxy database is connected. Returns: Returns a dictionary containing information about each model. Example Response: ```json { "data": [ { "model_name": "fake-openai-endpoint", "litellm_params": { "api_base": "https://exampleopenaiendpoint-production.up.railway.app/", "model": "openai/fake" }, "model_info": { "id": "112f74fab24a7a5245d2ced3536dd8f5f9192c57ee6e332af0f0512e08bed5af", "db_model": false } } ] } ```' operationId: model_info_v1_v1_model_info_get security: - APIKeyHeader: [] parameters: - name: litellm_model_id in: query required: false schema: anyOf: - type: string - type: 'null' title: Litellm Model Id - name: include_team_models in: query required: false schema: anyOf: - type: boolean - type: 'null' description: When true, filter to deployments the caller can use via direct access or team membership. default: false title: Include Team Models description: When true, filter to deployments the caller can use via direct access or team membership. - name: teamId in: query required: false schema: anyOf: - type: string - type: 'null' description: Filter models by team ID. Returns models with direct_access=True or teamId in access_via_team_ids title: Teamid description: Filter models by team ID. Returns models with direct_access=True or teamId in access_via_team_ids - name: healthy_only in: query required: false schema: anyOf: - type: boolean - type: 'null' default: false title: Healthy Only responses: '200': description: Successful Response content: application/json: schema: {} '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /model/info: get: tags: - Model Management summary: Model Info V1 description: 'Provides more info about each model in /models, including config.yaml descriptions (except api key and api base) Parameters: litellm_model_id: Optional[str] = None (this is the value of `x-litellm-model-id` returned in response headers) - When litellm_model_id is passed, it will return the info for that specific model - When litellm_model_id is not passed, it will return the info for all models - include_team_models: When true, filter to deployments the caller can use (same as /v2/model/info). - teamId: Filter to models accessible by the given team. - healthy_only: When true, hide models whose backing deployments are all marked unhealthy by background health checks, matching `/v1/models?healthy_only=true`. Set `general_settings.model_list_healthy_only: true` to apply this to every caller without the query parameter. Requires `background_health_checks: true`, plus either `model_list_healthy_only` or `enable_health_check_routing` to keep deployment health state cached; without health state the listing is returned unfiltered (fail open). Ignored when `litellm_model_id` is passed, since that is a direct lookup of one deployment rather than a listing. Hiding is presentation-only: a hidden model can still be called directly. Each model in the list response includes `model_info.access_via_team_ids` and `model_info.direct_access` when the proxy database is connected. Returns: Returns a dictionary containing information about each model. Example Response: ```json { "data": [ { "model_name": "fake-openai-endpoint", "litellm_params": { "api_base": "https://exampleopenaiendpoint-production.up.railway.app/", "model": "openai/fake" }, "model_info": { "id": "112f74fab24a7a5245d2ced3536dd8f5f9192c57ee6e332af0f0512e08bed5af", "db_model": false } } ] } ```' operationId: model_info_v1_model_info_get security: - APIKeyHeader: [] parameters: - name: litellm_model_id in: query required: false schema: anyOf: - type: string - type: 'null' title: Litellm Model Id - name: include_team_models in: query required: false schema: anyOf: - type: boolean - type: 'null' description: When true, filter to deployments the caller can use via direct access or team membership. default: false title: Include Team Models description: When true, filter to deployments the caller can use via direct access or team membership. - name: teamId in: query required: false schema: anyOf: - type: string - type: 'null' description: Filter models by team ID. Returns models with direct_access=True or teamId in access_via_team_ids title: Teamid description: Filter models by team ID. Returns models with direct_access=True or teamId in access_via_team_ids - name: healthy_only in: query required: false schema: anyOf: - type: boolean - type: 'null' default: false title: Healthy Only responses: '200': description: Successful Response content: application/json: schema: {} '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /v1/model/deprecations: get: tags: - Model Management summary: Model Deprecations description: 'List models with known deprecation/sunset dates, bucketed by urgency. Reads `deprecation_date` metadata from `model_prices_and_context_window.json` (and any per-deployment `model_info.deprecation_date` overrides) for the models configured on this proxy. Parameters: warn_within_days: Window (in days) used to bucket "imminent" models, 30 by default. Returns: A payload with three lists of `ModelDeprecationInfo` entries: - `deprecated`: deprecation date is in the past, so these requests may fail at any time. - `imminent`: deprecation date is within `warn_within_days` from today. - `upcoming`: deprecation date is further out. Example: ```shell curl -X GET ''http://localhost:4000/model/deprecations'' \ -H ''Authorization: Bearer sk-1234'' ```' operationId: model_deprecations_v1_model_deprecations_get security: - APIKeyHeader: [] parameters: - name: warn_within_days in: query required: false schema: type: integer default: 30 title: Warn Within Days responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/ModelDeprecationResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /model/deprecations: get: tags: - Model Management summary: Model Deprecations description: 'List models with known deprecation/sunset dates, bucketed by urgency. Reads `deprecation_date` metadata from `model_prices_and_context_window.json` (and any per-deployment `model_info.deprecation_date` overrides) for the models configured on this proxy. Parameters: warn_within_days: Window (in days) used to bucket "imminent" models, 30 by default. Returns: A payload with three lists of `ModelDeprecationInfo` entries: - `deprecated`: deprecation date is in the past, so these requests may fail at any time. - `imminent`: deprecation date is within `warn_within_days` from today. - `upcoming`: deprecation date is further out. Example: ```shell curl -X GET ''http://localhost:4000/model/deprecations'' \ -H ''Authorization: Bearer sk-1234'' ```' operationId: model_deprecations_model_deprecations_get security: - APIKeyHeader: [] parameters: - name: warn_within_days in: query required: false schema: type: integer default: 30 title: Warn Within Days responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/ModelDeprecationResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /model_group/info: get: tags: - Model Management summary: Model Group Info description: 'Get information about all the deployments on litellm proxy, including config.yaml descriptions (except api key and api base) - /model_group/info returns all model groups. End users of proxy should use /model_group/info since those models will be used for /chat/completions, /embeddings, etc. - /model_group/info?model_group=rerank-english-v3.0 returns all model groups for a specific model group (`model_name` in config.yaml) Example Request (All Models): ```shell curl -X ''GET'' ''http://localhost:4000/model_group/info'' -H ''accept: application/json'' -H ''x-api-key: sk-1234'' ``` Example Request (Specific Model Group): ```shell curl -X ''GET'' ''http://localhost:4000/model_group/info?model_group=rerank-english-v3.0'' -H ''accept: application/json'' -H ''Authorization: Bearer sk-1234'' ``` Example Request (Specific Wildcard Model Group): (e.g. `model_name: openai/*` on config.yaml) ```shell curl -X ''GET'' ''http://localhost:4000/model_group/info?model_group=openai/tts-1'' -H ''accept: application/json'' -H ''Authorization: Bearersk-1234'' ``` Learn how to use and set wildcard models here Example Response: ```json { "data": [ { "model_group": "rerank-english-v3.0", "providers": [ "cohere" ], "max_input_tokens": null, "max_output_tokens": null, "input_cost_per_token": 0.0, "output_cost_per_token": 0.0, "mode": null, "tpm": null, "rpm": null, "supports_parallel_function_calling": false, "supports_vision": false, "supports_function_calling": false, "supported_openai_params": [ "stream", "temperature", "max_tokens", "logit_bias", "top_p", "frequency_penalty", "presence_penalty", "stop", "n", "extra_headers" ] }, { "model_group": "gpt-3.5-turbo", "providers": [ "openai" ], "max_input_tokens": 16385.0, "max_output_tokens": 4096.0, "input_cost_per_token": 1.5e-06, "output_cost_per_token": 2e-06, "mode": "chat", "tpm": null, "rpm": null, "supports_parallel_function_calling": false, "supports_vision": false, "supports_function_calling": true, "supported_openai_params": [ "frequency_penalty", "logit_bias", "logprobs", "top_logprobs", "max_tokens", "max_completion_tokens", "n", "presence_penalty", "seed", "stop", "stream", "stream_options", "temperature", "top_p", "tools", "tool_choice", "function_call", "functions", "max_retries", "extra_headers", "parallel_tool_calls", "response_format" ] }, { "model_group": "llava-hf", "providers": [ "openai" ], "max_input_tokens": null, "max_output_tokens": null, "input_cost_per_token": 0.0, "output_cost_per_token": 0.0, "mode": null, "tpm": null, "rpm": null, "supports_parallel_function_calling": false, "supports_vision": true, "supports_function_calling": false, "supported_openai_params": [ "frequency_penalty", "logit_bias", "logprobs", "top_logprobs", "max_tokens", "max_completion_tokens", "n", "presence_penalty", "seed", "stop", "stream", "stream_options", "temperature", "top_p", "tools", "tool_choice", "function_call", "functions", "max_retries", "extra_headers", "parallel_tool_calls", "response_format" ] } ] } ```' operationId: model_group_info_model_group_info_get security: - APIKeyHeader: [] parameters: - name: model_group in: query required: false schema: anyOf: - type: string - type: 'null' title: Model Group responses: '200': description: Successful Response content: application/json: schema: {} '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /public/model_hub: get: tags: - Model Management summary: Public Model Hub operationId: public_model_hub_public_model_hub_get responses: '200': description: Successful Response content: application/json: schema: items: $ref: '#/components/schemas/ModelGroupInfoProxy' type: array title: Response Public Model Hub Public Model Hub Get /public/model_hub/info: get: tags: - Model Management summary: Public Model Hub Info operationId: public_model_hub_info_public_model_hub_info_get responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/PublicModelHubInfo' /public/litellm_model_cost_map: get: tags: - Model Management summary: Get Litellm Model Cost Map description: 'Public endpoint to get the LiteLLM model cost map. Returns pricing information for all supported models.' operationId: get_litellm_model_cost_map_public_litellm_model_cost_map_get responses: '200': description: Successful Response content: application/json: schema: {} /public/v1/model_hub: get: tags: - Model Management summary: Public Model Hub List description: 'The public model groups this proxy publishes, paged, sortable, searchable and filterable, for the public Model Hub page. No authentication. A rejected request answers with the parameters, sort fields and filter operators it would have accepted, so the accepted set stays discoverable from the endpoint itself rather than from a copy of the spec kept here. Example curl: ``` curl --location --globoff ''http://0.0.0.0:4000/public/v1/model_hub?sort=-input_cost_per_token&filter[mode][in]=chat&page_size=25'' ```' operationId: public_model_hub_list_public_v1_model_hub_get responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/ListResponse_ModelGroupInfoProxy_' security: - APIKeyHeader: [] /public/v1/model_hub/{facet}: get: tags: - Model Management summary: Public Model Hub Facet description: 'The distinct providers, modes or features across the published model groups, for the Model Hub''s filter dropdowns. No authentication. Carries the same filters and search as the list route, so a dropdown offers exactly the values the table can show: asking for providers under `filter[mode][in]=chat` lists only the providers that serve a chat model. Example curl: ``` curl --location --globoff ''http://0.0.0.0:4000/public/v1/model_hub/providers?filter[mode][in]=chat&page_size=50'' ```' operationId: public_model_hub_facet_public_v1_model_hub__facet__get security: - APIKeyHeader: [] parameters: - name: facet in: path required: true schema: enum: - providers - modes - features type: string title: Facet responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/FacetListResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /model/{model_id}/update: patch: tags: - Model Management summary: Patch Model description: 'PATCH Endpoint for partial model updates. Only updates the fields specified in the request while preserving other existing values. Follows proper PATCH semantics by only modifying provided fields. Args: model_id: The ID of the model to update patch_data: The fields to update and their new values user_api_key_dict: User authentication information Returns: Updated model information Raises: ProxyException: For various error conditions including authentication and database errors' operationId: patch_model_model__model_id__update_patch security: - APIKeyHeader: [] parameters: - name: model_id in: path required: true schema: type: string title: Model Id requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/updateDeployment' responses: '200': description: Successful Response content: application/json: schema: {} '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /model/block: post: tags: - Model Management summary: Block Model description: 'Block a DB-stored model deployment from serving requests. Parameters: - model_id: str - The model deployment id to block.' operationId: block_model_model_block_post security: - APIKeyHeader: [] parameters: - name: litellm-changed-by in: header required: false schema: anyOf: - type: string - type: 'null' description: The litellm-changed-by header enables tracking of actions performed by authorized users on behalf of other users, providing an audit trail for accountability title: Litellm-Changed-By description: The litellm-changed-by header enables tracking of actions performed by authorized users on behalf of other users, providing an audit trail for accountability requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/BlockModelRequest' responses: '200': description: Successful Response content: application/json: schema: anyOf: - $ref: '#/components/schemas/LiteLLM_ProxyModelTable' - type: 'null' title: Response Block Model Model Block Post '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /model/unblock: post: tags: - Model Management summary: Unblock Model description: 'Unblock a DB-stored model deployment so it can serve requests again. Parameters: - model_id: str - The model deployment id to unblock.' operationId: unblock_model_model_unblock_post security: - APIKeyHeader: [] parameters: - name: litellm-changed-by in: header required: false schema: anyOf: - type: string - type: 'null' description: The litellm-changed-by header enables tracking of actions performed by authorized users on behalf of other users, providing an audit trail for accountability title: Litellm-Changed-By description: The litellm-changed-by header enables tracking of actions performed by authorized users on behalf of other users, providing an audit trail for accountability requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/BlockModelRequest' responses: '200': description: Successful Response content: application/json: schema: anyOf: - $ref: '#/components/schemas/LiteLLM_ProxyModelTable' - type: 'null' title: Response Unblock Model Model Unblock Post '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /model/delete: post: tags: - Model Management summary: Delete Model description: Allows deleting models in the model list in the config.yaml operationId: delete_model_model_delete_post requestBody: content: application/json: schema: $ref: '#/components/schemas/ModelInfoDelete' required: true responses: '200': description: Successful Response content: application/json: schema: {} '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' security: - APIKeyHeader: [] /model/new: post: tags: - Model Management summary: Add New Model description: Allows adding new models to the model list in the config.yaml operationId: add_new_model_model_new_post requestBody: content: application/json: schema: $ref: '#/components/schemas/Deployment' required: true responses: '200': description: Successful Response content: application/json: schema: {} '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' security: - APIKeyHeader: [] /model/update: post: tags: - Model Management summary: Update Model description: Edit existing model params operationId: update_model_model_update_post requestBody: content: application/json: schema: $ref: '#/components/schemas/updateDeployment' required: true responses: '200': description: Successful Response content: application/json: schema: {} '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' security: - APIKeyHeader: [] /model_group/make_public: post: tags: - Model Management summary: Update Public Model Groups description: Update which model groups are public operationId: update_public_model_groups_model_group_make_public_post requestBody: content: application/json: schema: $ref: '#/components/schemas/UpdatePublicModelGroupsRequest' required: true responses: '200': description: Successful Response content: application/json: schema: {} '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' security: - APIKeyHeader: [] /model_hub/update_useful_links: post: tags: - Model Management summary: Update Useful Links description: Update useful links operationId: update_useful_links_model_hub_update_useful_links_post requestBody: content: application/json: schema: $ref: '#/components/schemas/UpdateUsefulLinksRequest' required: true responses: '200': description: Successful Response content: application/json: schema: {} '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' security: - APIKeyHeader: [] /auto_router/classifier/default_prompt: post: tags: - Model Management summary: Preview Auto Router Classifier Prompt description: Get the system prompt an auto-router's LLM classifier sends for an edited tier set operationId: preview_auto_router_classifier_prompt_auto_router_classifier_default_prompt_post security: - APIKeyHeader: [] requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/AutoRouterClassifierPromptPreviewRequest' responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/AutoRouterClassifierDefaultPromptResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' get: tags: - Model Management summary: Get Auto Router Classifier Default Prompt description: Get the built-in system prompt used by an auto-router's LLM classifier operationId: get_auto_router_classifier_default_prompt_auto_router_classifier_default_prompt_get security: - APIKeyHeader: [] parameters: - name: context_window_size in: query required: false schema: type: integer default: 3 title: Context Window Size - name: tier_labels in: query required: false schema: anyOf: - type: string - type: 'null' title: Tier Labels - name: classification_rubric in: query required: false schema: anyOf: - $ref: '#/components/schemas/ClassificationRubric' - type: 'null' title: Classification Rubric responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/AutoRouterClassifierDefaultPromptResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /access_group/new: post: tags: - Model Management summary: Create Model Group description: 'Create a new access group containing multiple model names. An access group is a named collection of model groups that can be referenced by teams/keys for simplified access control. Example: ```bash curl -X POST ''http://localhost:4000/access_group/new'' \ -H ''Authorization: Bearer sk-1234'' \ -H ''Content-Type: application/json'' \ -d ''{ "access_group": "production-models", "model_names": ["gpt-4", "claude-3-opus", "gemini-pro"] }'' ``` Parameters: - access_group: str - The access group name (e.g., "production-models") - model_names: List[str] - List of existing model groups to include Returns: - NewModelGroupResponse with the created access group details Raises: - HTTPException 400: If any model names don''t exist - HTTPException 500: If database operations fail' operationId: create_model_group_access_group_new_post requestBody: content: application/json: schema: $ref: '#/components/schemas/NewModelGroupRequest' required: true responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/NewModelGroupResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' security: - APIKeyHeader: [] /access_group/list: get: tags: - Model Management summary: List Access Groups description: 'List all access groups. Returns a list of all access groups with their model names, deployment counts, shared budget and the spend drawn against it. Example: ```bash curl -X GET ''http://localhost:4000/access_group/list'' \ -H ''Authorization: Bearer sk-1234'' ``` Returns: - ListAccessGroupsResponse with all access groups' operationId: list_access_groups_access_group_list_get responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/ListAccessGroupsResponse' security: - APIKeyHeader: [] /access_group/{access_group}/info: get: tags: - Model Management summary: Get Access Group Info description: 'Get information about a specific access group. Example: ```bash curl -X GET ''http://localhost:4000/access_group/production-models/info'' \ -H ''Authorization: Bearer sk-1234'' ``` Parameters: - access_group: str - The access group name (URL path parameter) Returns: - AccessGroupInfo with the access group details, its shared budget and its spend Raises: - HTTPException 404: If access group not found' operationId: get_access_group_info_access_group__access_group__info_get security: - APIKeyHeader: [] parameters: - name: access_group in: path required: true schema: type: string title: Access Group responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/AccessGroupInfo' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /access_group/{access_group}/update: put: tags: - Model Management summary: Update Access Group description: 'Update an access group''s model names. This will: 1. Remove the access group from all current deployments 2. Add the access group to all deployments for the new model_names list Example: ```bash curl -X PUT ''http://localhost:4000/access_group/production-models/update'' \ -H ''Authorization: Bearer sk-1234'' \ -H ''Content-Type: application/json'' \ -d ''{ "model_names": ["gpt-4", "claude-3-sonnet"] }'' ``` Parameters: - access_group: str - The access group name (URL path parameter) - model_names: List[str] - New list of model groups to include Returns: - NewModelGroupResponse with the updated access group details Raises: - HTTPException 400: If any model names don''t exist - HTTPException 404: If access group not found' operationId: update_access_group_access_group__access_group__update_put security: - APIKeyHeader: [] parameters: - name: access_group in: path required: true schema: type: string title: Access Group requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/UpdateModelGroupRequest' responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/NewModelGroupResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /access_group/{access_group}/delete: delete: tags: - Model Management summary: Delete Access Group description: 'Delete an access group. Removes the access group from all deployments that have it. Example: ```bash curl -X DELETE ''http://localhost:4000/access_group/production-models/delete'' \ -H ''Authorization: Bearer sk-1234'' ``` Parameters: - access_group: str - The access group name (URL path parameter) Returns: - DeleteModelGroupResponse with deletion details Raises: - HTTPException 404: If access group not found' operationId: delete_access_group_access_group__access_group__delete_delete security: - APIKeyHeader: [] parameters: - name: access_group in: path required: true schema: type: string title: Access Group responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/DeleteModelGroupResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /access_group/{access_group}/budget: get: tags: - Model Management summary: Get Access Group Budget description: 'Get the shared budget of an access group, and the spend drawn against it. Example: ```bash curl -X GET ''http://localhost:4000/access_group/production-models/budget'' \ -H ''Authorization: Bearer sk-1234'' ``` Parameters: - access_group: str - The access group name (URL path parameter) Returns: - AccessGroupBudgetResponse; budget is null when the group has no budget set Raises: - HTTPException 404: If access group not found' operationId: get_access_group_budget_access_group__access_group__budget_get security: - APIKeyHeader: [] parameters: - name: access_group in: path required: true schema: type: string title: Access Group responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/AccessGroupBudgetResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' put: tags: - Model Management summary: Set Access Group Budget description: 'Set or replace the shared budget of an access group. Idempotent. Every key that can reach a model in the group draws from this one budget. Example: ```bash curl -X PUT ''http://localhost:4000/access_group/production-models/budget'' \ -H ''Authorization: Bearer sk-1234'' \ -H ''Content-Type: application/json'' \ -d ''{ "max_budget": 100.0, "budget_duration": "30d" }'' ``` Parameters: - access_group: str - The access group name (URL path parameter) - max_budget: Optional[float] - Requests fail once the group''s shared spend exceeds this - soft_budget: Optional[float] - Fires an alert when reached; requests still succeed - budget_duration: Optional[str] - Frequency of resetting the group''s spend (e.g. ''30d'') - budget_id: Optional[str] - Link an existing budget instead of creating one Returns: - AccessGroupBudgetResponse with the stored budget and current spend Raises: - HTTPException 400: If no budget field is given, or budget_duration cannot be parsed - HTTPException 404: If access group not found' operationId: set_access_group_budget_access_group__access_group__budget_put security: - APIKeyHeader: [] parameters: - name: access_group in: path required: true schema: type: string title: Access Group requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/AccessGroupBudgetRequest' responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/AccessGroupBudgetResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' delete: tags: - Model Management summary: Delete Access Group Budget description: 'Clear the shared budget of an access group, leaving the group itself in place. Example: ```bash curl -X DELETE ''http://localhost:4000/access_group/production-models/budget'' \ -H ''Authorization: Bearer sk-1234'' ``` Parameters: - access_group: str - The access group name (URL path parameter) Returns: - DeleteAccessGroupBudgetResponse; budget_deleted is false when there was nothing to clear Raises: - HTTPException 404: If access group not found' operationId: delete_access_group_budget_access_group__access_group__budget_delete security: - APIKeyHeader: [] parameters: - name: access_group in: path required: true schema: type: string title: Access Group responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/DeleteAccessGroupBudgetResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /auto_router/validate_complexity_router_config: post: tags: - Model Management summary: Validate Complexity Router Config description: 'Validate a complexity-router config without saving it. Runs the same check every write path runs (the router''s own pydantic model), so a form can show the backend''s exact verdict while the operator is still editing rather than after a rejected save. Uses the same team opt-in and model-access checks as configuration writes for members. Nothing is created, routed, or billed.' operationId: validate_complexity_router_config_auto_router_validate_complexity_router_config_post requestBody: content: application/json: schema: $ref: '#/components/schemas/ComplexityRouterConfigValidationRequest' required: true responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/ComplexityRouterConfigValidationResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' security: - APIKeyHeader: [] /auto_router/test_routing: post: tags: - Model Management summary: Preview Auto Router Routing description: 'Route a single request through a complexity-router config and report where it landed. Answers "which model would this request get?" for a config that only exists in a form, so an auto router can be checked before it is created. The request is classified by the same pre-routing hook a live request runs, over the same messages, system prompt and tool definitions, then dropped: nothing is sent to the model it routed to, and no auto router is created. A heuristic config therefore spends nothing, while an `llm` classifier or semantic keyword matching bills its classifier/embedding call to the calling key, like Test Connection does. Send `messages` to classify a real turn, with `system` and `tools` beside it when the surface carries them top level, as Anthropic /v1/messages does. `prompt` is the single-ask shorthand and routes as one user turn with nothing around it. **Example Request:** ```json { "messages": [ {"role": "system", "content": "You are a database migration assistant"}, {"role": "user", "content": "the index is not unique"}, {"role": "assistant", "content": "Then two workers can both insert. Add a unique index"}, {"role": "user", "content": "ok do it"} ], "tools": [{"type": "function", "function": {"name": "Bash", "description": "Run a command"}}], "complexity_router_config": { "tiers": {"SIMPLE": ["gpt-4o-mini"], "REASONING": ["o3"]}, "classifier_type": "heuristic" } } ```' operationId: preview_auto_router_routing_auto_router_test_routing_post requestBody: content: application/json: schema: $ref: '#/components/schemas/AutoRouterRoutingTestRequest' required: true responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/AutoRouterRoutingTestResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' security: - APIKeyHeader: [] components: schemas: ChatCompletionReasoningSummaryTextBlock: properties: type: const: summary_text title: Type type: string text: title: Text type: string required: - type title: ChatCompletionReasoningSummaryTextBlock type: object ListLinks: properties: self: type: string title: Self first: type: string title: First prev: anyOf: - type: string - type: 'null' title: Prev next: anyOf: - type: string - type: 'null' title: Next last: type: string title: Last type: object required: - self - first - last title: ListLinks description: Page-mode counterpart to `PageLinks`. `first`/`last` are knowable here because the total count is. ValidationError: properties: loc: items: anyOf: - type: string - type: integer type: array title: Location msg: type: string title: Message type: type: string title: Error Type input: title: Input ctx: type: object title: Context type: object required: - loc - msg - type title: ValidationError HTTPValidationError: properties: detail: items: $ref: '#/components/schemas/ValidationError' type: array title: Detail type: object title: HTTPValidationError ComplexityTierModel: properties: model_name: type: string title: Model Name litellm_params: additionalProperties: true type: object title: Litellm Params type: object required: - model_name title: ComplexityTierModel ChoiceLogprobs: properties: content: anyOf: - items: $ref: '#/components/schemas/ChatCompletionTokenLogprob' type: array - type: 'null' title: Content additionalProperties: true type: object title: ChoiceLogprobs ChatCompletionTokenLogprob: properties: token: type: string title: Token bytes: anyOf: - items: type: integer type: array - type: 'null' title: Bytes logprob: type: number title: Logprob top_logprobs: items: $ref: '#/components/schemas/TopLogprob' type: array title: Top Logprobs additionalProperties: true type: object required: - token - logprob - top_logprobs title: ChatCompletionTokenLogprob AccessGroupBudget: properties: budget_id: type: string title: Budget Id max_budget: anyOf: - type: number - type: 'null' title: Max Budget soft_budget: anyOf: - type: number - type: 'null' title: Soft Budget budget_duration: anyOf: - type: string - type: 'null' title: Budget Duration budget_reset_at: anyOf: - type: string format: date-time - type: 'null' title: Budget Reset At type: object required: - budget_id title: AccessGroupBudget AutoRouterClassifierDefaultPromptResponse: properties: system_prompt: type: string title: System Prompt type: object required: - system_prompt title: AutoRouterClassifierDefaultPromptResponse description: 'The built-in system prompt an auto-router''s LLM classifier uses when none is configured. Served so the dashboard''s prompt editor prefills the rubric the proxy actually sends, rather than a copy in the frontend that drifts the moment the rubric is edited.' KeywordTierRule: properties: keywords: items: type: string type: array minItems: 1 title: Keywords description: Keywords/phrases that trigger this rule (lexical or semantic match) tier: type: string title: Tier description: 'Tier to route to when this rule matches: a built-in tier name, or with tier_definitions set, one of the defined tier names' type: object required: - keywords - tier title: KeywordTierRule description: 'A deterministic override: if any keyword matches, route to this tier.' ChatCompletionMessageCustomToolCall: properties: id: type: string title: Id type: type: string const: custom title: Type default: custom custom: $ref: '#/components/schemas/ChatCompletionCustomToolCallPayload' additionalProperties: true type: object required: - id - custom title: ChatCompletionMessageCustomToolCall PageMeta: properties: page: type: integer title: Page page_size: type: integer title: Page Size has_more: type: boolean title: Has More type: object required: - page - page_size - has_more title: PageMeta description: '`has_more` rather than `total_count`, which would need a COUNT(*) over the whole match set per keystroke.' ModelInfoDelete: properties: id: type: string title: Id type: object required: - id title: ModelInfoDelete ChatCompletionReasoningItem: description: Represents an OpenAI Responses API reasoning item for round-tripping in conversation history. properties: type: const: reasoning title: Type type: string id: title: Id type: string encrypted_content: anyOf: - type: string - type: 'null' title: Encrypted Content summary: items: $ref: '#/components/schemas/ChatCompletionReasoningSummaryTextBlock' title: Summary type: array required: - type title: ChatCompletionReasoningItem type: object JevClassifierConfig: properties: model: type: string title: Model default: jev-latest api_key: anyOf: - type: string - type: 'null' title: Api Key description: TypeSafe API key, falling back to TYPESAFE_API_KEY api_base: anyOf: - type: string - type: 'null' title: Api Base description: TypeSafe API base, falling back to TYPESAFE_API_BASE and then https://api.typesafe.ai timeout_ms: type: integer minimum: 1.0 title: Timeout Ms default: 3000 instructions: anyOf: - type: string - type: 'null' title: Instructions description: Replaces the built-in Jev question instructions circuit_breaker_enabled: type: boolean title: Circuit Breaker Enabled default: true circuit_breaker_cooldown_seconds: type: number exclusiveMinimum: 0.0 title: Circuit Breaker Cooldown Seconds default: 30.0 additionalProperties: false type: object title: JevClassifierConfig updateLiteLLMParams: properties: input_cost_per_token: anyOf: - type: number - type: 'null' title: Input Cost Per Token output_cost_per_token: anyOf: - type: number - type: 'null' title: Output Cost Per Token input_cost_per_character: anyOf: - type: number - type: 'null' title: Input Cost Per Character output_cost_per_character: anyOf: - type: number - type: 'null' title: Output Cost Per Character cache_read_input_token_cost: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost cache_creation_input_token_cost: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost tiered_pricing: anyOf: - items: additionalProperties: true type: object type: array - type: 'null' title: Tiered Pricing input_cost_per_second: anyOf: - type: number - type: 'null' title: Input Cost Per Second output_cost_per_second: anyOf: - type: number - type: 'null' title: Output Cost Per Second output_cost_per_second_1080p: anyOf: - type: number - type: 'null' title: Output Cost Per Second 1080P output_cost_per_second_480p: anyOf: - type: number - type: 'null' title: Output Cost Per Second 480P output_cost_per_second_720p: anyOf: - type: number - type: 'null' title: Output Cost Per Second 720P output_cost_per_second_4k: anyOf: - type: number - type: 'null' title: Output Cost Per Second 4K input_cost_per_pixel: anyOf: - type: number - type: 'null' title: Input Cost Per Pixel output_cost_per_pixel: anyOf: - type: number - type: 'null' title: Output Cost Per Pixel input_cost_per_token_flex: anyOf: - type: number - type: 'null' title: Input Cost Per Token Flex input_cost_per_token_priority: anyOf: - type: number - type: 'null' title: Input Cost Per Token Priority input_cost_per_token_ultrafast: anyOf: - type: number - type: 'null' title: Input Cost Per Token Ultrafast cache_creation_input_token_cost_above_1hr: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Above 1Hr cache_creation_input_token_cost_above_200k_tokens: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Above 200K Tokens cache_creation_input_token_cost_above_272k_tokens: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Above 272K Tokens cache_creation_input_token_cost_above_272k_tokens_priority: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Above 272K Tokens Priority cache_creation_input_token_cost_above_272k_tokens_flex: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Above 272K Tokens Flex cache_creation_input_token_cost_flex: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Flex cache_creation_input_token_cost_priority: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Priority cache_creation_input_token_cost_ultrafast: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Ultrafast cache_creation_input_audio_token_cost: anyOf: - type: number - type: 'null' title: Cache Creation Input Audio Token Cost cache_read_input_token_cost_flex: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Flex cache_read_input_token_cost_priority: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Priority cache_read_input_token_cost_ultrafast: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Ultrafast cache_read_input_token_cost_above_200k_tokens: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Above 200K Tokens cache_read_input_token_cost_above_200k_tokens_priority: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Above 200K Tokens Priority cache_read_input_token_cost_above_272k_tokens_priority: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Above 272K Tokens Priority cache_read_input_token_cost_above_272k_tokens_flex: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Above 272K Tokens Flex cache_read_input_audio_token_cost: anyOf: - type: number - type: 'null' title: Cache Read Input Audio Token Cost input_cost_per_character_above_128k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Character Above 128K Tokens input_cost_per_audio_token: anyOf: - type: number - type: 'null' title: Input Cost Per Audio Token input_cost_per_token_cache_hit: anyOf: - type: number - type: 'null' title: Input Cost Per Token Cache Hit input_cost_per_token_above_128k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Token Above 128K Tokens input_cost_per_token_above_200k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Token Above 200K Tokens input_cost_per_token_above_200k_tokens_priority: anyOf: - type: number - type: 'null' title: Input Cost Per Token Above 200K Tokens Priority input_cost_per_token_above_272k_tokens_priority: anyOf: - type: number - type: 'null' title: Input Cost Per Token Above 272K Tokens Priority input_cost_per_token_above_272k_tokens_flex: anyOf: - type: number - type: 'null' title: Input Cost Per Token Above 272K Tokens Flex input_cost_per_query: anyOf: - type: number - type: 'null' title: Input Cost Per Query input_cost_per_image: anyOf: - type: number - type: 'null' title: Input Cost Per Image input_cost_per_image_above_128k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Image Above 128K Tokens input_cost_per_audio_per_second: anyOf: - type: number - type: 'null' title: Input Cost Per Audio Per Second input_cost_per_audio_per_second_above_128k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Audio Per Second Above 128K Tokens input_cost_per_video_per_second: anyOf: - type: number - type: 'null' title: Input Cost Per Video Per Second input_cost_per_video_per_second_above_128k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Video Per Second Above 128K Tokens input_cost_per_video_per_second_above_15s_interval: anyOf: - type: number - type: 'null' title: Input Cost Per Video Per Second Above 15S Interval input_cost_per_video_per_second_above_8s_interval: anyOf: - type: number - type: 'null' title: Input Cost Per Video Per Second Above 8S Interval input_cost_per_token_batches: anyOf: - type: number - type: 'null' title: Input Cost Per Token Batches output_cost_per_token_batches: anyOf: - type: number - type: 'null' title: Output Cost Per Token Batches output_cost_per_token_flex: anyOf: - type: number - type: 'null' title: Output Cost Per Token Flex output_cost_per_token_priority: anyOf: - type: number - type: 'null' title: Output Cost Per Token Priority output_cost_per_token_ultrafast: anyOf: - type: number - type: 'null' title: Output Cost Per Token Ultrafast output_cost_per_audio_token: anyOf: - type: number - type: 'null' title: Output Cost Per Audio Token output_cost_per_token_above_128k_tokens: anyOf: - type: number - type: 'null' title: Output Cost Per Token Above 128K Tokens output_cost_per_token_above_200k_tokens: anyOf: - type: number - type: 'null' title: Output Cost Per Token Above 200K Tokens output_cost_per_token_above_200k_tokens_priority: anyOf: - type: number - type: 'null' title: Output Cost Per Token Above 200K Tokens Priority output_cost_per_token_above_272k_tokens_priority: anyOf: - type: number - type: 'null' title: Output Cost Per Token Above 272K Tokens Priority output_cost_per_token_above_272k_tokens_flex: anyOf: - type: number - type: 'null' title: Output Cost Per Token Above 272K Tokens Flex output_cost_per_character_above_128k_tokens: anyOf: - type: number - type: 'null' title: Output Cost Per Character Above 128K Tokens output_cost_per_image: anyOf: - type: number - type: 'null' title: Output Cost Per Image output_cost_per_image_token: anyOf: - type: number - type: 'null' title: Output Cost Per Image Token output_cost_per_video_token: anyOf: - type: number - type: 'null' title: Output Cost Per Video Token output_cost_per_reasoning_token: anyOf: - type: number - type: 'null' title: Output Cost Per Reasoning Token output_cost_per_reasoning_token_flex: anyOf: - type: number - type: 'null' title: Output Cost Per Reasoning Token Flex output_cost_per_reasoning_token_priority: anyOf: - type: number - type: 'null' title: Output Cost Per Reasoning Token Priority output_cost_per_video_per_second: anyOf: - type: number - type: 'null' title: Output Cost Per Video Per Second output_cost_per_audio_per_second: anyOf: - type: number - type: 'null' title: Output Cost Per Audio Per Second search_context_cost_per_query: anyOf: - additionalProperties: true type: object - type: 'null' title: Search Context Cost Per Query google_maps_grounding_cost_per_query: anyOf: - type: number - type: 'null' title: Google Maps Grounding Cost Per Query citation_cost_per_token: anyOf: - type: number - type: 'null' title: Citation Cost Per Token cache_read_input_token_cost_above_272k_tokens: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Above 272K Tokens cache_read_input_token_cost_above_512k_tokens: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Above 512K Tokens input_cost_per_image_token: anyOf: - type: number - type: 'null' title: Input Cost Per Image Token input_cost_per_video_token: anyOf: - type: number - type: 'null' title: Input Cost Per Video Token input_cost_per_token_above_272k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Token Above 272K Tokens input_cost_per_token_above_512k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Token Above 512K Tokens output_cost_per_token_above_272k_tokens: anyOf: - type: number - type: 'null' title: Output Cost Per Token Above 272K Tokens output_cost_per_token_above_512k_tokens: anyOf: - type: number - type: 'null' title: Output Cost Per Token Above 512K Tokens output_vector_size: anyOf: - type: integer - type: 'null' title: Output Vector Size ocr_cost_per_page: anyOf: - type: number - type: 'null' title: Ocr Cost Per Page ocr_cost_per_credit: anyOf: - type: number - type: 'null' title: Ocr Cost Per Credit annotation_cost_per_page: anyOf: - type: number - type: 'null' title: Annotation Cost Per Page regional_processing_uplift_multiplier_eu: anyOf: - type: number - type: 'null' title: Regional Processing Uplift Multiplier Eu regional_processing_uplift_multiplier_us: anyOf: - type: number - type: 'null' title: Regional Processing Uplift Multiplier Us regional_endpoint_uplift_multiplier: anyOf: - type: number - type: 'null' title: Regional Endpoint Uplift Multiplier api_key: anyOf: - type: string - type: 'null' title: Api Key api_base: anyOf: - type: string - type: 'null' title: Api Base api_version: anyOf: - type: string - type: 'null' title: Api Version azure_ad_token: anyOf: - type: string - type: 'null' title: Azure Ad Token vertex_project: anyOf: - type: string - type: 'null' title: Vertex Project vertex_location: anyOf: - type: string - type: 'null' title: Vertex Location vertex_credentials: anyOf: - type: string - additionalProperties: true type: object - type: 'null' title: Vertex Credentials region_name: anyOf: - type: string - type: 'null' title: Region Name gcs_bucket_name: anyOf: - type: string - type: 'null' title: Gcs Bucket Name aws_access_key_id: anyOf: - type: string - type: 'null' title: Aws Access Key Id aws_secret_access_key: anyOf: - type: string - type: 'null' title: Aws Secret Access Key aws_session_token: anyOf: - type: string - type: 'null' title: Aws Session Token aws_region_name: anyOf: - type: string - type: 'null' title: Aws Region Name aws_session_name: anyOf: - type: string - type: 'null' title: Aws Session Name aws_profile_name: anyOf: - type: string - type: 'null' title: Aws Profile Name aws_role_name: anyOf: - type: string - type: 'null' title: Aws Role Name aws_web_identity_token: anyOf: - type: string - type: 'null' title: Aws Web Identity Token aws_sts_endpoint: anyOf: - type: string - type: 'null' title: Aws Sts Endpoint aws_external_id: anyOf: - type: string - type: 'null' title: Aws External Id aws_session_tags: anyOf: - items: $ref: '#/components/schemas/AwsSessionTag' type: array - type: 'null' title: Aws Session Tags aws_bedrock_runtime_endpoint: anyOf: - type: string - type: 'null' title: Aws Bedrock Runtime Endpoint aws_bedrock_project_id: anyOf: - type: string - type: 'null' title: Aws Bedrock Project Id s3_bucket_name: anyOf: - type: string - type: 'null' title: S3 Bucket Name s3_region_name: anyOf: - type: string - type: 'null' title: S3 Region Name s3_encryption_key_id: anyOf: - type: string - type: 'null' title: S3 Encryption Key Id aws_batch_role_arn: anyOf: - type: string - type: 'null' title: Aws Batch Role Arn s3_output_bucket_name: anyOf: - type: string - type: 'null' title: S3 Output Bucket Name bedrock_tags: anyOf: - items: {} type: array - type: 'null' title: Bedrock Tags watsonx_region_name: anyOf: - type: string - type: 'null' title: Watsonx Region Name custom_llm_provider: anyOf: - type: string - type: 'null' title: Custom Llm Provider tpm: anyOf: - type: integer - type: 'null' title: Tpm rpm: anyOf: - type: integer - type: 'null' title: Rpm itpm: anyOf: - type: integer - type: 'null' title: Itpm otpm: anyOf: - type: integer - type: 'null' title: Otpm timeout: anyOf: - type: number - type: string - type: 'null' title: Timeout stream_timeout: anyOf: - type: number - type: string - type: 'null' title: Stream Timeout max_retries: anyOf: - type: integer - type: 'null' title: Max Retries drop_params: anyOf: - type: boolean - type: string - type: 'null' title: Drop Params organization: anyOf: - type: string - type: 'null' title: Organization configurable_clientside_auth_params: anyOf: - items: anyOf: - type: string - $ref: '#/components/schemas/ConfigurableClientsideParamsCustomAuth-Input' type: array - type: 'null' title: Configurable Clientside Auth Params litellm_credential_name: anyOf: - type: string - type: 'null' title: Litellm Credential Name litellm_trace_id: anyOf: - type: string - type: 'null' title: Litellm Trace Id max_file_size_mb: anyOf: - type: number - type: 'null' title: Max File Size Mb default_api_key_tpm_limit: anyOf: - type: integer - type: 'null' title: Default Api Key Tpm Limit default_api_key_rpm_limit: anyOf: - type: integer - type: 'null' title: Default Api Key Rpm Limit max_budget: anyOf: - type: number - type: 'null' title: Max Budget budget_duration: anyOf: - type: string - type: 'null' title: Budget Duration keepalive_seconds: anyOf: - type: number - type: 'null' title: Keepalive Seconds allow_client_keepalive_override: anyOf: - type: boolean - type: 'null' title: Allow Client Keepalive Override default: false use_in_pass_through: anyOf: - type: boolean - type: 'null' title: Use In Pass Through default: false use_litellm_proxy: anyOf: - type: boolean - type: 'null' title: Use Litellm Proxy default: false use_chat_completions_api: anyOf: - type: boolean - type: 'null' title: Use Chat Completions Api use_xai_oauth: anyOf: - type: boolean - type: 'null' title: Use Xai Oauth description: Use stored xAI OAuth credentials when no xAI API key is configured. default: false merge_reasoning_content_in_choices: anyOf: - type: boolean - type: 'null' title: Merge Reasoning Content In Choices default: false model_info: anyOf: - additionalProperties: true type: object - type: 'null' title: Model Info mock_response: anyOf: - type: string - $ref: '#/components/schemas/ModelResponse' - {} - type: 'null' title: Mock Response tags: anyOf: - items: type: string type: array - type: 'null' title: Tags tag_regex: anyOf: - items: type: string type: array - type: 'null' title: Tag Regex auto_router_config_path: anyOf: - type: string - type: 'null' title: Auto Router Config Path auto_router_config: anyOf: - type: string - type: 'null' title: Auto Router Config auto_router_default_model: anyOf: - type: string - type: 'null' title: Auto Router Default Model auto_router_embedding_model: anyOf: - type: string - type: 'null' title: Auto Router Embedding Model auto_router_max_input_chars: anyOf: - type: integer - type: 'null' title: Auto Router Max Input Chars auto_router_routing_compression: anyOf: - type: string - type: 'null' title: Auto Router Routing Compression auto_router_model_compression: anyOf: - type: string - type: 'null' title: Auto Router Model Compression complexity_router_config: anyOf: - additionalProperties: true type: object - type: 'null' title: Complexity Router Config complexity_router_default_model: anyOf: - type: string - type: 'null' title: Complexity Router Default Model adaptive_router_default_model: anyOf: - type: string - type: 'null' title: Adaptive Router Default Model adaptive_router_config: anyOf: - additionalProperties: true type: object - type: 'null' title: Adaptive Router Config quality_router_config: anyOf: - additionalProperties: true type: object - type: 'null' title: Quality Router Config quality_router_default_model: anyOf: - type: string - type: 'null' title: Quality Router Default Model vector_store_id: anyOf: - type: string - type: 'null' title: Vector Store Id milvus_text_field: anyOf: - type: string - type: 'null' title: Milvus Text Field milvus_db_name: anyOf: - type: string - type: 'null' title: Milvus Db Name milvus_partition_names: anyOf: - items: type: string type: array - type: 'null' title: Milvus Partition Names valkey_host: anyOf: - type: string - type: 'null' title: Valkey Host valkey_port: anyOf: - type: integer - type: 'null' title: Valkey Port valkey_password: anyOf: - type: string - type: 'null' title: Valkey Password valkey_ssl: anyOf: - type: boolean - type: 'null' title: Valkey Ssl valkey_text_field: anyOf: - type: string - type: 'null' title: Valkey Text Field valkey_embedding_field: anyOf: - type: string - type: 'null' title: Valkey Embedding Field model: anyOf: - type: string - type: 'null' title: Model additionalProperties: true type: object title: updateLiteLLMParams TierDomainStatistic: properties: tier: type: integer maximum: 4.0 minimum: 1.0 title: Tier successes: type: number minimum: 0.0 title: Successes observations: type: number exclusiveMinimum: 0.0 title: Observations request_type: $ref: '#/components/schemas/RequestType' type: object required: - tier - successes - observations - request_type title: TierDomainStatistic ConfigurableClientsideParamsCustomAuth-Output: properties: api_base: type: string title: Api Base type: object required: - api_base title: ConfigurableClientsideParamsCustomAuth ReminderMarkerPair: properties: open: type: string title: Open description: Opening delimiter, e.g. '' close: type: string title: Close description: Closing delimiter, e.g. '' type: object required: - open - close title: ReminderMarkerPair description: 'One open/close delimiter pair a harness wraps injected context in. Normalizing here rather than at the scan is what makes matching case-insensitive: markers reach the scan already lowered, so it lowercases only the haystack and never the needles. Stripping keeps YAML indentation whitespace from becoming part of the delimiter.' StandardLoggingRoutingDecision: properties: router_model_name: type: string title: Router Model Name router_type: type: string enum: - complexity - adaptive - quality title: Router Type routed_model: type: string title: Routed Model cause: type: string enum: - heuristic_scorer - heuristic_v2 - reasoning_override - llm_classifier - jev_classifier - heuristic_first_short_circuit - hybrid_short_circuit - classifier_plugin - classifier_fallback - default_model_fallback - literal_keyword_match - semantic_keyword_match - plan_mode - housekeeping - modality_escalation - modality_pin_override - health_failover - health_default_fallback - session_affinity_pin - session_affinity_escalation - user_turn_continuation - default_fallback - keyword - quality_tier - bandit title: Cause tier: type: string title: Tier tier_label: type: string title: Tier Label request_type: type: string title: Request Type score: type: number title: Score signals: items: type: string type: array title: Signals matched_keyword: type: string title: Matched Keyword escalation_keyword: type: string title: Escalation Keyword classifier_model: type: string title: Classifier Model classifier_cost: type: number title: Classifier Cost classifier_probabilities: additionalProperties: type: number type: object title: Classifier Probabilities classifier_confidence: type: number title: Classifier Confidence escalated: type: boolean title: Escalated context_escalated: type: boolean title: Context Escalated context_escalation_original_tier: type: string title: Context Escalation Original Tier tier_boundaries: $ref: '#/components/schemas/StandardLoggingRoutingDecisionTierBoundaries' reasoning_override_min_score: type: number title: Reasoning Override Min Score conversation_continuing: type: boolean title: Conversation Continuing savings_baseline_model: type: string title: Savings Baseline Model savings_baseline_deployment_id: type: string title: Savings Baseline Deployment Id tier_litellm_params: additionalProperties: true type: object title: Tier Litellm Params type: object title: StandardLoggingRoutingDecision description: Per-request provenance for a pre-routing strategy (auto-router) decision. Deployment: properties: model_name: type: string title: Model Name litellm_params: $ref: '#/components/schemas/LiteLLM_Params' model_info: $ref: '#/components/schemas/ModelInfo' additionalProperties: true type: object required: - model_name - litellm_params - model_info title: Deployment LiteLLM_ProxyModelTable: properties: model_id: type: string title: Model Id model_name: type: string title: Model Name litellm_params: additionalProperties: true type: object title: Litellm Params model_info: anyOf: - additionalProperties: true type: object - type: 'null' title: Model Info blocked: type: boolean title: Blocked default: false created_at: anyOf: - type: string format: date-time - type: 'null' title: Created At created_by: anyOf: - type: string - type: 'null' title: Created By updated_at: anyOf: - type: string format: date-time - type: 'null' title: Updated At updated_by: anyOf: - type: string - type: 'null' title: Updated By type: object required: - model_id - model_name - litellm_params title: LiteLLM_ProxyModelTable ModelInfo: properties: input_cost_per_token: anyOf: - type: number - type: 'null' title: Input Cost Per Token output_cost_per_token: anyOf: - type: number - type: 'null' title: Output Cost Per Token input_cost_per_character: anyOf: - type: number - type: 'null' title: Input Cost Per Character output_cost_per_character: anyOf: - type: number - type: 'null' title: Output Cost Per Character cache_read_input_token_cost: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost cache_creation_input_token_cost: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost tiered_pricing: anyOf: - items: additionalProperties: true type: object type: array - type: 'null' title: Tiered Pricing id: anyOf: - type: string - type: 'null' title: Id db_model: type: boolean title: Db Model default: false updated_at: anyOf: - type: string format: date-time - type: 'null' title: Updated At updated_by: anyOf: - type: string - type: 'null' title: Updated By created_at: anyOf: - type: string format: date-time - type: 'null' title: Created At created_by: anyOf: - type: string - type: 'null' title: Created By base_model: anyOf: - type: string - type: 'null' title: Base Model tier: anyOf: - type: string enum: - free - paid - type: 'null' title: Tier team_id: anyOf: - type: string - type: 'null' title: Team Id team_public_model_name: anyOf: - type: string - type: 'null' title: Team Public Model Name blocked: anyOf: - type: boolean - type: 'null' title: Blocked ptu_count: anyOf: - type: integer - type: 'null' title: Ptu Count cost_per_ptu_per_hour: anyOf: - type: number - type: 'null' title: Cost Per Ptu Per Hour ptu_effective_from: anyOf: - type: string format: date-time - type: 'null' title: Ptu Effective From ptu_effective_to: anyOf: - type: string format: date-time - type: 'null' title: Ptu Effective To allow_fail_open: anyOf: - type: boolean - type: 'null' title: Allow Fail Open enable_tag_filtering: anyOf: - type: boolean - type: 'null' title: Enable Tag Filtering internal_router_model: anyOf: - type: boolean - type: 'null' title: Internal Router Model additionalProperties: true type: object required: - id title: ModelInfo StandardLoggingRoutingDecisionTierBoundaries: properties: simple_medium: type: number title: Simple Medium medium_complex: type: number title: Medium Complex complex_reasoning: type: number title: Complex Reasoning type: object required: - simple_medium - medium_complex - complex_reasoning title: StandardLoggingRoutingDecisionTierBoundaries description: 'Snapshot of the complexity scorer''s tier boundaries at decision time, so a historical spend log row stays explainable after the router config changes.' BlockModelRequest: properties: model_id: type: string title: Model Id type: object required: - model_id title: BlockModelRequest NewModelGroupRequest: properties: access_group: type: string title: Access Group model_names: anyOf: - items: type: string type: array - type: 'null' title: Model Names model_ids: anyOf: - items: type: string type: array - type: 'null' title: Model Ids type: object required: - access_group title: NewModelGroupRequest ChatCompletionAnnotation: properties: type: type: string const: url_citation title: Type url_citation: $ref: '#/components/schemas/ChatCompletionAnnotationURLCitation' additionalProperties: true type: object title: ChatCompletionAnnotation ModelResponse: properties: id: type: string title: Id created: type: integer title: Created model: anyOf: - type: string - type: 'null' title: Model object: type: string title: Object system_fingerprint: anyOf: - type: string - type: 'null' title: System Fingerprint choices: items: $ref: '#/components/schemas/Choices' type: array title: Choices additionalProperties: true type: object required: - id - created - object - choices title: ModelResponse ChatCompletionCachedContent: properties: type: const: ephemeral title: Type type: string ttl: enum: - 5m - 1h title: Ttl type: string required: - type title: ChatCompletionCachedContent type: object PageLinks: properties: self: type: string title: Self prev: anyOf: - type: string - type: 'null' title: Prev next: anyOf: - type: string - type: 'null' title: Next type: object required: - self title: PageLinks description: 'Hypermedia for a paginated list. No `first`/`last`: without a total count the last page is unknown.' ClassifierVisionConfig: properties: enabled: type: boolean title: Enabled description: 'Forward image content to the classifier. Requires a classifier model declared supports_vision, on the deployment''s model_info or in the model cost map; images stay stripped otherwise, so a classifier that cannot read them is never sent one. Declare model_info.supports_vision on the deployment to enable a model the cost map does not describe. Only inline data: URIs are forwarded. A request whose images are http(s) URLs still classifies on its text alone, because some providers fetch such a URL from the proxy rather than the provider, which would let a caller aim a proxy-side request at an address of their choosing.' default: false max_images: type: integer minimum: 1.0 title: Max Images description: How many images from the newest user turn to forward, in wire order. Bounds the added cost of a turn that attaches many images. Images on earlier turns are never forwarded. default: 1 type: object title: ClassifierVisionConfig description: 'Whether the LLM classifier sees the images on the request it is classifying. Off by default because images cost far more than the text ask they arrive with, and the classifier runs on every request. A turn whose complexity lives in the image ("what is wrong in this stack trace screenshot") is invisible to a text-only classifier, which is what this buys.' UpdatePublicModelGroupsRequest: properties: model_groups: items: type: string type: array title: Model Groups description: List of model group names to make public additionalProperties: false type: object required: - model_groups title: UpdatePublicModelGroupsRequest description: Request model for updating public model groups updateDeployment: properties: model_name: anyOf: - type: string - type: 'null' title: Model Name litellm_params: anyOf: - $ref: '#/components/schemas/updateLiteLLMParams' - type: 'null' model_info: anyOf: - $ref: '#/components/schemas/ModelInfo' - type: 'null' blocked: anyOf: - type: boolean - type: 'null' title: Blocked type: object title: updateDeployment ChatCompletionAudioResponse: properties: id: type: string title: Id data: type: string title: Data expires_at: type: integer title: Expires At transcript: type: string title: Transcript additionalProperties: true type: object required: - id - data - expires_at - transcript title: ChatCompletionAudioResponse AdaptiveRouterWeights: properties: quality: type: number maximum: 1.0 minimum: 0.0 title: Quality default: 0.7 cost: type: number maximum: 1.0 minimum: 0.0 title: Cost default: 0.3 type: object title: AdaptiveRouterWeights RequestComplexityRouterConfig: properties: tiers: additionalProperties: anyOf: - type: string - items: type: string type: array type: object title: Tiers description: Mapping of complexity tiers to a model or model pool. A list is randomly picked from when adaptive=False, and used as a soft-floor home pool when adaptive=True tier_model_configs: additionalProperties: items: $ref: '#/components/schemas/ComplexityTierModel' type: array type: object title: Tier Model Configs enable_non_reasoning_tier: type: boolean title: Enable Non Reasoning Tier description: 'Add NON_REASONING as a fifth built-in tier below SIMPLE, for operational agent traffic that relays or reformats information rather than reasoning about it. Off by default: turning it on adds a rung to this router''s ladder, a bullet to the LLM classifier''s rubric, and a value the classifier may return, all of which move tier decisions and spend on an already-deployed router. Requires an LLM, Jev, or custom classifier plugin, since the heuristic scorers cannot produce the tier, and a model in `tiers` under the NON_REASONING key. Escalation still walks up from it, and it is never the savings baseline or a `heuristic_v2` prediction.' default: false tier_definitions: anyOf: - items: $ref: '#/components/schemas/TierDefinition' type: array - type: 'null' title: Tier Definitions description: 'Operator-defined tier set replacing the built-in SIMPLE/MEDIUM/COMPLEX/REASONING. Each entry''s name becomes a value the LLM classifier can return and its description becomes that tier''s rubric bullet; entries named after a built-in tier may omit the description and inherit the built-in criteria. List order is ascending severity and decides which tier wins when several keyword_tier_rules match. Requires classifier_type ''llm'', ''jev'' or ''custom'', a fallback_tier, and `tiers` keys matching the defined names exactly. Escalation, adaptive selection, session affinity, plugins, tier_labels, and the calibration-example rubric presets are unavailable with a custom tier set: the first four are built on the built-in tier ladder, and the last two rename or exemplify tiers the set replaces.' fallback_tier: anyOf: - type: string - type: 'null' title: Fallback Tier description: Tier routed to when the LLM classifier fails (timeout, provider error, or an unparseable reply). Required with tier_definitions and must name a defined tier; the heuristic scorer cannot produce custom tiers, so this replaces the heuristic fallback for custom tier sets. classification_prompt: anyOf: - type: string - type: 'null' title: Classification Prompt description: Replaces the classification instructions that open the LLM classifier rubric, and nothing else. The per-tier bullets follow it, the calibration examples follow those, and the trust-boundary paragraph telling the classifier to ignore tier requests embedded in quoted caller text is always appended after them and cannot be overridden. Requires an LLM classifier and cannot be combined with classifier_llm_config.system_prompt. With built-in tiers the rubric preset still supplies the tier criteria and, unless classification_examples replaces them, the calibration examples. classification_examples: anyOf: - type: string - type: 'null' title: Classification Examples description: 'Replaces the calibration examples of the LLM classifier rubric, and nothing else. Written as example lines only: the router renders the ''Calibration examples:'' heading above them, after the per-tier bullets. Requires an LLM classifier and cannot be combined with classifier_llm_config.system_prompt. With built-in tiers the rubric preset still supplies the tier criteria and, unless classification_prompt replaces them, the classification instructions; a custom tier set ships no examples of its own, so the section renders only when this is set.' tier_labels: additionalProperties: type: string propertyNames: $ref: '#/components/schemas/ComplexityTier' type: object title: Tier Labels description: 'Display names for the complexity tiers, so a deployment can use its own vocabulary (e.g. Cheap/Standard/Premium/Deep) in the dashboard, spend logs, and the LLM classifier rubric. Purely operator-facing: config keys stay canonical (tiers, keyword_tier_rules[].tier, tier_boundaries), API callers never see these names, and the heuristic scorer never reads them. Unlisted tiers keep their canonical name. Partial maps are allowed.' tier_boundaries: additionalProperties: type: number type: object title: Tier Boundaries description: Score boundaries between tiers. These keys (simple_medium, medium_complex, complex_reasoning) name the gaps between the default tier names and are not renameable by tier_labels; they are scorer knobs persisted by name on every routing decision reasoning_override_min_score: anyOf: - type: number - type: 'null' title: Reasoning Override Min Score description: Minimum weighted score a request must reach before 2+ reasoning markers may promote it to the reasoning tier. Unset tracks tier_boundaries.simple_medium, so the override never rescues a request the scorer placed in the cheapest tier; 0 restores the unconditional override token_thresholds: additionalProperties: type: integer type: object title: Token Thresholds description: Token count thresholds for simple/complex classification dimension_weights: additionalProperties: type: number type: object title: Dimension Weights description: Weights for each scoring dimension custom_dimensions: items: $ref: '#/components/schemas/CustomDimension' type: array maxItems: 16 title: Custom Dimensions description: 'Named dimensions added to the heuristic-v1 score. Each contributes its inline weight once when any keyword matches the current ask or a case-insensitive regex matches its first 2048 characters; scoring_mode ''match_count'' instead grades half weight for one distinct matcher and full for two or more. Regex quantifiers repeat one character or class at most 64 times. Unbounded quantifiers, repeated groups, backreferences and lookarounds are rejected. Conservative work limits include alternation paths, repeat lengths and subsequent matching: 2048 units per pattern, 8192 across the router. Only heuristic, heuristic_first and hybrid accept this field. Uses the existing heuristic tuning quota.' default: [] code_keywords: anyOf: - items: type: string type: array - type: 'null' title: Code Keywords description: Keywords indicating code-related content reasoning_keywords: anyOf: - items: type: string type: array - type: 'null' title: Reasoning Keywords description: Keywords indicating reasoning-required content technical_keywords: anyOf: - items: type: string type: array - type: 'null' title: Technical Keywords description: Keywords indicating technical content custom_technical_keywords: anyOf: - items: type: string type: array - type: 'null' title: Custom Technical Keywords description: Domain-specific technical keywords appended to the effective base list (technical_keywords if set, otherwise DEFAULT_TECHNICAL_KEYWORDS). Order is preserved; duplicates are removed case-insensitively against the base list and within this list. simple_keywords: anyOf: - items: type: string type: array - type: 'null' title: Simple Keywords description: Keywords indicating simple/basic queries default_model: anyOf: - type: string - type: 'null' title: Default Model description: Default model to use if tier cannot be determined return_raw_model_name: type: boolean title: Return Raw Model Name description: Return the resolved raw model name in the response model field instead of the client-requested complexity-router alias default: false classifier_type: type: string enum: - heuristic - heuristic_v2 - llm - custom - heuristic_first - hybrid - jev title: Classifier Type description: 'Classification strategy: local regex/keyword scoring, the bundled trained four-tier heuristic, an LLM call, a custom classifier plugin, ''heuristic_first'', which scores locally and only pays for the LLM classifier when the local scorer does not confidently land a cheap tier, or ''hybrid'', which trusts the local scorer everywhere except when its score lands near a tier boundary, or ''jev'', a TypeSafe AI Jev structured choice call' default: heuristic heuristic_v2_artifact: anyOf: - $ref: '#/components/schemas/TrainedTierArtifact' - type: string const: ultrafeedback title: Heuristic V2 Artifact description: Success-probability artifact used by classifier_type 'heuristic_v2'. The bundled UltraFeedback artifact is selected by default; an inline trained artifact may replace it default: ultrafeedback classifier_llm_config: anyOf: - $ref: '#/components/schemas/ClassifierLLMConfig' - type: 'null' description: Configuration for the LLM classifier; required when classifier_type is 'llm', 'heuristic_first' or 'hybrid' jev_classifier_config: anyOf: - $ref: '#/components/schemas/JevClassifierConfig' - type: 'null' heuristic_first_max_tier: anyOf: - type: string - type: 'null' title: Heuristic First Max Tier description: 'The highest tier the local scorer may decide on its own; required when classifier_type is ''heuristic_first'' and rejected otherwise. A request whose heuristic tier is at or below this one skips the LLM classifier and routes straight to that heuristic tier, so the classifier call is only paid for on traffic the scorer could not place cheaply. The scorer must also have produced at least one signal: a prompt where no dimension fired scores 0.0 and would otherwise land SIMPLE by default rather than by evidence, which is how a chained router would silently send unclassified traffic to the cheapest model. Names a built-in tier, and may not name the highest one, since that would make the LLM classifier unreachable.' hybrid_boundary_margin: anyOf: - type: number maximum: 1.0 minimum: 0.0 - type: 'null' title: Hybrid Boundary Margin description: How close to a tier boundary a heuristic score has to land before the LLM classifier breaks the tie; required when classifier_type is 'hybrid' and rejected otherwise. Everything further than this from every active boundary routes on the scorer's own tier with no classifier call, at any tier, which is what separates 'hybrid' from 'heuristic_first' and its cheap-tier ceiling. A prompt where no dimension fired still goes to the classifier, since the scorer has no opinion to be near a boundary with. 0 escalates only scores sitting exactly on a boundary. classifier_plugin: type: 'null' title: Classifier Plugin description: Not settable over HTTP; the classifier plugin is a runtime object classifier_plugin_timeout_ms: type: integer exclusiveMinimum: 0.0 title: Classifier Plugin Timeout Ms description: Timeout budget for the classifier plugin call, in milliseconds. On expiry the fallback path decides the tier. Only applies when classifier_type is 'custom'. default: 3000 classifier_fallback: type: string enum: - heuristic - default_model title: Classifier Fallback description: 'What classifies the request when the LLM classifier errors, times out, or returns an unparseable response. ''heuristic'' runs the local complexity scorer, which is right when the classifier grades complexity too. ''default_model'' skips scoring and routes to default_model, which is what a classifier on some other taxonomy wants: a prompt that grades data sensitivity has no use for a complexity score, and scoring one produces a tier unrelated to what the operator configured. Requires default_model when set to ''default_model''. Only applies when classifier_type is ''llm'', ''custom'', or ''heuristic_first''.' default: heuristic classifier_context_window_size: type: integer minimum: 0.0 title: Classifier Context Window Size description: Number of prior user turns (tool output and harness reminders excluded) to include as context in the LLM or JEV classifier input, so a follow-up like 'now do the same for the streaming path' is classified against what it refers to. Counts turns of both roles when classifier_context_include_assistant_turns is enabled. These turns are sent to the classifier model (the configured TypeSafe endpoint for JEV), which may be a different deployment or provider than the routed completion model; that call carries the current user ask and, except for Claude Code requests, the extracted system-role text in full. Claude Code system text is omitted to avoid classifying harness instructions; the routed completion still receives it. Set to 0 to omit prior turns and the conversation-depth summary; the current ask and selected system text are still sent. Applies to LLM and JEV classification. default: 3 classifier_context_budget_chars: type: integer minimum: 0.0 title: Classifier Context Budget Chars description: Maximum characters of prior-turn text quoted to the LLM or JEV classifier, across the whole context window, per classification call. Turns are taken newest first and quoted whole while they fit, so a conversation small enough to quote entirely is never cut; once the budget runs out the older turns are dropped whole and only the turn straddling the boundary is truncated, into whatever space is left. The current ask and, except for Claude Code requests, the extracted system-role text sit outside this budget and are sent in full, as does the numbering each quoted turn carries. A budget under 120 leaves no room to quote a turn and suppresses the block; set classifier_context_window_size to 0 to turn context off deliberately. Applies to LLM and JEV classification. default: 8000 classifier_context_per_turn_chars: anyOf: - type: integer exclusiveMinimum: 0.0 - type: 'null' title: Classifier Context Per Turn Chars description: Optional cap on each individual prior turn's text, applied before classifier_context_budget_chars bounds the block. Unset by default, so one long turn may spend the whole budget, which is usually what a follow-up needs; set it when no single turn should dominate the context the classifier sees. A capped turn keeps its opening and its ending with the middle elided. Applies to LLM and JEV classification. classifier_context_include_assistant_turns: type: boolean title: Classifier Context Include Assistant Turns description: 'Include assistant turns in the classifier context window, so difficulty stated by the model rather than by the user stays visible: a plan the assistant calls complex, which the user approves with ''yes'', is classified on the work being approved instead of on the word ''yes''. When enabled, classifier_context_window_size counts the last N turns of the conversation across both roles rather than the last N user turns, and assistant text is sent to the classifier model, which may be a different deployment or provider than the routed completion model. Assistant replies spend classifier_context_budget_chars alongside user turns, so raise it if the oldest turns stop being quoted once replies join the window. Off by default because enabling it shifts tier decisions, and therefore spend, for an already-deployed router. Applies to LLM and JEV classification.' default: false adaptive: type: boolean title: Adaptive description: Enable adaptive bandit selection with soft complexity floors default: false adaptive_weights: $ref: '#/components/schemas/AdaptiveRouterWeights' description: Quality vs cost weights for adaptive selection (used when adaptive=True) tier_distance_penalty: type: number minimum: 0.0 title: Tier Distance Penalty description: Score penalty per tier-step away from the classified tier when adaptive=True default: 0.5 adaptive_eligible: type: string enum: - all - classified_tier title: Adaptive Eligible description: 'When adaptive=True: ''all'' scores every pool model with a tier-distance penalty (soft floors); ''classified_tier'' Thompson-samples only inside the classified tier''s pool' default: all escalation_keywords: anyOf: - items: type: string type: array - type: 'null' title: Escalation Keywords description: Case-sensitive phrases a user can include to force a bump to the next-higher complexity tier when they aren't satisfied with results (they can force a stronger model, but not choose which one). Defaults to ['LITELLM ESCALATE'] when unset; set to an empty list to disable. keyword_tier_rules: anyOf: - items: $ref: '#/components/schemas/KeywordTierRule' type: array - type: 'null' title: Keyword Tier Rules description: Rules that force a specific tier when their keywords match the prompt stall_escalation_enabled: type: boolean title: Stall Escalation Enabled description: 'Escalate mid-task to the next-higher configured tier when the assistant''s own recent tool calls look stuck: the newest tool call repeats, or errors, at least stall_escalation_repeat_threshold times across the last stall_escalation_window calls. Both tests are anchored on the newest call, so a task that tried the same thing a few times and then moved on is not escalated on the strength of those older calls alone, while a retry loop broken up by an unrelated lookup still counts. One tier at most, on the same ladder escalation_keywords bumps along, and never above the highest configured tier. Detection re-runs on every classified turn from the tool calls visible in that request, so it needs no state and nothing survives past the task. Mutually exclusive with session_affinity and classification_mode=''user_turn'', which both replay a held routing decision instead of classifying most turns, so this would never see the tool calls to look at. Off by default.' default: false stall_escalation_window: type: integer exclusiveMinimum: 0.0 title: Stall Escalation Window description: How many of the assistant's most recent tool calls stall detection looks at, oldest ones dropped as new calls happen. Counted across the whole visible conversation rather than reset at the newest human ask, so evidence from before a plain follow-up message like 'try again' is still visible on the turn after it. default: 6 stall_escalation_repeat_threshold: type: integer minimum: 2.0 title: Stall Escalation Repeat Threshold description: How many of the last stall_escalation_window tool calls must repeat the newest call, or must have errored alongside it, before the task counts as stalled. Must not exceed stall_escalation_window, or the condition could never be reached. default: 3 plan_mode_min_tier: anyOf: - type: string - type: 'null' title: Plan Mode Min Tier description: 'When set, requests carrying a coding-agent plan-mode sentinel (Claude Code plan mode, VS Code Copilot Plan mode, Copilot CLI''s exit_plan_mode tool) are routed to at least this tier: the classified tier still wins when it is higher, and the floor also overrides a session-affinity pin to a lower tier for exactly the turns carrying the sentinel, without rewriting the pin -- the first turn after plan mode exits routes as if plan mode had never happened. Names a built-in tier, or with tier_definitions set, one of the defined tier names (list order is ascending severity, same as keyword_tier_rules). Unset disables detection entirely. The sentinels ride in client-injected prompt text, so a caller who pastes one can spend up to this tier''s models -- never down, and never outside the configured pools.' plan_mode_patterns: anyOf: - items: type: string type: array - type: 'null' title: Plan Mode Patterns description: Additional case-sensitive literal sentinels that mark a request as plan mode, on top of the built-in Claude Code and Copilot ones. For clients whose plan-mode wording the built-ins don't cover, or after a client release changes its strings. max_tokens_from_tier_model: type: boolean title: Max Tokens From Tier Model description: 'Set max_tokens on every routed request to the output ceiling of the tier model it lands on, replacing whatever the caller sent. A caller behind an auto-router cannot pick one value that fits every tier: the smallest tier''s ceiling starves a bigger tier''s thinking budget, and a bigger tier''s ceiling is rejected by the smallest. The ceiling is the smallest max_output_tokens across the tier model''s deployments, read from each deployment''s model_info and then the model cost map; a tier model with a deployment whose ceiling is unknown keeps the caller''s value. A max_tokens, max_completion_tokens or max_output_tokens in the tier''s own litellm_params still wins. Set false to forward the caller''s value unchanged.' default: true route_housekeeping_to_cheapest_tier: type: boolean title: Route Housekeeping To Cheapest Tier description: 'Route a coding agent''s own housekeeping calls to the cheapest configured tier without classifying them. A client names the conversation by quoting the whole session and asking for a title, so the ask reads as the session''s engineering work and lands on the most expensive tier, which is the reverse of what the call is worth. Detection is a literal match against client-owned sentinels on the newest ask only, so it cannot fire on an earlier turn, and it never lowers what anyone else asked for: a keyword_tier_rule or a session pin still decides instead, and an escalation keyword or the plan-mode floor still raises the tier from here. Only the classifier is displaced, and its call is skipped, so a matched request costs nothing to route. Set false to classify these calls like any other.' default: true housekeeping_patterns: anyOf: - items: type: string type: array - type: 'null' title: Housekeeping Patterns description: Additional case-sensitive literal sentinels that mark a request as client housekeeping, on top of the built-in conversation-title ones. For clients whose wording the built-ins don't cover, or after a client release changes its strings. enable_context_window_escalation: type: boolean title: Enable Context Window Escalation description: Escalate a request off a tier whose models provably cannot hold its prompt, before dispatch. The classifier scores complexity and never prompt size, so a long agentic session whose newest ask is trivial lands on a small-window tier and the provider rejects it with a context-window 400 that nothing retries. When every model of the decided tier has a declared window smaller than the estimated prompt, the request moves to the lowest configured tier with a model whose declared window fits; when only some of the tier's models fit, the pick is restricted to those and the tier keeps the request. Models with no resolvable window are never escalated away from and never escalated onto. Set false to dispatch on complexity alone, as before. default: true context_window_escalation_buffer: type: number maximum: 1.0 exclusiveMinimum: 0.0 title: Context Window Escalation Buffer description: Fraction of a model's declared context window the estimated prompt must fit within. The token count is an estimate, so fitting against the full window would dispatch prompts that the provider's own tokenizer then rejects; 0.95 leaves room for that drift plus the response tokens. default: 0.95 modality_routing: type: boolean title: Modality Routing description: Route image-bearing requests only to models that can accept image input. The classifier reads text alone, so an image request whose text classifies cheap otherwise lands on a text-only model and fails with a provider 400. When enabled, a routed model explicitly declared supports_vision false (deployment model_info or the model cost map; unmapped names stay routable) is replaced by the nearest HIGHER tier holding a capable model, then default_model, else a clear 400. A kept session-affinity pin still wins even when an image arrives, unless modality_pin_override is also enabled. default: false modality_pin_override: type: boolean title: Modality Pin Override description: Let modality_routing replace a kept session-affinity pin on the turns that carry an image. Without this, a session pinned to a text-only model fails every image turn with a provider 400, since the pin is exempt from the modality gate. When enabled, such a turn routes to a capable model for that request only and the stored pin is left untouched, so the next text turn replays the session's own model; the override is reported as cause modality_pin_override and is never itself pinned. Inert unless modality_routing is also enabled. default: false semantic_keyword_matching: type: boolean title: Semantic Keyword Matching description: Match keyword_tier_rules by embedding similarity instead of literal text default: false embedding_model: anyOf: - type: string - type: 'null' title: Embedding Model description: Embedding model (LiteLLM model name) used when semantic_keyword_matching is enabled match_threshold: type: number maximum: 1.0 minimum: 0.0 title: Match Threshold description: Minimum cosine similarity for a semantic keyword match default: 0.5 classification_mode: type: string enum: - every_request - user_turn title: Classification Mode description: 'When to run the complexity classifier. ''every_request'' (the default) classifies every inference request, including the tool-result continuation turns of an agentic loop. ''user_turn'' classifies only requests whose newest turn is a new human ask and replays the session''s held routing decision on continuation turns, which cuts classifier spend and eliminates mid-loop model switches. Continuations with no held decision to replay (no resolvable session_id, expired pin, fresh restart) still classify. Unlike session_affinity, a new human ask always re-classifies, so a session can still move tiers between asks. Suppressed when plugins are configured, for the same reason session_affinity is: a replayed decision would bypass the plugin pipeline.' default: every_request session_affinity: type: boolean title: Session Affinity description: 'When True and a session_id is resolvable on the request, pin the model chosen on the session''s first turn and reuse it for every later turn, skipping re-classification. Off by default so every turn is classified on its own merits and routed to the cheapest adequate tier. Set True to keep a multi-turn session on one model, which preserves provider prompt caches and avoids cross-model conversation-history errors. Always implies the deployment pin regardless of deployment_affinity: the session sticks to one deployment of the pinned model, since freezing the model while re-shuffling its deployments would still go cache-cold.' default: false deployment_affinity: type: boolean title: Deployment Affinity description: 'When True and a session_id is resolvable on the request, pin the deployment chosen inside each routed model group and reuse it whenever the session returns to that group, without pinning which group the session routes to. Independent of session_affinity, which pins the model group instead (and always carries this deployment pin with it): with session_affinity off, every turn is still classified on its own merits while a session that escalates to a stronger tier and comes back still lands on the deployment it used before, which is what keeps a provider prompt cache warm. Pins are held per model group, so switching tiers does not disturb the pin left behind in the previous group. On by default because re-shuffling a conversation across deployments of the same model discards that cache for no benefit; set False to keep every turn load-balanced across the group, which is what a deployment set with tight per-deployment rate limits wants. Inert when no session_id is resolvable, since there is nothing to key a pin on, and suppressed when plugins are configured, for the same reason session_affinity is.' default: true session_affinity_ttl_seconds: type: integer exclusiveMinimum: 0.0 title: Session Affinity Ttl Seconds description: TTL for the session affinity pin; refreshed on every cache hit. Bounds both the session_affinity model pin and the deployment_affinity deployment pin, so it measures idle time for the session's routing decisions rather than total session length default: 3600 plugins: type: 'null' title: Plugins description: Not settable over HTTP; routing plugins are runtime objects reminder_markers: anyOf: - items: $ref: '#/components/schemas/ReminderMarkerPair' type: array minItems: 1 - type: 'null' title: Reminder Markers description: Override the delimiter pairs used to recognize and strip harness-injected reminder blocks before classification. A harness that wraps injected context differently per agent type (main, subagent, cron) lists every pair it emits. Replaces, rather than adds to, the built-in system-reminder pair and the Codex envelope pairs enabled for Codex user agents, so list every built-in pair your harness also emits. Matching is case-insensitive. additionalProperties: true type: object title: RequestComplexityRouterConfig description: 'The part of a complexity-router config a request can carry. `plugins` holds live RoutingPlugin objects, which no JSON body can express and which have no OpenAPI schema, so it is closed off here rather than left as an arbitrary-type field.' AutoRouterRoutingTestRequest: properties: prompt: anyOf: - type: string - type: 'null' title: Prompt description: A single ask to route, as an end user would send it. Mutually exclusive with messages messages: anyOf: - items: additionalProperties: true type: object type: array - type: 'null' title: Messages description: The full message list to route, exactly as the serving path would receive it. Mutually exclusive with prompt system: anyOf: - type: string - items: additionalProperties: true type: object type: array - type: 'null' title: System description: The top-level system prompt an Anthropic /v1/messages body carries beside its messages tools: anyOf: - items: additionalProperties: true type: object type: array - type: 'null' title: Tools description: The tool definitions the request advertises, which decide whether the plan-mode floor applies complexity_router_config: $ref: '#/components/schemas/RequestComplexityRouterConfig' description: The complexity router config to route against, in the shape /model/new accepts saved_model_id: anyOf: - type: string minLength: 1 - type: 'null' title: Saved Model Id description: Test this saved deployment's server-side configuration instead of the supplied config and default model default_model: anyOf: - type: string - type: 'null' title: Default Model description: Model to route to when no tier resolves, i.e. complexity_router_default_model router_name: type: string title: Router Name description: Name reported as the router in the routing decision. Display only default: auto_router_routing_test team_id: anyOf: - type: string - type: 'null' title: Team Id description: Team the router is being created for. Required for a team admin, who may only test their own team's routers type: object required: - complexity_router_config title: AutoRouterRoutingTestRequest description: 'A single request to classify against a complexity-router config that need not be saved yet. Carries the same fields the serving path carries, so a dry run classifies what a real turn would classify. `messages`, `system` and `tools` are forwarded to the routing hook untranslated, which is why they are typed loosely: the hook reads whatever dialect the surface produced, and validating them against one surface''s schema would reject the others.' TierDataset: properties: name: type: string minLength: 1 title: Name url: type: string minLength: 1 title: Url license: type: string minLength: 1 title: License rows: type: integer exclusiveMinimum: 0.0 title: Rows success_definition: type: string minLength: 1 title: Success Definition default: quality score meets the dataset success threshold type: object required: - name - url - license - rows title: TierDataset AwsSessionTag: properties: Key: type: string title: Key Value: type: string title: Value additionalProperties: true type: object required: - Key - Value title: AwsSessionTag PublicModelHubInfo: properties: docs_title: type: string title: Docs Title custom_docs_description: anyOf: - type: string - type: 'null' title: Custom Docs Description litellm_version: type: string title: Litellm Version useful_links: anyOf: - additionalProperties: anyOf: - type: string - additionalProperties: true type: object type: object - type: 'null' title: Useful Links type: object required: - docs_title - custom_docs_description - litellm_version - useful_links title: PublicModelHubInfo ComplexityRouterConfigValidationRequest: properties: complexity_router_config: additionalProperties: true type: object title: Complexity Router Config team_id: anyOf: - type: string - type: 'null' title: Team Id description: Team the router is being created for. Required for a team admin, who may only validate their own team's routers type: object required: - complexity_router_config title: ComplexityRouterConfigValidationRequest description: 'A complexity-router config to validate without saving, so a form can surface the backend''s own verdict inline instead of a raw 400 at write time.' ChatCompletionThinkingBlock: properties: type: const: thinking title: Type type: string thinking: title: Thinking type: string signature: anyOf: - type: string - type: 'null' title: Signature cache_control: anyOf: - additionalProperties: true type: object - $ref: '#/components/schemas/ChatCompletionCachedContent' - type: 'null' title: Cache Control required: - type title: ChatCompletionThinkingBlock type: object ComplexityRouterConfigValidationResponse: properties: valid: type: boolean title: Valid error: anyOf: - type: string - type: 'null' title: Error type: object required: - valid title: ComplexityRouterConfigValidationResponse ModelDeprecationInfo: properties: model_name: type: string title: Model Name description: The public name of the model on the proxy (model_group). litellm_model: anyOf: - type: string - type: 'null' title: Litellm Model description: The underlying litellm model string the deprecation date is sourced from. deprecation_date: type: string format: date title: Deprecation Date description: The date (UTC) when the model becomes deprecated. days_until_deprecation: type: integer title: Days Until Deprecation description: Days remaining until the deprecation date. Negative if the model is already deprecated. status: type: string enum: - upcoming - imminent - deprecated title: Status description: '''deprecated'' if the date has passed, ''imminent'' if it falls within warn_within_days, ''upcoming'' otherwise.' litellm_provider: anyOf: - type: string - type: 'null' title: Litellm Provider description: The provider this model belongs to. type: object required: - model_name - deprecation_date - days_until_deprecation - status title: ModelDeprecationInfo Choices: properties: finish_reason: type: string enum: - stop - content_filter - function_call - tool_calls - length - guardrail_intervened - eos - finish_reason_unspecified - malformed_function_call title: Finish Reason index: type: integer title: Index message: $ref: '#/components/schemas/Message' logprobs: anyOf: - $ref: '#/components/schemas/ChoiceLogprobs' - {} - type: 'null' title: Logprobs provider_specific_fields: anyOf: - additionalProperties: true type: object - type: 'null' title: Provider Specific Fields additionalProperties: true type: object required: - finish_reason - index - message title: Choices ChatCompletionMessageToolCall: properties: {} additionalProperties: true type: object title: ChatCompletionMessageToolCall FacetListResponse: properties: data: items: type: string type: array title: Data meta: $ref: '#/components/schemas/PageMeta' links: $ref: '#/components/schemas/PageLinks' type: object required: - data - meta - links title: FacetListResponse description: The distinct values one column takes over a filtered query. `data` holds bare values, not entity rows. UpdateModelGroupRequest: properties: model_names: anyOf: - items: type: string type: array - type: 'null' title: Model Names model_ids: anyOf: - items: type: string type: array - type: 'null' title: Model Ids type: object title: UpdateModelGroupRequest AutoRouterClassifierPromptPreviewRequest: properties: tier_definitions: anyOf: - items: $ref: '#/components/schemas/TierDefinition' type: array - type: 'null' title: Tier Definitions tier_labels: anyOf: - additionalProperties: type: string propertyNames: $ref: '#/components/schemas/ComplexityTier' type: object - type: 'null' title: Tier Labels classification_rubric: anyOf: - $ref: '#/components/schemas/ClassificationRubric' - type: 'null' context_window_size: type: integer minimum: 0.0 title: Context Window Size default: 3 classification_prompt: anyOf: - type: string - type: 'null' title: Classification Prompt classification_examples: anyOf: - type: string - type: 'null' title: Classification Examples type: object title: AutoRouterClassifierPromptPreviewRequest description: 'A POST rather than query params: the classification sections are the operator''s own text, which must not reach access logs through a URL.' ListMeta: properties: total_count: type: integer title: Total Count page: type: integer title: Page page_size: type: integer title: Page Size total_pages: type: integer title: Total Pages type: object required: - total_count - page - page_size - total_pages title: ListMeta description: 'Page-mode counterpart to `PageMeta`: an entity list pays for the COUNT(*) so the table can show a page count.' TopLogprob: properties: token: type: string title: Token bytes: anyOf: - items: type: integer type: array - type: 'null' title: Bytes logprob: type: number title: Logprob additionalProperties: true type: object required: - token - logprob title: TopLogprob AccessGroupBudgetResponse: properties: access_group: type: string title: Access Group spend: type: number title: Spend budget: anyOf: - $ref: '#/components/schemas/AccessGroupBudget' - type: 'null' type: object required: - access_group - spend title: AccessGroupBudgetResponse AutoRouterRoutingTestResponse: properties: routed_model: type: string title: Routed Model description: The model group the router picked routed_model_configured: type: boolean title: Routed Model Configured description: Whether routed_model is a model group available to the caller, scoped to team_id when given. Never confirms models the caller could not use routing_decision: $ref: '#/components/schemas/StandardLoggingRoutingDecision' description: The decision record this request would have written to its log row type: object required: - routed_model - routed_model_configured - routing_decision title: AutoRouterRoutingTestResponse description: Where one prompt would have been routed, and why. TierGlobalStatistic: properties: tier: type: integer maximum: 4.0 minimum: 1.0 title: Tier successes: type: number minimum: 0.0 title: Successes observations: type: number exclusiveMinimum: 0.0 title: Observations type: object required: - tier - successes - observations title: TierGlobalStatistic ImageURLListItem: properties: image_url: $ref: '#/components/schemas/ImageURLObject' index: type: integer title: Index type: type: string const: image_url title: Type additionalProperties: true type: object required: - image_url - index - type title: ImageURLListItem ComplexityTier: type: string enum: - NON_REASONING - SIMPLE - MEDIUM - COMPLEX - REASONING title: ComplexityTier description: Complexity tiers for routing decisions. ChatCompletionCustomToolCallPayload: properties: name: type: string title: Name input: type: string title: Input additionalProperties: true type: object required: - name - input title: ChatCompletionCustomToolCallPayload RequestType: type: string enum: - code_generation - code_understanding - technical_design - analytical_reasoning - writing - factual_lookup - general title: RequestType description: Fixed v0 taxonomy. User-extensible types come in v1. DeleteModelGroupResponse: properties: access_group: type: string title: Access Group models_updated: type: integer title: Models Updated message: type: string title: Message type: object required: - access_group - models_updated - message title: DeleteModelGroupResponse UpdateUsefulLinksRequest: properties: useful_links: additionalProperties: anyOf: - type: string - additionalProperties: true type: object type: object title: Useful Links type: object required: - useful_links title: UpdateUsefulLinksRequest ChatCompletionAnnotationURLCitation: properties: end_index: type: integer title: End Index start_index: type: integer title: Start Index title: type: string title: Title url: type: string title: Url additionalProperties: true type: object title: ChatCompletionAnnotationURLCitation ConfigurableClientsideParamsCustomAuth-Input: properties: api_base: type: string title: Api Base additionalProperties: true type: object required: - api_base title: ConfigurableClientsideParamsCustomAuth ListAccessGroupsResponse: properties: access_groups: items: $ref: '#/components/schemas/AccessGroupInfo' type: array title: Access Groups type: object required: - access_groups title: ListAccessGroupsResponse ModelGroupInfoProxy: properties: model_group: type: string title: Model Group providers: items: type: string type: array title: Providers max_input_tokens: anyOf: - type: number - type: 'null' title: Max Input Tokens max_output_tokens: anyOf: - type: number - type: 'null' title: Max Output Tokens input_cost_per_token: anyOf: - type: number - type: 'null' title: Input Cost Per Token output_cost_per_token: anyOf: - type: number - type: 'null' title: Output Cost Per Token input_cost_per_pixel: anyOf: - type: number - type: 'null' title: Input Cost Per Pixel mode: anyOf: - type: string - type: string enum: - chat - embedding - completion - image_generation - audio_transcription - rerank - moderations - type: 'null' title: Mode default: chat tpm: anyOf: - type: integer - type: 'null' title: Tpm rpm: anyOf: - type: integer - type: 'null' title: Rpm itpm: anyOf: - type: integer - type: 'null' title: Itpm otpm: anyOf: - type: integer - type: 'null' title: Otpm supports_parallel_function_calling: type: boolean title: Supports Parallel Function Calling default: false supports_vision: type: boolean title: Supports Vision default: false supports_web_search: type: boolean title: Supports Web Search default: false supports_url_context: type: boolean title: Supports Url Context default: false supports_reasoning: type: boolean title: Supports Reasoning default: false supports_function_calling: type: boolean title: Supports Function Calling default: false supported_reasoning_efforts: anyOf: - items: type: string type: array - type: 'null' title: Supported Reasoning Efforts supported_openai_params: anyOf: - items: type: string type: array - type: 'null' title: Supported Openai Params default: [] configurable_clientside_auth_params: anyOf: - items: anyOf: - type: string - $ref: '#/components/schemas/ConfigurableClientsideParamsCustomAuth-Output' type: array - type: 'null' title: Configurable Clientside Auth Params is_public_model_group: type: boolean title: Is Public Model Group default: false health_status: anyOf: - type: string - type: 'null' title: Health Status health_response_time: anyOf: - type: number - type: 'null' title: Health Response Time health_checked_at: anyOf: - type: string - type: 'null' title: Health Checked At type: object required: - model_group - providers title: ModelGroupInfoProxy LiteLLM_Params: properties: input_cost_per_token: anyOf: - type: number - type: 'null' title: Input Cost Per Token output_cost_per_token: anyOf: - type: number - type: 'null' title: Output Cost Per Token input_cost_per_character: anyOf: - type: number - type: 'null' title: Input Cost Per Character output_cost_per_character: anyOf: - type: number - type: 'null' title: Output Cost Per Character cache_read_input_token_cost: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost cache_creation_input_token_cost: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost tiered_pricing: anyOf: - items: additionalProperties: true type: object type: array - type: 'null' title: Tiered Pricing input_cost_per_second: anyOf: - type: number - type: 'null' title: Input Cost Per Second output_cost_per_second: anyOf: - type: number - type: 'null' title: Output Cost Per Second output_cost_per_second_1080p: anyOf: - type: number - type: 'null' title: Output Cost Per Second 1080P output_cost_per_second_480p: anyOf: - type: number - type: 'null' title: Output Cost Per Second 480P output_cost_per_second_720p: anyOf: - type: number - type: 'null' title: Output Cost Per Second 720P output_cost_per_second_4k: anyOf: - type: number - type: 'null' title: Output Cost Per Second 4K input_cost_per_pixel: anyOf: - type: number - type: 'null' title: Input Cost Per Pixel output_cost_per_pixel: anyOf: - type: number - type: 'null' title: Output Cost Per Pixel input_cost_per_token_flex: anyOf: - type: number - type: 'null' title: Input Cost Per Token Flex input_cost_per_token_priority: anyOf: - type: number - type: 'null' title: Input Cost Per Token Priority input_cost_per_token_ultrafast: anyOf: - type: number - type: 'null' title: Input Cost Per Token Ultrafast cache_creation_input_token_cost_above_1hr: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Above 1Hr cache_creation_input_token_cost_above_200k_tokens: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Above 200K Tokens cache_creation_input_token_cost_above_272k_tokens: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Above 272K Tokens cache_creation_input_token_cost_above_272k_tokens_priority: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Above 272K Tokens Priority cache_creation_input_token_cost_above_272k_tokens_flex: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Above 272K Tokens Flex cache_creation_input_token_cost_flex: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Flex cache_creation_input_token_cost_priority: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Priority cache_creation_input_token_cost_ultrafast: anyOf: - type: number - type: 'null' title: Cache Creation Input Token Cost Ultrafast cache_creation_input_audio_token_cost: anyOf: - type: number - type: 'null' title: Cache Creation Input Audio Token Cost cache_read_input_token_cost_flex: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Flex cache_read_input_token_cost_priority: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Priority cache_read_input_token_cost_ultrafast: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Ultrafast cache_read_input_token_cost_above_200k_tokens: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Above 200K Tokens cache_read_input_token_cost_above_200k_tokens_priority: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Above 200K Tokens Priority cache_read_input_token_cost_above_272k_tokens_priority: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Above 272K Tokens Priority cache_read_input_token_cost_above_272k_tokens_flex: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Above 272K Tokens Flex cache_read_input_audio_token_cost: anyOf: - type: number - type: 'null' title: Cache Read Input Audio Token Cost input_cost_per_character_above_128k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Character Above 128K Tokens input_cost_per_audio_token: anyOf: - type: number - type: 'null' title: Input Cost Per Audio Token input_cost_per_token_cache_hit: anyOf: - type: number - type: 'null' title: Input Cost Per Token Cache Hit input_cost_per_token_above_128k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Token Above 128K Tokens input_cost_per_token_above_200k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Token Above 200K Tokens input_cost_per_token_above_200k_tokens_priority: anyOf: - type: number - type: 'null' title: Input Cost Per Token Above 200K Tokens Priority input_cost_per_token_above_272k_tokens_priority: anyOf: - type: number - type: 'null' title: Input Cost Per Token Above 272K Tokens Priority input_cost_per_token_above_272k_tokens_flex: anyOf: - type: number - type: 'null' title: Input Cost Per Token Above 272K Tokens Flex input_cost_per_query: anyOf: - type: number - type: 'null' title: Input Cost Per Query input_cost_per_image: anyOf: - type: number - type: 'null' title: Input Cost Per Image input_cost_per_image_above_128k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Image Above 128K Tokens input_cost_per_audio_per_second: anyOf: - type: number - type: 'null' title: Input Cost Per Audio Per Second input_cost_per_audio_per_second_above_128k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Audio Per Second Above 128K Tokens input_cost_per_video_per_second: anyOf: - type: number - type: 'null' title: Input Cost Per Video Per Second input_cost_per_video_per_second_above_128k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Video Per Second Above 128K Tokens input_cost_per_video_per_second_above_15s_interval: anyOf: - type: number - type: 'null' title: Input Cost Per Video Per Second Above 15S Interval input_cost_per_video_per_second_above_8s_interval: anyOf: - type: number - type: 'null' title: Input Cost Per Video Per Second Above 8S Interval input_cost_per_token_batches: anyOf: - type: number - type: 'null' title: Input Cost Per Token Batches output_cost_per_token_batches: anyOf: - type: number - type: 'null' title: Output Cost Per Token Batches output_cost_per_token_flex: anyOf: - type: number - type: 'null' title: Output Cost Per Token Flex output_cost_per_token_priority: anyOf: - type: number - type: 'null' title: Output Cost Per Token Priority output_cost_per_token_ultrafast: anyOf: - type: number - type: 'null' title: Output Cost Per Token Ultrafast output_cost_per_audio_token: anyOf: - type: number - type: 'null' title: Output Cost Per Audio Token output_cost_per_token_above_128k_tokens: anyOf: - type: number - type: 'null' title: Output Cost Per Token Above 128K Tokens output_cost_per_token_above_200k_tokens: anyOf: - type: number - type: 'null' title: Output Cost Per Token Above 200K Tokens output_cost_per_token_above_200k_tokens_priority: anyOf: - type: number - type: 'null' title: Output Cost Per Token Above 200K Tokens Priority output_cost_per_token_above_272k_tokens_priority: anyOf: - type: number - type: 'null' title: Output Cost Per Token Above 272K Tokens Priority output_cost_per_token_above_272k_tokens_flex: anyOf: - type: number - type: 'null' title: Output Cost Per Token Above 272K Tokens Flex output_cost_per_character_above_128k_tokens: anyOf: - type: number - type: 'null' title: Output Cost Per Character Above 128K Tokens output_cost_per_image: anyOf: - type: number - type: 'null' title: Output Cost Per Image output_cost_per_image_token: anyOf: - type: number - type: 'null' title: Output Cost Per Image Token output_cost_per_video_token: anyOf: - type: number - type: 'null' title: Output Cost Per Video Token output_cost_per_reasoning_token: anyOf: - type: number - type: 'null' title: Output Cost Per Reasoning Token output_cost_per_reasoning_token_flex: anyOf: - type: number - type: 'null' title: Output Cost Per Reasoning Token Flex output_cost_per_reasoning_token_priority: anyOf: - type: number - type: 'null' title: Output Cost Per Reasoning Token Priority output_cost_per_video_per_second: anyOf: - type: number - type: 'null' title: Output Cost Per Video Per Second output_cost_per_audio_per_second: anyOf: - type: number - type: 'null' title: Output Cost Per Audio Per Second search_context_cost_per_query: anyOf: - additionalProperties: true type: object - type: 'null' title: Search Context Cost Per Query google_maps_grounding_cost_per_query: anyOf: - type: number - type: 'null' title: Google Maps Grounding Cost Per Query citation_cost_per_token: anyOf: - type: number - type: 'null' title: Citation Cost Per Token cache_read_input_token_cost_above_272k_tokens: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Above 272K Tokens cache_read_input_token_cost_above_512k_tokens: anyOf: - type: number - type: 'null' title: Cache Read Input Token Cost Above 512K Tokens input_cost_per_image_token: anyOf: - type: number - type: 'null' title: Input Cost Per Image Token input_cost_per_video_token: anyOf: - type: number - type: 'null' title: Input Cost Per Video Token input_cost_per_token_above_272k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Token Above 272K Tokens input_cost_per_token_above_512k_tokens: anyOf: - type: number - type: 'null' title: Input Cost Per Token Above 512K Tokens output_cost_per_token_above_272k_tokens: anyOf: - type: number - type: 'null' title: Output Cost Per Token Above 272K Tokens output_cost_per_token_above_512k_tokens: anyOf: - type: number - type: 'null' title: Output Cost Per Token Above 512K Tokens output_vector_size: anyOf: - type: integer - type: 'null' title: Output Vector Size ocr_cost_per_page: anyOf: - type: number - type: 'null' title: Ocr Cost Per Page ocr_cost_per_credit: anyOf: - type: number - type: 'null' title: Ocr Cost Per Credit annotation_cost_per_page: anyOf: - type: number - type: 'null' title: Annotation Cost Per Page regional_processing_uplift_multiplier_eu: anyOf: - type: number - type: 'null' title: Regional Processing Uplift Multiplier Eu regional_processing_uplift_multiplier_us: anyOf: - type: number - type: 'null' title: Regional Processing Uplift Multiplier Us regional_endpoint_uplift_multiplier: anyOf: - type: number - type: 'null' title: Regional Endpoint Uplift Multiplier api_key: anyOf: - type: string - type: 'null' title: Api Key api_base: anyOf: - type: string - type: 'null' title: Api Base api_version: anyOf: - type: string - type: 'null' title: Api Version azure_ad_token: anyOf: - type: string - type: 'null' title: Azure Ad Token vertex_project: anyOf: - type: string - type: 'null' title: Vertex Project vertex_location: anyOf: - type: string - type: 'null' title: Vertex Location vertex_credentials: anyOf: - type: string - additionalProperties: true type: object - type: 'null' title: Vertex Credentials region_name: anyOf: - type: string - type: 'null' title: Region Name gcs_bucket_name: anyOf: - type: string - type: 'null' title: Gcs Bucket Name aws_access_key_id: anyOf: - type: string - type: 'null' title: Aws Access Key Id aws_secret_access_key: anyOf: - type: string - type: 'null' title: Aws Secret Access Key aws_session_token: anyOf: - type: string - type: 'null' title: Aws Session Token aws_region_name: anyOf: - type: string - type: 'null' title: Aws Region Name aws_session_name: anyOf: - type: string - type: 'null' title: Aws Session Name aws_profile_name: anyOf: - type: string - type: 'null' title: Aws Profile Name aws_role_name: anyOf: - type: string - type: 'null' title: Aws Role Name aws_web_identity_token: anyOf: - type: string - type: 'null' title: Aws Web Identity Token aws_sts_endpoint: anyOf: - type: string - type: 'null' title: Aws Sts Endpoint aws_external_id: anyOf: - type: string - type: 'null' title: Aws External Id aws_session_tags: anyOf: - items: $ref: '#/components/schemas/AwsSessionTag' type: array - type: 'null' title: Aws Session Tags aws_bedrock_runtime_endpoint: anyOf: - type: string - type: 'null' title: Aws Bedrock Runtime Endpoint aws_bedrock_project_id: anyOf: - type: string - type: 'null' title: Aws Bedrock Project Id s3_bucket_name: anyOf: - type: string - type: 'null' title: S3 Bucket Name s3_region_name: anyOf: - type: string - type: 'null' title: S3 Region Name s3_encryption_key_id: anyOf: - type: string - type: 'null' title: S3 Encryption Key Id aws_batch_role_arn: anyOf: - type: string - type: 'null' title: Aws Batch Role Arn s3_output_bucket_name: anyOf: - type: string - type: 'null' title: S3 Output Bucket Name bedrock_tags: anyOf: - items: {} type: array - type: 'null' title: Bedrock Tags watsonx_region_name: anyOf: - type: string - type: 'null' title: Watsonx Region Name custom_llm_provider: anyOf: - type: string - type: 'null' title: Custom Llm Provider tpm: anyOf: - type: integer - type: 'null' title: Tpm rpm: anyOf: - type: integer - type: 'null' title: Rpm itpm: anyOf: - type: integer - type: 'null' title: Itpm otpm: anyOf: - type: integer - type: 'null' title: Otpm timeout: anyOf: - type: number - type: string - type: 'null' title: Timeout stream_timeout: anyOf: - type: number - type: string - type: 'null' title: Stream Timeout max_retries: anyOf: - type: integer - type: 'null' title: Max Retries drop_params: anyOf: - type: boolean - type: string - type: 'null' title: Drop Params organization: anyOf: - type: string - type: 'null' title: Organization configurable_clientside_auth_params: anyOf: - items: anyOf: - type: string - $ref: '#/components/schemas/ConfigurableClientsideParamsCustomAuth-Input' type: array - type: 'null' title: Configurable Clientside Auth Params litellm_credential_name: anyOf: - type: string - type: 'null' title: Litellm Credential Name litellm_trace_id: anyOf: - type: string - type: 'null' title: Litellm Trace Id max_file_size_mb: anyOf: - type: number - type: 'null' title: Max File Size Mb default_api_key_tpm_limit: anyOf: - type: integer - type: 'null' title: Default Api Key Tpm Limit default_api_key_rpm_limit: anyOf: - type: integer - type: 'null' title: Default Api Key Rpm Limit max_budget: anyOf: - type: number - type: 'null' title: Max Budget budget_duration: anyOf: - type: string - type: 'null' title: Budget Duration keepalive_seconds: anyOf: - type: number - type: 'null' title: Keepalive Seconds allow_client_keepalive_override: anyOf: - type: boolean - type: 'null' title: Allow Client Keepalive Override default: false use_in_pass_through: anyOf: - type: boolean - type: 'null' title: Use In Pass Through default: false use_litellm_proxy: anyOf: - type: boolean - type: 'null' title: Use Litellm Proxy default: false use_chat_completions_api: anyOf: - type: boolean - type: 'null' title: Use Chat Completions Api use_xai_oauth: anyOf: - type: boolean - type: 'null' title: Use Xai Oauth description: Use stored xAI OAuth credentials when no xAI API key is configured. default: false merge_reasoning_content_in_choices: anyOf: - type: boolean - type: 'null' title: Merge Reasoning Content In Choices default: false model_info: anyOf: - additionalProperties: true type: object - type: 'null' title: Model Info mock_response: anyOf: - type: string - $ref: '#/components/schemas/ModelResponse' - {} - type: 'null' title: Mock Response tags: anyOf: - items: type: string type: array - type: 'null' title: Tags tag_regex: anyOf: - items: type: string type: array - type: 'null' title: Tag Regex auto_router_config_path: anyOf: - type: string - type: 'null' title: Auto Router Config Path auto_router_config: anyOf: - type: string - type: 'null' title: Auto Router Config auto_router_default_model: anyOf: - type: string - type: 'null' title: Auto Router Default Model auto_router_embedding_model: anyOf: - type: string - type: 'null' title: Auto Router Embedding Model auto_router_max_input_chars: anyOf: - type: integer - type: 'null' title: Auto Router Max Input Chars auto_router_routing_compression: anyOf: - type: string - type: 'null' title: Auto Router Routing Compression auto_router_model_compression: anyOf: - type: string - type: 'null' title: Auto Router Model Compression complexity_router_config: anyOf: - additionalProperties: true type: object - type: 'null' title: Complexity Router Config complexity_router_default_model: anyOf: - type: string - type: 'null' title: Complexity Router Default Model adaptive_router_default_model: anyOf: - type: string - type: 'null' title: Adaptive Router Default Model adaptive_router_config: anyOf: - additionalProperties: true type: object - type: 'null' title: Adaptive Router Config quality_router_config: anyOf: - additionalProperties: true type: object - type: 'null' title: Quality Router Config quality_router_default_model: anyOf: - type: string - type: 'null' title: Quality Router Default Model vector_store_id: anyOf: - type: string - type: 'null' title: Vector Store Id milvus_text_field: anyOf: - type: string - type: 'null' title: Milvus Text Field milvus_db_name: anyOf: - type: string - type: 'null' title: Milvus Db Name milvus_partition_names: anyOf: - items: type: string type: array - type: 'null' title: Milvus Partition Names valkey_host: anyOf: - type: string - type: 'null' title: Valkey Host valkey_port: anyOf: - type: integer - type: 'null' title: Valkey Port valkey_password: anyOf: - type: string - type: 'null' title: Valkey Password valkey_ssl: anyOf: - type: boolean - type: 'null' title: Valkey Ssl valkey_text_field: anyOf: - type: string - type: 'null' title: Valkey Text Field valkey_embedding_field: anyOf: - type: string - type: 'null' title: Valkey Embedding Field model: type: string title: Model additionalProperties: true type: object required: - model title: LiteLLM_Params description: LiteLLM Params with 'model' requirement - used for completions ModelDeprecationResponse: properties: deprecated: items: $ref: '#/components/schemas/ModelDeprecationInfo' type: array title: Deprecated description: Models whose deprecation date has already passed. imminent: items: $ref: '#/components/schemas/ModelDeprecationInfo' type: array title: Imminent description: Models whose deprecation date is within warn_within_days from today and require immediate migration planning. upcoming: items: $ref: '#/components/schemas/ModelDeprecationInfo' type: array title: Upcoming description: Models with a future deprecation date outside the warn window. warn_within_days: type: integer title: Warn Within Days description: The window (in days) used to bucket 'imminent' models. checked_at: type: string format: date-time title: Checked At description: UTC timestamp when the deprecation snapshot was generated. type: object required: - warn_within_days - checked_at title: ModelDeprecationResponse ChatCompletionRedactedThinkingBlock: properties: type: const: redacted_thinking title: Type type: string data: title: Data type: string cache_control: anyOf: - additionalProperties: true type: object - $ref: '#/components/schemas/ChatCompletionCachedContent' - type: 'null' title: Cache Control required: - type title: ChatCompletionRedactedThinkingBlock type: object CustomDimension: properties: name: type: string maxLength: 64 minLength: 1 pattern: ^[A-Za-z][A-Za-z0-9_]*$ title: Name weight: type: number maximum: 1.0 exclusiveMinimum: 0.0 title: Weight keywords: items: type: string maxLength: 256 minLength: 1 type: array maxItems: 32 title: Keywords default: [] patterns: items: type: string maxLength: 256 minLength: 1 type: array maxItems: 32 title: Patterns default: [] scoring_mode: type: string enum: - binary - match_count title: Scoring Mode description: '''binary'' scores 1 when any matcher hits. ''match_count'' scores 0.5 when one distinct matcher hits and 1 when two or more do; repeated occurrences of one matcher never raise it. Keywords are distinct case-insensitively, patterns by source, and a keyword and a pattern are always distinct from each other.' default: binary additionalProperties: false type: object required: - name - weight title: CustomDimension Message: properties: content: anyOf: - type: string - type: 'null' title: Content role: type: string enum: - assistant - user - system - tool - function title: Role tool_calls: anyOf: - items: anyOf: - $ref: '#/components/schemas/ChatCompletionMessageToolCall' - $ref: '#/components/schemas/ChatCompletionMessageCustomToolCall' type: array - type: 'null' title: Tool Calls function_call: anyOf: - $ref: '#/components/schemas/FunctionCall' - type: 'null' audio: anyOf: - $ref: '#/components/schemas/ChatCompletionAudioResponse' - type: 'null' images: anyOf: - items: $ref: '#/components/schemas/ImageURLListItem' type: array - type: 'null' title: Images reasoning_content: anyOf: - type: string - type: 'null' title: Reasoning Content thinking_blocks: anyOf: - items: anyOf: - $ref: '#/components/schemas/ChatCompletionThinkingBlock' - $ref: '#/components/schemas/ChatCompletionRedactedThinkingBlock' type: array - type: 'null' title: Thinking Blocks reasoning_items: anyOf: - items: $ref: '#/components/schemas/ChatCompletionReasoningItem' type: array - type: 'null' title: Reasoning Items provider_specific_fields: anyOf: - additionalProperties: true type: object - type: 'null' title: Provider Specific Fields annotations: anyOf: - items: $ref: '#/components/schemas/ChatCompletionAnnotation' type: array - type: 'null' title: Annotations additionalProperties: true type: object required: - content - role - tool_calls - function_call title: Message FunctionCall: properties: arguments: type: string title: Arguments name: anyOf: - type: string - type: 'null' title: Name additionalProperties: true type: object required: - arguments title: FunctionCall TrainedTierArtifact: properties: schema_version: type: integer const: 1 title: Schema Version default: 1 global_statistics: items: $ref: '#/components/schemas/TierGlobalStatistic' type: array title: Global Statistics domain_statistics: items: $ref: '#/components/schemas/TierDomainStatistic' type: array title: Domain Statistics default: [] cohort_statistics: items: $ref: '#/components/schemas/TierCohortStatistic' type: array title: Cohort Statistics default: [] domain_prior_mass: type: number exclusiveMinimum: 0.0 title: Domain Prior Mass default: 200.0 cohort_prior_mass: type: number exclusiveMinimum: 0.0 title: Cohort Prior Mass default: 20.0 routing_threshold: type: number maximum: 1.0 minimum: 0.0 title: Routing Threshold default: 0.75 datasets: items: $ref: '#/components/schemas/TierDataset' type: array title: Datasets default: [] success_definition: type: string minLength: 1 title: Success Definition default: quality score meets the dataset success threshold split_method: type: string minLength: 1 title: Split Method default: 'sha256(prompt): 70% train, 15% validation, 15% test' type: object required: - global_statistics title: TrainedTierArtifact ListResponse_ModelGroupInfoProxy_: properties: data: items: $ref: '#/components/schemas/ModelGroupInfoProxy' type: array title: Data meta: $ref: '#/components/schemas/ListMeta' links: $ref: '#/components/schemas/ListLinks' type: object required: - data - meta - links title: ListResponse[ModelGroupInfoProxy] ImageURLObject: properties: url: type: string title: Url detail: anyOf: - type: string - type: 'null' title: Detail additionalProperties: true type: object required: - url title: ImageURLObject AccessGroupBudgetRequest: properties: budget_id: anyOf: - type: string - type: 'null' title: Budget Id max_budget: anyOf: - type: number minimum: 0.0 - type: 'null' title: Max Budget soft_budget: anyOf: - type: number minimum: 0.0 - type: 'null' title: Soft Budget budget_duration: anyOf: - type: string - type: 'null' title: Budget Duration additionalProperties: false type: object title: AccessGroupBudgetRequest NewModelGroupResponse: properties: access_group: type: string title: Access Group model_names: anyOf: - items: type: string type: array - type: 'null' title: Model Names model_ids: anyOf: - items: type: string type: array - type: 'null' title: Model Ids models_updated: type: integer title: Models Updated type: object required: - access_group - models_updated title: NewModelGroupResponse DeleteAccessGroupBudgetResponse: properties: access_group: type: string title: Access Group budget_deleted: type: boolean title: Budget Deleted message: type: string title: Message type: object required: - access_group - budget_deleted - message title: DeleteAccessGroupBudgetResponse TierCohortStatistic: properties: tier: type: integer maximum: 4.0 minimum: 1.0 title: Tier successes: type: number minimum: 0.0 title: Successes observations: type: number exclusiveMinimum: 0.0 title: Observations cohort: type: string minLength: 1 title: Cohort type: object required: - tier - successes - observations - cohort title: TierCohortStatistic ClassifierLLMConfig: properties: model: type: string title: Model description: Model name (from the router's model_list) to call for classification vision: $ref: '#/components/schemas/ClassifierVisionConfig' description: Whether the classifier sees images on the request, and how many reasoning_effort: anyOf: - type: string enum: - none - minimal - low - medium - high - xhigh - max - type: 'null' title: Reasoning Effort description: Reasoning effort override for classifier calls. Leave unset to use the classifier deployment or provider default. timeout_ms: type: integer title: Timeout Ms description: Timeout budget for the classification call, in milliseconds default: 3000 circuit_breaker_enabled: type: boolean title: Circuit Breaker Enabled description: Whether one classifier timeout temporarily sends requests through classifier_fallback. Enabled by default so an unhealthy classifier cannot repeat its timeout across sessions. default: true circuit_breaker_cooldown_seconds: type: number exclusiveMinimum: 0.0 title: Circuit Breaker Cooldown Seconds description: How long to skip this router's LLM classifier after a classification call times out. Requests use classifier_fallback during the cooldown. When it expires, one request probes the classifier while concurrent requests keep using the fallback; a successful probe closes the circuit and a failed probe restarts the cooldown. default: 30.0 classification_rubric: anyOf: - $ref: '#/components/schemas/ClassificationRubric' - type: 'null' description: Which calibration examples the built-in rubric carries. 'agentic' anchors routine installs, builds, multi-file edits, and standard debugging at MEDIUM, so ordinary engineering does not route to the most expensive tier; it suits agent, terminal, and coding-assistant traffic as well as mixed traffic. 'chat' omits those engineering anchors, for a deployment serving only conversational traffic. 'business' carries business/sales anchors and business-flavored tier criteria that keep routine drafting and summarizing off the expensive tiers and reserve the top tier for committing to decisions under tradeoffs; it suits sales, support, and go-to-market traffic. Every preset keeps the same four tiers, so this moves where the boundary sits without changing the taxonomy. Leave unset for 'legacy', the rubric as it shipped before calibration examples existed, so an existing router's tier decisions and spend do not move on upgrade. Mutually exclusive with system_prompt, which replaces the rubric this would select. Only applies when classifier_type is 'llm'. system_prompt: anyOf: - type: string - type: 'null' title: System Prompt description: 'Replaces the built-in complexity rubric as the classifier''s entire system role. When set, neither the default rubric nor the context-window closing line is appended, so the prompt owns the whole taxonomy and the tier names SIMPLE/MEDIUM/COMPLEX/REASONING become whatever buckets it defines: a prompt that classifies data sensitivity routes on that instead of on difficulty. Two consequences of full replacement. The default rubric''s closing paragraph is the classifier''s prompt-injection defense, telling it that the caller''s quoted system prompt and prior turns are material to judge and never instructions; a replacement that omits it lets a caller ask for a tier and get it. And the heuristic fallback still scores complexity, so a router on some other taxonomy wants classifier_fallback=''default_model''. Leave unset for the built-in rubric. Only applies when classifier_type is ''llm''.' type: object required: - model title: ClassifierLLMConfig description: Configuration for the LLM-based complexity classifier. AccessGroupInfo: properties: access_group: type: string title: Access Group model_names: items: type: string type: array title: Model Names deployment_count: type: integer title: Deployment Count spend: anyOf: - type: number - type: 'null' title: Spend budget: anyOf: - $ref: '#/components/schemas/AccessGroupBudget' - type: 'null' type: object required: - access_group - model_names - deployment_count title: AccessGroupInfo TierDefinition: properties: name: type: string title: Name description: Tier name; becomes a value the LLM classifier can return and a key of `tiers` description: anyOf: - type: string - type: 'null' title: Description description: What belongs in this tier; rendered as this tier's bullet in the classifier rubric. Required unless the name is a built-in tier (NON_REASONING, SIMPLE, MEDIUM, COMPLEX, REASONING), which inherits the built-in criteria when omitted type: object required: - name title: TierDefinition description: 'An operator-defined tier: the name the LLM classifier must return and its rubric description.' ClassificationRubric: type: string enum: - legacy - agentic - chat - business title: ClassificationRubric description: Which calibration examples, and for BUSINESS which tier criteria, the built-in classifier rubric carries. securitySchemes: APIKeyHeader: type: apiKey description: Bearer token in: header name: x-litellm-api-key