---
openapi: 3.1.0
info:
title: Inference Gateway API
description: |
The API for interacting with various language models and other AI services.
OpenAI, Groq, Ollama, and other providers are supported.
OpenAI compatible API for using with existing clients.
Unified API for all providers.
contact:
name: Inference Gateway
url: https://inference-gateway.github.io/docs/
version: 1.0.0
license:
name: Apache-2.0
url: https://github.com/inference-gateway/inference-gateway/blob/main/LICENSE
servers:
- url: http://localhost:8080
description: Default server without version prefix for healthcheck and proxy and points
x-server-tags:
- Health
- Proxy
- url: http://localhost:8080/v1
description: Default server with version prefix for listing models and chat completions
x-server-tags:
- Models
- Completions
- Responses
- url: https://api.inference-gateway.local/v1
description: Local server with version prefix for listing models and chat completions
x-server-tags:
- Models
- Completions
- Responses
tags:
- name: Models
description: List and describe the various models available in the API.
- name: Completions
description: Generate completions from the models.
- name: Responses
description: Generate model responses using the OpenAI-compatible Responses API.
- name: MCP
description: List and manage MCP tools.
- name: Proxy
description: Proxy requests to provider endpoints.
- name: Metrics
description: Push metrics to the gateway (OTLP/HTTP).
- name: Health
description: Health check
paths:
/models:
get:
operationId: listModels
tags:
- Models
description: |
Lists the currently available models, and provides basic information
about each one such as the owner and availability.
summary:
Lists the currently available models, and provides basic information
about each one such as the owner and availability.
security:
- bearerAuth: []
parameters:
- name: provider
in: query
required: false
schema:
$ref: '#/components/schemas/Provider'
description: Specific provider to query (optional)
responses:
'200':
description: List of available models
content:
application/json:
schema:
$ref: '#/components/schemas/ListModelsResponse'
examples:
allProviders:
summary: Models from all providers
value:
object: 'list'
data:
- id: 'openai/gpt-4o'
object: 'model'
created: 1686935002
owned_by: 'openai'
served_by: 'openai'
- id: 'openai/llama-3.3-70b-versatile'
object: 'model'
created: 1723651281
owned_by: 'groq'
served_by: 'groq'
- id: 'cohere/claude-3-opus-20240229'
object: 'model'
created: 1708905600
owned_by: 'anthropic'
served_by: 'anthropic'
- id: 'cohere/command-r'
object: 'model'
created: 1707868800
owned_by: 'cohere'
served_by: 'cohere'
- id: 'ollama/phi3:3.8b'
object: 'model'
created: 1718441600
owned_by: 'ollama'
served_by: 'ollama'
- id: 'ollama_cloud/gpt-oss:20b'
object: 'model'
created: 1730419200
owned_by: 'ollama_cloud'
served_by: 'ollama_cloud'
- id: 'mistral/mistral-large-latest'
object: 'model'
created: 1698019200
owned_by: 'mistral'
served_by: 'mistral'
singleProvider:
summary: Models from a specific provider
value:
object: 'list'
data:
- id: 'openai/gpt-4o'
object: 'model'
created: 1686935002
owned_by: 'openai'
served_by: 'openai'
- id: 'openai/gpt-4-turbo'
object: 'model'
created: 1687882410
owned_by: 'openai'
served_by: 'openai'
- id: 'openai/gpt-3.5-turbo'
object: 'model'
created: 1677649963
owned_by: 'openai'
served_by: 'openai'
'401':
$ref: '#/components/responses/Unauthorized'
'500':
$ref: '#/components/responses/InternalError'
/chat/completions:
post:
operationId: createChatCompletion
tags:
- Completions
description: |
Generates a chat completion based on the provided input.
The completion can be streamed to the client as it is generated.
summary: Create a chat completion
security:
- bearerAuth: []
parameters:
- name: provider
in: query
required: false
schema:
$ref: '#/components/schemas/Provider'
description: Specific provider to use (default determined by model)
requestBody:
$ref: '#/components/requestBodies/CreateChatCompletionRequest'
responses:
'200':
description: Successful response
content:
application/json:
schema:
$ref: '#/components/schemas/CreateChatCompletionResponse'
text/event-stream:
schema:
description: |
Server-Sent Events stream. Each frame is an `SSEvent` whose
`data` field contains the JSON-serialized payload for that
event. For content/message chunk events the payload is a
`CreateChatCompletionStreamResponse`. The `oneOf` here makes
the streaming payload schemas reachable from this operation
so that code generators emit types for them.
oneOf:
- $ref: '#/components/schemas/SSEvent'
- $ref: '#/components/schemas/CreateChatCompletionStreamResponse'
'400':
$ref: '#/components/responses/BadRequest'
'401':
$ref: '#/components/responses/Unauthorized'
'500':
$ref: '#/components/responses/InternalError'
/responses:
post:
operationId: createResponse
tags:
- Responses
description: |
Creates a model response using the OpenAI-compatible Responses API.
The request accepts either a single text input or a list of input
items (allowing batched, multi-turn input in one request), and the
result can be streamed to the client as it is generated.
Not every provider implements the Responses API. Requests routed to a
provider that does not support it return `400 Bad Request` with an
explanatory error message; use `/chat/completions` for those providers.
summary: Create a model response
security:
- bearerAuth: []
parameters:
- name: provider
in: query
required: false
schema:
$ref: '#/components/schemas/Provider'
description: Specific provider to use (default determined by model)
requestBody:
$ref: '#/components/requestBodies/CreateResponseRequest'
responses:
'200':
description: Successful response
content:
application/json:
schema:
$ref: '#/components/schemas/Response'
text/event-stream:
schema:
description: |
Server-Sent Events stream. Each frame is an `SSEvent` whose
`data` field contains the JSON-serialized payload for that
event. For Responses streaming the payload is a
`ResponseStreamEvent`. The `oneOf` here makes the streaming
payload schemas reachable from this operation so that code
generators emit types for them.
oneOf:
- $ref: '#/components/schemas/SSEvent'
- $ref: '#/components/schemas/ResponseStreamEvent'
'400':
$ref: '#/components/responses/ResponsesNotSupported'
'401':
$ref: '#/components/responses/Unauthorized'
'500':
$ref: '#/components/responses/InternalError'
/mcp/tools:
get:
operationId: listTools
tags:
- MCP
description: |
Lists the currently available MCP tools. Only accessible when EXPOSE_MCP is enabled.
summary: Lists the currently available MCP tools
security:
- bearerAuth: []
responses:
'200':
description: Successful response
content:
application/json:
schema:
$ref: '#/components/schemas/ListToolsResponse'
'401':
$ref: '#/components/responses/Unauthorized'
'403':
$ref: '#/components/responses/MCPNotExposed'
'500':
$ref: '#/components/responses/InternalError'
/metrics:
post:
operationId: pushMetrics
tags:
- Metrics
description: |
OTLP/HTTP metrics push endpoint. Accepts an OTLP ExportMetricsServiceRequest
encoded as protobuf or JSON. Only accessible when TELEMETRY_ENABLE and
TELEMETRY_METRICS_PUSH_ENABLE are enabled.
summary: Push metrics to the gateway (OTLP/HTTP)
security:
- bearerAuth: []
requestBody:
required: true
description: OTLP ExportMetricsServiceRequest payload
content:
application/x-protobuf:
schema:
type: string
format: binary
application/json:
schema:
type: object
responses:
'200':
description: OTLP ExportMetricsServiceResponse, possibly with partial success details
content:
application/x-protobuf:
schema:
type: string
format: binary
application/json:
schema:
type: object
'400':
$ref: '#/components/responses/BadRequest'
'401':
$ref: '#/components/responses/Unauthorized'
'403':
description: Metrics push is not enabled
'413':
description: Payload too large
'415':
description: Unsupported content type
/proxy/{provider}/{path}:
parameters:
- name: provider
in: path
required: true
schema:
$ref: '#/components/schemas/Provider'
- name: path
in: path
required: true
style: simple
explode: false
schema:
type: string
description: The remaining path to proxy to the provider
get:
operationId: proxyGet
tags:
- Proxy
description: |
Proxy GET request to provider
The request body depends on the specific provider and endpoint being called.
If you decide to use this approach, please follow the provider-specific documentations.
summary: Proxy GET request to provider
responses:
'200':
$ref: '#/components/responses/ProviderResponse'
'400':
$ref: '#/components/responses/BadRequest'
'401':
$ref: '#/components/responses/Unauthorized'
'500':
$ref: '#/components/responses/InternalError'
security:
- bearerAuth: []
post:
operationId: proxyPost
tags:
- Proxy
description: |
Proxy POST request to provider
The request body depends on the specific provider and endpoint being called.
If you decide to use this approach, please follow the provider-specific documentations.
summary: Proxy POST request to provider
requestBody:
$ref: '#/components/requestBodies/ProviderRequest'
responses:
'200':
$ref: '#/components/responses/ProviderResponse'
'400':
$ref: '#/components/responses/BadRequest'
'401':
$ref: '#/components/responses/Unauthorized'
'500':
$ref: '#/components/responses/InternalError'
security:
- bearerAuth: []
put:
operationId: proxyPut
tags:
- Proxy
description: |
Proxy PUT request to provider
The request body depends on the specific provider and endpoint being called.
If you decide to use this approach, please follow the provider-specific documentations.
summary: Proxy PUT request to provider
requestBody:
$ref: '#/components/requestBodies/ProviderRequest'
responses:
'200':
$ref: '#/components/responses/ProviderResponse'
'400':
$ref: '#/components/responses/BadRequest'
'401':
$ref: '#/components/responses/Unauthorized'
'500':
$ref: '#/components/responses/InternalError'
security:
- bearerAuth: []
delete:
operationId: proxyDelete
tags:
- Proxy
description: |
Proxy DELETE request to provider
The request body depends on the specific provider and endpoint being called.
If you decide to use this approach, please follow the provider-specific documentations.
summary: Proxy DELETE request to provider
responses:
'200':
$ref: '#/components/responses/ProviderResponse'
'400':
$ref: '#/components/responses/BadRequest'
'401':
$ref: '#/components/responses/Unauthorized'
'500':
$ref: '#/components/responses/InternalError'
security:
- bearerAuth: []
patch:
operationId: proxyPatch
tags:
- Proxy
description: |
Proxy PATCH request to provider
The request body depends on the specific provider and endpoint being called.
If you decide to use this approach, please follow the provider-specific documentations.
summary: Proxy PATCH request to provider
requestBody:
$ref: '#/components/requestBodies/ProviderRequest'
responses:
'200':
$ref: '#/components/responses/ProviderResponse'
'400':
$ref: '#/components/responses/BadRequest'
'401':
$ref: '#/components/responses/Unauthorized'
'500':
$ref: '#/components/responses/InternalError'
security:
- bearerAuth: []
/health:
get:
operationId: healthCheck
tags:
- Health
description: |
Health check endpoint
Returns a 200 status code if the service is healthy
summary: Health check
responses:
'200':
description: Health check successful
components:
requestBodies:
ProviderRequest:
required: true
description: |
ProviderRequest depends on the specific provider and endpoint being called
If you decide to use this approach, please follow the provider-specific documentations.
content:
application/json:
schema:
type: object
properties:
model:
type: string
messages:
type: array
items:
type: object
properties:
role:
type: string
content:
type: string
temperature:
type: number
format: float
default: 0.7
examples:
openai:
summary: OpenAI chat completion request
value:
model: 'gpt-3.5-turbo'
messages:
- role: 'user'
content: 'Hello! How can I assist you today?'
temperature: 0.7
anthropic:
summary: Anthropic Claude request
value:
model: 'claude-3-opus-20240229'
messages:
- role: 'user'
content: 'Explain quantum computing'
temperature: 0.5
mistral:
summary: Mistral AI request
value:
model: 'mistral-large-latest'
messages:
- role: 'user'
content: 'Write a Python function to calculate fibonacci numbers'
temperature: 0.3
CreateChatCompletionRequest:
required: true
description: |
ProviderRequest depends on the specific provider and endpoint being called
If you decide to use this approach, please follow the provider-specific documentations.
content:
application/json:
schema:
$ref: '#/components/schemas/CreateChatCompletionRequest'
CreateResponseRequest:
required: true
description: |
Request payload for the Responses API. Mirrors the OpenAI
`POST /v1/responses` request body.
content:
application/json:
schema:
$ref: '#/components/schemas/CreateResponseRequest'
responses:
BadRequest:
description: Bad request
content:
application/json:
schema:
$ref: '#/components/schemas/Error'
Unauthorized:
description: Unauthorized
content:
application/json:
schema:
$ref: '#/components/schemas/Error'
InternalError:
description: Internal server error
content:
application/json:
schema:
$ref: '#/components/schemas/Error'
MCPNotExposed:
description: MCP tools endpoint is not exposed
content:
application/json:
schema:
$ref: '#/components/schemas/Error'
example:
error: 'MCP tools endpoint is not exposed. Set EXPOSE_MCP=true to enable.'
ResponsesNotSupported:
description: |
The selected provider does not implement the Responses API. The
gateway returns this when a request is routed to a provider without
Responses support.
content:
application/json:
schema:
$ref: '#/components/schemas/Error'
example:
error: 'The Responses API is not supported by this provider yet.'
ProviderResponse:
description: |
ProviderResponse depends on the specific provider and endpoint being called
If you decide to use this approach, please follow the provider-specific documentations.
content:
application/json:
schema:
$ref: '#/components/schemas/ProviderSpecificResponse'
examples:
openai:
summary: OpenAI API response
value:
{
'id': 'chatcmpl-123',
'object': 'chat.completion',
'created': 1677652288,
'model': 'gpt-3.5-turbo',
'choices':
[
{
'index': 0,
'message':
{
'role': 'assistant',
'content': 'Hello! How can I help you today?',
},
'finish_reason': 'stop',
},
],
}
mistral:
summary: Mistral AI response
value:
{
'id': 'cmpl-123',
'object': 'chat.completion',
'created': 1677652288,
'model': 'mistral-large-latest',
'choices':
[
{
'index': 0,
'message':
{
'role': 'assistant',
'content': 'def fibonacci(n):\n if n <= 1:\n return n\n return fibonacci(n-1) + fibonacci(n-2)',
},
'finish_reason': 'stop',
},
],
}
securitySchemes:
bearerAuth:
type: http
scheme: bearer
bearerFormat: JWT
description: |
Authentication is optional by default.
To enable authentication, set AUTH_ENABLE to true.
When enabled, requests must include a valid JWT token in the Authorization header.
schemas:
Provider:
type: string
enum:
- ollama
- ollama_cloud
- groq
- llamacpp
- openai
- cloudflare
- cohere
- anthropic
- deepseek
- google
- mistral
- minimax
- moonshot
- nvidia
- zai
x-provider-configs:
ollama:
id: 'ollama'
url: 'http://ollama:8080/v1'
auth_type: 'none'
supports_vision: true
endpoints:
models:
name: 'list_models'
method: 'GET'
endpoint: '/models'
chat:
name: 'chat_completions'
method: 'POST'
endpoint: '/chat/completions'
ollama_cloud:
id: 'ollama_cloud'
url: 'https://ollama.com/v1'
auth_type: 'bearer'
supports_vision: true
endpoints:
models:
name: 'list_models'
method: 'GET'
endpoint: '/models'
chat:
name: 'chat_completions'
method: 'POST'
endpoint: '/chat/completions'
anthropic:
id: 'anthropic'
url: 'https://api.anthropic.com/v1'
auth_type: 'xheader'
supports_vision: true
extra_headers:
anthropic-version: '2023-06-01'
endpoints:
models:
name: 'list_models'
method: 'GET'
endpoint: '/models'
chat:
name: 'chat_completions'
method: 'POST'
endpoint: '/chat/completions'
cohere:
id: 'cohere'
url: 'https://api.cohere.ai'
auth_type: 'bearer'
supports_vision: true
endpoints:
models:
name: 'list_models'
method: 'GET'
endpoint: '/v1/models'
chat:
name: 'chat_completions'
method: 'POST'
endpoint: '/compatibility/v1/chat/completions'
groq:
id: 'groq'
url: 'https://api.groq.com/openai/v1'
auth_type: 'bearer'
supports_vision: true
endpoints:
models:
name: 'list_models'
method: 'GET'
endpoint: '/models'
chat:
name: 'chat_completions'
method: 'POST'
endpoint: '/chat/completions'
llamacpp:
id: 'llamacpp'
url: 'http://llamacpp:8080/v1'
auth_type: 'bearer'
supports_vision: true
endpoints:
models:
name: 'list_models'
method: 'GET'
endpoint: '/models'
chat:
name: 'chat_completions'
method: 'POST'
endpoint: '/chat/completions'
openai:
id: 'openai'
url: 'https://api.openai.com/v1'
auth_type: 'bearer'
supports_vision: true
endpoints:
models:
name: 'list_models'
method: 'GET'
endpoint: '/models'
chat:
name: 'chat_completions'
method: 'POST'
endpoint: '/chat/completions'
responses:
name: 'responses'
method: 'POST'
endpoint: '/responses'
cloudflare:
id: 'cloudflare'
url: 'https://api.cloudflare.com/client/v4/accounts/{ACCOUNT_ID}/ai'
auth_type: 'bearer'
supports_vision: false
endpoints:
models:
name: 'list_models'
method: 'GET'
endpoint: '/finetunes/public?limit=1000'
chat:
name: 'chat_completions'
method: 'POST'
endpoint: '/v1/chat/completions'
deepseek:
id: 'deepseek'
url: 'https://api.deepseek.com'
auth_type: 'bearer'
supports_vision: false
endpoints:
models:
name: 'list_models'
method: 'GET'
endpoint: '/models'
chat:
name: 'chat_completions'
method: 'POST'
endpoint: '/chat/completions'
google:
id: 'google'
url: 'https://generativelanguage.googleapis.com/v1beta/openai'
auth_type: 'bearer'
supports_vision: true
endpoints:
models:
name: 'list_models'
method: 'GET'
endpoint: '/models'
chat:
name: 'chat_completions'
method: 'POST'
endpoint: '/chat/completions'
mistral:
id: 'mistral'
url: 'https://api.mistral.ai/v1'
auth_type: 'bearer'
supports_vision: true
endpoints:
models:
name: 'list_models'
method: 'GET'
endpoint: '/models'
chat:
name: 'chat_completions'
method: 'POST'
endpoint: '/chat/completions'
minimax:
id: 'minimax'
url: 'https://api.minimax.io/v1'
auth_type: 'bearer'
supports_vision: true
endpoints:
models:
name: 'list_models'
method: 'GET'
endpoint: '/models'
chat:
name: 'chat_completions'
method: 'POST'
endpoint: '/chat/completions'
moonshot:
id: 'moonshot'
url: 'https://api.moonshot.ai/v1'
auth_type: 'bearer'
supports_vision: true
endpoints:
models:
name: 'list_models'
method: 'GET'
endpoint: '/models'
chat:
name: 'chat_completions'
method: 'POST'
endpoint: '/chat/completions'
nvidia:
id: 'nvidia'
url: 'https://integrate.api.nvidia.com/v1'
auth_type: 'bearer'
supports_vision: true
endpoints:
models:
name: 'list_models'
method: 'GET'
endpoint: '/models'
chat:
name: 'chat_completions'
method: 'POST'
endpoint: '/chat/completions'
zai:
id: 'zai'
url: 'https://api.z.ai/v1'
auth_type: 'bearer'
supports_vision: true
endpoints:
models:
name: 'list_models'
method: 'GET'
endpoint: '/models'
chat:
name: 'chat_completions'
method: 'POST'
endpoint: '/chat/completions'
ProviderSpecificResponse:
type: object
description: |
Provider-specific response format. Examples:
OpenAI GET /v1/models?provider=openai response:
```json
{
"provider": "openai",
"object": "list",
"data": [
{
"id": "gpt-4",
"object": "model",
"created": 1687882410,
"owned_by": "openai",
"served_by": "openai"
}
]
}
```
Anthropic GET /v1/models?provider=anthropic response:
```json
{
"provider": "anthropic",
"object": "list",
"data": [
{
"id": "gpt-4",
"object": "model",
"created": 1687882410,
"owned_by": "openai",
"served_by": "openai"
}
]
}
```
ProviderAuthType:
type: string
description: Authentication type for providers
enum:
- bearer
- xheader
- query
- none
SSEvent:
type: object
properties:
event:
type: string
enum:
- message-start
- stream-start
- content-start
- content-delta
- content-end
- message-end
- stream-end
data:
type: string
format: byte
retry:
type: integer
Endpoints:
type: object
properties:
models:
type: string
chat:
type: string
responses:
type: string
required:
- models
- chat
Error:
type: object
properties:
error:
type: string
MessageRole:
type: string
description: Role of the message sender
enum:
- system
- user
- assistant
- tool
x-enum-varnames:
- System
- User
- Assistant
- Tool
Message:
type: object
description: Message structure for provider requests
properties:
role:
$ref: '#/components/schemas/MessageRole'
content:
$ref: '#/components/schemas/MessageContent'
tool_calls:
type: array
items:
$ref: '#/components/schemas/ChatCompletionMessageToolCall'
tool_call_id:
type: string
reasoning_content:
type: string
description: The reasoning content of the chunk message.
reasoning:
type: string
description: The reasoning of the chunk message. Same as reasoning_content.
required:
- role
- content
MessageContent:
description: Message content - either text or multimodal content parts
oneOf:
- type: string
description: Text content (backward compatibility)
- type: array
items:
$ref: '#/components/schemas/ContentPart'
description: Array of content parts for multimodal messages
ContentPart:
type: object
description: A content part within a multimodal message
oneOf:
- $ref: '#/components/schemas/TextContentPart'
- $ref: '#/components/schemas/ImageContentPart'
TextContentPart:
type: object
description: Text content part
properties:
type:
type: string
enum:
- text
description: Content type identifier
text:
type: string
description: The text content
required:
- type
- text
ImageContentPart:
type: object
description: Image content part
properties:
type:
type: string
enum:
- image_url
description: Content type identifier
image_url:
$ref: '#/components/schemas/ImageURL'
required:
- type
- image_url
ImageURL:
type: object
description: Image URL configuration
properties:
url:
type: string
description: URL of the image (data URLs supported)
detail:
type: string
enum:
- auto
- low
- high
x-enum-varnames:
- ImageURLDetailAuto
- ImageURLDetailLow
- ImageURLDetailHigh
default: auto
description: Image detail level for vision processing
required:
- url
Model:
type: object
description: Common model information
properties:
id:
type: string
object:
type: string
created:
type: integer
format: int64
owned_by:
type: string
served_by:
$ref: '#/components/schemas/Provider'
required:
- id
- object
- created
- owned_by
- served_by
ListModelsResponse:
type: object
description: Response structure for listing models
properties:
provider:
$ref: '#/components/schemas/Provider'
object:
type: string
data:
type: array
items:
$ref: '#/components/schemas/Model'
default: []
required:
- object
- data
ListToolsResponse:
type: object
description: Response structure for listing MCP tools
properties:
object:
type: string
description: Always "list"
example: 'list'
data:
type: array
items:
$ref: '#/components/schemas/MCPTool'
default: []
description: Array of available MCP tools
required:
- object
- data
MCPTool:
type: object
description: An MCP tool definition
properties:
name:
type: string
description: The name of the tool
example: 'read_file'
description:
type: string
description: A description of what the tool does
example: 'Read content from a file'
server:
type: string
description: The MCP server that provides this tool
example: 'http://mcp-filesystem-server:8083/mcp'
input_schema:
type: object
description: JSON schema for the tool's input parameters
example:
type: 'object'
properties:
file_path:
type: 'string'
description: 'Path to the file to read'
required:
- file_path
additionalProperties: true
required:
- name
- description
- server
FunctionObject:
type: object
properties:
description:
type: string
description:
A description of what the function does, used by the model to
choose when and how to call the function.
name:
type: string
description:
The name of the function to be called. Must be a-z, A-Z, 0-9, or
contain underscores and dashes, with a maximum length of 64.
parameters:
$ref: '#/components/schemas/FunctionParameters'
strict:
type: boolean
default: false
description:
Whether to enable strict schema adherence when generating the
function call. If set to true, the model will follow the exact
schema defined in the `parameters` field. Only a subset of JSON
Schema is supported when `strict` is `true`. Learn more about
Structured Outputs in the [function calling
guide](docs/guides/function-calling).
required:
- name
ChatCompletionTool:
type: object
properties:
type:
$ref: '#/components/schemas/ChatCompletionToolType'
function:
$ref: '#/components/schemas/FunctionObject'
required:
- type
- function
FunctionParameters:
type: object
description: >-
The parameters the functions accepts, described as a JSON Schema object.
See the [guide](/docs/guides/function-calling) for examples, and the
[JSON Schema
reference](https://json-schema.org/understanding-json-schema/) for
documentation about the format.
Omitting `parameters` defines a function with an empty parameter list.
additionalProperties: true
ChatCompletionToolType:
type: string
description: The type of the tool. Currently, only `function` is supported.
enum:
- function
CompletionUsage:
type: object
description: Usage statistics for the completion request.
properties:
completion_tokens:
type: integer
default: 0
format: int64
description: Number of tokens in the generated completion.
prompt_tokens:
type: integer
default: 0
format: int64
description: Number of tokens in the prompt.
total_tokens:
type: integer
default: 0
format: int64
description: Total number of tokens used in the request (prompt + completion).
required:
- prompt_tokens
- completion_tokens
- total_tokens
ChatCompletionStreamOptions:
description: >
Options for streaming response. Only set this when you set `stream:
true`.
type: object
properties:
include_usage:
type: boolean
description: >
If set, an additional chunk will be streamed before the `data:
[DONE]` message. The `usage` field on this chunk shows the token
usage statistics for the entire request, and the `choices` field
will always be an empty array. All other chunks will also include a
`usage` field, but with a null value.
required:
- include_usage
CreateChatCompletionRequest:
type: object
properties:
model:
type: string
description: Model ID to use
messages:
description: >
A list of messages comprising the conversation so far.
type: array
minItems: 1
items:
$ref: '#/components/schemas/Message'
max_tokens:
description: >
The maximum number of tokens that can be generated in the chat
completion. This value can be used to control costs for text
generated via API. This value is now deprecated in favor of
`max_completion_tokens`, and is not compatible with o-series models.
type: integer
deprecated: true
max_completion_tokens:
description: >
An upper bound for the number of tokens that can be generated
for a completion, including visible output tokens and reasoning tokens.
type: integer
temperature:
description: >
What sampling temperature to use, between 0 and 2. Higher values
like 0.8 will make the output more random, while lower values
like 0.2 will make it more focused and deterministic.
type: number
minimum: 0
maximum: 2
default: 1
top_p:
description: >
An alternative to sampling with temperature, called nucleus
sampling, where the model considers the results of the tokens
with top_p probability mass.
type: number
minimum: 0
maximum: 1
default: 1
frequency_penalty:
description: >
Number between -2.0 and 2.0. Positive values penalize new tokens
based on their existing frequency in the text so far, decreasing
the model's likelihood to repeat the same line verbatim.
type: number
minimum: -2
maximum: 2
default: 0
presence_penalty:
description: >
Number between -2.0 and 2.0. Positive values penalize new tokens
based on whether they appear in the text so far, increasing the
model's likelihood to talk about new topics.
type: number
minimum: -2
maximum: 2
default: 0
n:
description: >
How many chat completion choices to generate for each input message.
type: integer
minimum: 1
maximum: 128
default: 1
stop:
description: >
Up to 4 sequences where the API will stop generating further tokens.
oneOf:
- type: string
- type: array
minItems: 1
maxItems: 4
items:
type: string
seed:
description: >
If specified, our system will make a best effort to sample
deterministically, such that repeated requests with the same `seed`
and parameters should return the same result. Determinism is not
guaranteed, and you should refer to the `system_fingerprint`
response parameter to monitor changes in the backend.
type: integer
logprobs:
description: >
Whether to return log probabilities of the output tokens or not. If
true, returns the log probabilities of each output token returned in
the `content` of `message`.
type: boolean
default: false
top_logprobs:
description: >
An integer between 0 and 20 specifying the number of most likely
tokens to return at each token position, each with an associated log
probability. `logprobs` must be set to `true` if this parameter is
used.
type: integer
minimum: 0
maximum: 20
response_format:
description: >
An object specifying the format that the model must output. Setting
to `{ "type": "json_schema", "json_schema": {...} }` enables
Structured Outputs which guarantees the model will match your
supplied JSON schema. Setting to `{ "type": "json_object" }` enables
the older JSON mode, which ensures the message the model generates is
valid JSON.
oneOf:
- $ref: '#/components/schemas/ResponseFormatText'
- $ref: '#/components/schemas/ResponseFormatJsonSchema'
- $ref: '#/components/schemas/ResponseFormatJsonObject'
logit_bias:
description: >
Modify the likelihood of specified tokens appearing in the
completion. Accepts a JSON object that maps tokens (specified by
their token ID in the tokenizer) to an associated bias value from
-100 to 100. The bias is added to the logits generated by the model
prior to sampling.
type: object
additionalProperties:
type: integer
user:
description: >
A unique identifier representing your end-user, which can help to
monitor and detect abuse.
type: string
stream:
description: >
If set to true, the model response data will be streamed to the
client as it is generated using [server-sent
events](https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events#Event_stream_format).
type: boolean
default: false
stream_options:
$ref: '#/components/schemas/ChatCompletionStreamOptions'
tools:
type: array
description: >
A list of tools the model may call. Currently, only functions
are supported as a tool. Use this to provide a list of functions
the model may generate JSON inputs for. A max of 128 functions
are supported.
items:
$ref: '#/components/schemas/ChatCompletionTool'
tool_choice:
$ref: '#/components/schemas/ChatCompletionToolChoiceOption'
parallel_tool_calls:
type: boolean
default: true
description: >
Whether to enable parallel function calling during tool use.
reasoning_format:
type: string
description: >
The format of the reasoning content. Can be `raw` or `parsed`.
When specified as raw some reasoning models will output tags.
When specified as parsed the model will output the reasoning under
`reasoning` or `reasoning_content` attribute.
reasoning_effort:
type: string
description: >
Constrains effort on reasoning for reasoning models. Currently
supported values are `minimal`, `low`, `medium`, and `high`.
Reducing reasoning effort can result in faster responses and fewer
tokens used on reasoning in a response.
x-enum-varnames:
- Minimal
- Low
- Medium
- High
enum:
- minimal
- low
- medium
- high
required:
- model
- messages
ResponseFormatText:
type: object
description: Default response format. Used to generate text responses.
properties:
type:
type: string
description: The type of response format being defined. Always `text`.
enum:
- text
required:
- type
ResponseFormatJsonObject:
type: object
description: >
JSON object response format. An older method of generating JSON
responses. Using `json_schema` is recommended for models that support
it. Note that the model will not generate JSON without a system or user
message instructing it to do so.
properties:
type:
type: string
description: The type of response format being defined. Always `json_object`.
enum:
- json_object
required:
- type
ResponseFormatJsonSchema:
type: object
description: >
JSON Schema response format. Used to generate structured JSON responses.
properties:
type:
type: string
description: The type of response format being defined. Always `json_schema`.
enum:
- json_schema
json_schema:
type: object
description: Structured Outputs configuration options, including a JSON Schema.
properties:
description:
type: string
description: >
A description of what the response format is for, used by the
model to determine how to respond in the format.
name:
type: string
description: >
The name of the response format. Must be a-z, A-Z, 0-9, or
contain underscores and dashes, with a maximum length of 64.
schema:
$ref: '#/components/schemas/ResponseFormatJsonSchemaSchema'
strict:
type: boolean
default: false
description: >
Whether to enable strict schema adherence when generating the
output. If set to true, the model will always follow the exact
schema defined in the `schema` field. Only a subset of JSON
Schema is supported when `strict` is `true`.
required:
- name
required:
- type
- json_schema
ResponseFormatJsonSchemaSchema:
type: object
description: >
The schema for the response format, described as a JSON Schema object.
additionalProperties: true
ChatCompletionToolChoiceOption:
description: >
Controls which (if any) tool is called by the model. `none` means the
model will not call any tool and instead generates a message. `auto`
means the model can pick between generating a message or calling one or
more tools. `required` means the model must call one or more tools.
Specifying a particular tool via `{"type": "function", "function":
{"name": "my_function"}}` forces the model to call that tool.
`none` is the default when no tools are present. `auto` is the default
if tools are present.
oneOf:
- type: string
description: >
`none` means the model will not call any tool and instead generates
a message. `auto` means the model can pick between generating a
message or calling one or more tools. `required` means the model
must call one or more tools.
enum:
- none
- auto
- required
- $ref: '#/components/schemas/ChatCompletionNamedToolChoice'
ChatCompletionNamedToolChoice:
type: object
description: >
Specifies a tool the model should use. Use to force the model to call a
specific function.
properties:
type:
$ref: '#/components/schemas/ChatCompletionToolType'
function:
type: object
properties:
name:
type: string
description: The name of the function to call.
required:
- name
required:
- type
- function
ChatCompletionMessageToolCallFunction:
type: object
description: The function that the model called.
properties:
name:
type: string
description: The name of the function to call.
arguments:
type: string
description:
The arguments to call the function with, as generated by the model
in JSON format. Note that the model does not always generate
valid JSON, and may hallucinate parameters not defined by your
function schema. Validate the arguments in your code before
calling your function.
required:
- name
- arguments
ChatCompletionMessageToolCall:
type: object
properties:
id:
type: string
description: The ID of the tool call.
type:
$ref: '#/components/schemas/ChatCompletionToolType'
function:
$ref: '#/components/schemas/ChatCompletionMessageToolCallFunction'
extra_content:
$ref: '#/components/schemas/ToolCallExtraContent'
required:
- id
- type
- function
ChatCompletionChoice:
type: object
properties:
finish_reason:
$ref: '#/components/schemas/FinishReason'
index:
type: integer
description: The index of the choice in the list of choices.
message:
$ref: '#/components/schemas/Message'
logprobs:
description: Log probability information for the choice.
type: object
nullable: true
properties:
content:
description: A list of message content tokens with log probability information.
type: array
items:
$ref: '#/components/schemas/ChatCompletionTokenLogprob'
refusal:
description: A list of message refusal tokens with log probability information.
type: array
items:
$ref: '#/components/schemas/ChatCompletionTokenLogprob'
required:
- content
- refusal
required:
- finish_reason
- index
- message
ChatCompletionStreamChoice:
type: object
required:
- delta
- finish_reason
- index
properties:
delta:
$ref: '#/components/schemas/ChatCompletionStreamResponseDelta'
logprobs:
description: Log probability information for the choice.
type: object
properties:
content:
description: A list of message content tokens with log probability information.
type: array
items:
$ref: '#/components/schemas/ChatCompletionTokenLogprob'
refusal:
description: A list of message refusal tokens with log probability information.
type: array
items:
$ref: '#/components/schemas/ChatCompletionTokenLogprob'
required:
- content
- refusal
finish_reason:
$ref: '#/components/schemas/FinishReason'
index:
type: integer
description: The index of the choice in the list of choices.
CreateChatCompletionResponse:
type: object
description:
Represents a chat completion response returned by model, based on
the provided input.
properties:
id:
type: string
description: A unique identifier for the chat completion.
choices:
type: array
description:
A list of chat completion choices. Can be more than one if `n` is
greater than 1.
items:
$ref: '#/components/schemas/ChatCompletionChoice'
created:
type: integer
description:
The Unix timestamp (in seconds) of when the chat completion was
created.
model:
type: string
description: The model used for the chat completion.
object:
type: string
description: The object type, which is always `chat.completion`.
x-stainless-const: true
usage:
$ref: '#/components/schemas/CompletionUsage'
required:
- choices
- created
- id
- model
- object
ChatCompletionStreamResponseDelta:
type: object
description: A chat completion delta generated by streamed model responses.
properties:
content:
type: string
description: The contents of the chunk message.
reasoning_content:
type: string
description: The reasoning content of the chunk message.
reasoning:
type: string
description: The reasoning of the chunk message. Same as reasoning_content.
tool_calls:
type: array
items:
$ref: '#/components/schemas/ChatCompletionMessageToolCallChunk'
role:
$ref: '#/components/schemas/MessageRole'
refusal:
type: string
description: The refusal message generated by the model.
required:
- content
- role
ChatCompletionMessageToolCallChunk:
type: object
properties:
index:
type: integer
id:
type: string
description: The ID of the tool call.
type:
type: string
description: The type of the tool. Currently, only `function` is supported.
function:
$ref: '#/components/schemas/ChatCompletionMessageToolCallFunction'
extra_content:
$ref: '#/components/schemas/ToolCallExtraContent'
required:
- index
ToolCallExtraContent:
type: object
description: |
Provider-specific opaque data attached to a tool call. The contents are
not interpreted by the gateway, but must be echoed back verbatim on the
next request that references this tool call. Currently used by Google
Gemini extended-thinking models to carry the per-call `thought_signature`.
Other providers may ignore the field.
properties:
google:
type: object
description: Google Gemini-specific extra content.
properties:
thought_signature:
type: string
description: |
Opaque signature returned with reasoning-enabled tool calls.
Must be echoed back verbatim in the next request that includes
this tool call, or Google will reject the request.
additionalProperties: true
ChatCompletionTokenLogprob:
type: object
properties:
token: &a1
description: The token.
type: string
logprob: &a2
description:
The log probability of this token, if it is within the top 20 most
likely tokens. Otherwise, the value `-9999.0` is used to signify
that the token is very unlikely.
type: number
bytes: &a3
description:
A list of integers representing the UTF-8 bytes representation of
the token. Useful in instances where characters are represented by
multiple tokens and their byte representations must be combined to
generate the correct text representation. Can be `null` if there is
no bytes representation for the token.
type: array
items:
type: integer
top_logprobs:
description:
List of the most likely tokens and their log probability, at this
token position. In rare cases, there may be fewer than the number of
requested `top_logprobs` returned.
type: array
items:
type: object
properties:
token: *a1
logprob: *a2
bytes: *a3
required:
- token
- logprob
- bytes
required:
- token
- logprob
- bytes
- top_logprobs
FinishReason:
type: string
description: >
The reason the model stopped generating tokens. This will be
`stop` if the model hit a natural stop point or a provided
stop sequence,
`length` if the maximum number of tokens specified in the
request was reached,
`content_filter` if content was omitted due to a flag from our
content filters,
`tool_calls` if the model called a tool.
enum:
- stop
- length
- tool_calls
- content_filter
- function_call
x-enum-varnames:
- Stop
- Length
- ToolCalls
- ContentFilter
- FunctionCall
CreateChatCompletionStreamResponse:
type: object
description: |
Represents a streamed chunk of a chat completion response returned
by the model, based on the provided input.
properties:
id:
type: string
description:
A unique identifier for the chat completion. Each chunk has the
same ID.
choices:
type: array
description: >
A list of chat completion choices. Can contain more than one
elements if `n` is greater than 1. Can also be empty for the
last chunk if you set `stream_options: {"include_usage": true}`.
items:
$ref: '#/components/schemas/ChatCompletionStreamChoice'
created:
type: integer
description:
The Unix timestamp (in seconds) of when the chat completion was
created. Each chunk has the same timestamp.
model:
type: string
description: The model to generate the completion.
system_fingerprint:
type: string
description: >
This fingerprint represents the backend configuration that the model
runs with.
Can be used in conjunction with the `seed` request parameter to
understand when backend changes have been made that might impact
determinism.
object:
type: string
description: The object type, which is always `chat.completion.chunk`.
usage:
$ref: '#/components/schemas/CompletionUsage'
reasoning_format:
type: string
description: >
The format of the reasoning content. Can be `raw` or `parsed`.
When specified as raw some reasoning models will output tags.
When specified as parsed the model will output the reasoning under reasoning_content.
required:
- choices
- created
- id
- model
- object
CreateResponseRequest:
type: object
description: |
Request body for creating a model response via the Responses API.
properties:
model:
type: string
description: Model ID used to generate the response.
input:
$ref: '#/components/schemas/ResponseInput'
instructions:
type: string
nullable: true
description: >
A system (or developer) message inserted into the model's context.
When used with `previous_response_id`, instructions from previous
responses are not carried over.
max_output_tokens:
type: integer
nullable: true
description: >
An upper bound for the number of tokens that can be generated for a
response, including visible output tokens and reasoning tokens.
stream:
type: boolean
default: false
description: >
If set to true, the model response data is streamed to the client
as it is generated using
[server-sent events](https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events#Event_stream_format).
temperature:
type: number
format: float
nullable: true
default: 1
description: >
What sampling temperature to use, between 0 and 2. Higher values
make the output more random; lower values make it more focused.
top_p:
type: number
format: float
nullable: true
default: 1
description: >
An alternative to sampling with temperature, called nucleus
sampling, where the model considers the tokens with `top_p`
probability mass.
tools:
type: array
description: >
An array of tools the model may call while generating a response.
items:
$ref: '#/components/schemas/ResponseTool'
tool_choice:
$ref: '#/components/schemas/ResponseToolChoice'
reasoning:
$ref: '#/components/schemas/ResponseReasoning'
text:
$ref: '#/components/schemas/ResponseTextConfig'
previous_response_id:
type: string
nullable: true
description: >
The unique ID of the previous response to the model. Use this to
create multi-turn conversations.
store:
type: boolean
default: true
description: >
Whether to store the generated model response for later retrieval.
background:
type: boolean
default: false
description: >
Whether to run the model response in the background. Useful for
long-running or batched requests.
parallel_tool_calls:
type: boolean
default: true
description: Whether to allow the model to run tool calls in parallel.
metadata:
type: object
additionalProperties:
type: string
description: >
Set of up to 16 key-value pairs that can be attached to the object
and returned when retrieving the response.
user:
type: string
description: >
A stable identifier for your end-users, used to help detect and
prevent abuse.
required:
- model
- input
ResponseInput:
description: >
Text, image, or file inputs to the model. Either a single text prompt
or a list of input items representing a (possibly batched) conversation.
oneOf:
- type: string
description: A text input to the model, equivalent to a user message.
- type: array
description: A list of input items.
items:
$ref: '#/components/schemas/ResponseInputItem'
ResponseInputItem:
type: object
description: >
A single input item. Most commonly an input message with a role and
content.
properties:
type:
type: string
default: message
description: The type of the input item. Defaults to `message`.
role:
$ref: '#/components/schemas/ResponseRole'
content:
$ref: '#/components/schemas/ResponseInputMessageContent'
required:
- role
- content
ResponseRole:
type: string
description: The role of the message input.
enum:
- user
- assistant
- system
- developer
x-enum-varnames:
- ResponseRoleUser
- ResponseRoleAssistant
- ResponseRoleSystem
- ResponseRoleDeveloper
ResponseInputMessageContent:
description: >
Text or multimodal content for an input message. Either a string or a
list of content parts.
oneOf:
- type: string
description: A text input to the model.
- type: array
items:
$ref: '#/components/schemas/ResponseInputContentPart'
ResponseInputContentPart:
type: object
description: A content part within an input message.
oneOf:
- $ref: '#/components/schemas/ResponseInputText'
- $ref: '#/components/schemas/ResponseInputImage'
ResponseInputText:
type: object
description: A text input to the model.
properties:
type:
type: string
enum:
- input_text
description: The type of the input item. Always `input_text`.
text:
type: string
description: The text input to the model.
required:
- type
- text
ResponseInputImage:
type: object
description: An image input to the model.
properties:
type:
type: string
enum:
- input_image
description: The type of the input item. Always `input_image`.
image_url:
type: string
description: The URL of the image (data URLs supported).
detail:
type: string
enum:
- auto
- low
- high
x-enum-varnames:
- ResponseInputImageDetailAuto
- ResponseInputImageDetailLow
- ResponseInputImageDetailHigh
default: auto
description: The detail level of the image to send to the model.
required:
- type
ResponseTool:
type: object
description: >
A tool the model may call. Only function tools are modeled here. Note
the Responses API uses a flattened function tool shape (`name`,
`description`, and `parameters` at the top level) rather than nesting
them under a `function` object as `/chat/completions` does.
properties:
type:
type: string
enum:
- function
x-enum-varnames:
- ResponseToolTypeFunction
description: The type of the tool. Currently only `function`.
name:
type: string
description: The name of the function to call.
description:
type: string
description: >
A description of the function, used by the model to decide when and
how to call it.
parameters:
$ref: '#/components/schemas/FunctionParameters'
strict:
type: boolean
default: false
description: Whether to enforce strict parameter validation.
required:
- type
- name
ResponseToolChoice:
description: >
How the model should select which tool (or tools) to use. Either a mode
string (`none`, `auto`, `required`) or an object forcing a specific
tool.
oneOf:
- type: string
enum:
- none
- auto
- required
description: The tool-choice mode.
- type: object
description: Forces the model to call a specific function tool.
properties:
type:
type: string
enum:
- function
x-enum-varnames:
- ResponseToolChoiceTypeFunction
name:
type: string
required:
- type
- name
ResponseReasoning:
type: object
description: Configuration options for reasoning models.
properties:
effort:
type: string
enum:
- minimal
- low
- medium
- high
default: medium
nullable: true
x-enum-varnames:
- ResponseReasoningEffortMinimal
- ResponseReasoningEffortLow
- ResponseReasoningEffortMedium
- ResponseReasoningEffortHigh
description: >
Constrains the effort on reasoning for reasoning models. Reducing
effort can result in faster responses and fewer reasoning tokens.
summary:
type: string
enum:
- auto
- concise
- detailed
nullable: true
description: >
A summary of the reasoning performed by the model, useful for
debugging and understanding the model's reasoning process.
ResponseTextConfig:
type: object
description: >
Configuration options for a text response from the model. Can be plain
text or structured JSON data.
properties:
format:
type: object
description: An object specifying the format that the model must output.
properties:
type:
type: string
enum:
- text
- json_schema
- json_object
x-enum-varnames:
- ResponseTextConfigFormatTypeText
- ResponseTextConfigFormatTypeJSONSchema
- ResponseTextConfigFormatTypeJSONObject
description: The type of response format being defined.
name:
type: string
description: The name of the response format (used with `json_schema`).
schema:
$ref: '#/components/schemas/FunctionParameters'
strict:
type: boolean
default: false
description: Whether to enable strict schema adherence.
required:
- type
Response:
type: object
description: Represents a model response returned by the Responses API.
properties:
id:
type: string
description: Unique identifier for this response.
object:
type: string
description: The object type, which is always `response`.
created_at:
type: integer
format: int64
description: Unix timestamp (in seconds) of when the response was created.
status:
$ref: '#/components/schemas/ResponseStatus'
model:
type: string
description: The model used to generate the response.
output:
type: array
description: An array of content items generated by the model.
items:
$ref: '#/components/schemas/ResponseOutputItem'
error:
$ref: '#/components/schemas/ResponseError'
incomplete_details:
$ref: '#/components/schemas/ResponseIncompleteDetails'
instructions:
type: string
nullable: true
description: The system/developer message used to generate the response.
max_output_tokens:
type: integer
nullable: true
description: An upper bound for the number of generated tokens.
previous_response_id:
type: string
nullable: true
description: The unique ID of the previous response, if any.
reasoning:
$ref: '#/components/schemas/ResponseReasoning'
temperature:
type: number
format: float
nullable: true
top_p:
type: number
format: float
nullable: true
tool_choice:
$ref: '#/components/schemas/ResponseToolChoice'
tools:
type: array
items:
$ref: '#/components/schemas/ResponseTool'
text:
$ref: '#/components/schemas/ResponseTextConfig'
metadata:
type: object
additionalProperties:
type: string
usage:
$ref: '#/components/schemas/ResponseUsage'
required:
- id
- object
- created_at
- status
- model
- output
ResponseStatus:
type: string
description: The status of the response generation.
enum:
- completed
- failed
- in_progress
- cancelled
- queued
- incomplete
ResponseError:
type: object
nullable: true
description: An error object returned when the model fails to generate a response.
properties:
code:
type: string
description: The error code for the response.
message:
type: string
description: A human-readable description of the error.
required:
- code
- message
ResponseIncompleteDetails:
type: object
nullable: true
description: Details about why the response is incomplete.
properties:
reason:
type: string
description: The reason why the response is incomplete.
ResponseOutputItem:
type: object
description: >
An output item generated by the model: an output message, a function
tool call, or a reasoning item.
oneOf:
- $ref: '#/components/schemas/ResponseOutputMessage'
- $ref: '#/components/schemas/ResponseFunctionToolCall'
- $ref: '#/components/schemas/ResponseReasoningItem'
ResponseOutputMessage:
type: object
description: An output message from the model.
properties:
type:
type: string
enum:
- message
description: The type of the output item. Always `message`.
id:
type: string
description: The unique ID of the output message.
role:
type: string
enum:
- assistant
x-enum-varnames:
- ResponseOutputMessageRoleAssistant
description: The role of the output message. Always `assistant`.
status:
type: string
enum:
- in_progress
- completed
- incomplete
description: The status of the message.
content:
type: array
items:
$ref: '#/components/schemas/ResponseOutputContent'
required:
- type
- id
- role
- content
ResponseOutputContent:
type: object
description: A content part of an output message.
oneOf:
- $ref: '#/components/schemas/ResponseOutputText'
- $ref: '#/components/schemas/ResponseOutputRefusal'
ResponseOutputText:
type: object
description: A text output from the model.
properties:
type:
type: string
enum:
- output_text
description: The type of the output text. Always `output_text`.
text:
type: string
description: The text output from the model.
required:
- type
- text
ResponseOutputRefusal:
type: object
description: A refusal generated by the model.
properties:
type:
type: string
enum:
- refusal
description: The type of the refusal. Always `refusal`.
refusal:
type: string
description: The refusal explanation from the model.
required:
- type
- refusal
ResponseFunctionToolCall:
type: object
description: A tool call to a function generated by the model.
properties:
type:
type: string
enum:
- function_call
x-enum-varnames:
- ResponseFunctionToolCallTypeFunctionCall
description: The type of the output item. Always `function_call`.
id:
type: string
description: The unique ID of the function tool call.
call_id:
type: string
description: >
The unique ID of the function tool call generated by the model,
used to associate the call with its output.
name:
type: string
description: The name of the function to run.
arguments:
type: string
description: A JSON string of the arguments to pass to the function.
status:
type: string
enum:
- in_progress
- completed
- incomplete
description: The status of the function tool call.
required:
- type
- call_id
- name
- arguments
ResponseReasoningItem:
type: object
description: A reasoning item describing the model's chain of thought.
properties:
type:
type: string
enum:
- reasoning
description: The type of the output item. Always `reasoning`.
id:
type: string
description: The unique ID of the reasoning item.
summary:
type: array
description: Reasoning summary content.
items:
$ref: '#/components/schemas/ResponseReasoningSummaryPart'
status:
type: string
enum:
- in_progress
- completed
- incomplete
description: The status of the reasoning item.
required:
- type
- id
- summary
ResponseReasoningSummaryPart:
type: object
description: A summary part of a reasoning item.
properties:
type:
type: string
enum:
- summary_text
description: The type of the summary. Always `summary_text`.
text:
type: string
description: A summary of the reasoning output from the model.
required:
- type
- text
ResponseUsage:
type: object
description: Token usage details for the response.
properties:
input_tokens:
type: integer
format: int64
default: 0
description: The number of input tokens.
input_tokens_details:
type: object
description: A detailed breakdown of the input tokens.
properties:
cached_tokens:
type: integer
format: int64
default: 0
description: The number of tokens retrieved from the cache.
output_tokens:
type: integer
format: int64
default: 0
description: The number of output tokens.
output_tokens_details:
type: object
description: A detailed breakdown of the output tokens.
properties:
reasoning_tokens:
type: integer
format: int64
default: 0
description: The number of reasoning tokens.
total_tokens:
type: integer
format: int64
default: 0
description: The total number of tokens used (input + output).
required:
- input_tokens
- output_tokens
- total_tokens
ResponseStreamEvent:
type: object
description: >
A server-sent event emitted while streaming a response. The Responses
API emits a sequence of typed events (for example `response.created`,
`response.output_text.delta`, and `response.completed`). This schema
models the common event envelope; which fields are populated depends on
the event `type`.
properties:
type:
type: string
description: >
The type of the streamed event, for example
`response.output_text.delta` or `response.completed`.
sequence_number:
type: integer
description: The sequence number of this event.
response:
$ref: '#/components/schemas/Response'
item_id:
type: string
description: The ID of the output item this event relates to.
output_index:
type: integer
description: The index of the output item in the response's output array.
content_index:
type: integer
description: The index of the content part within the output item.
delta:
type: string
description: The incremental text delta for `*.delta` events.
text:
type: string
description: The finalized text for `*.done` events.
required:
- type
Config:
x-config:
sections:
- general:
title: 'General settings'
settings:
- name: environment
env: 'ENVIRONMENT'
type: string
default: 'production'
description: 'The environment'
- name: allowed_models
env: 'ALLOWED_MODELS'
type: string
default: ''
description: 'Comma-separated list of models to allow. If empty, all models will be available'
- name: disallowed_models
env: 'DISALLOWED_MODELS'
type: string
default: ''
description: 'Comma-separated list of models to disallow. If empty, no models will be blocked. Takes lower precedence than ALLOWED_MODELS'
- name: enable_vision
env: 'ENABLE_VISION'
type: bool
default: 'false'
description: 'Enable vision/multimodal support for all providers. When disabled, image inputs will be rejected even if the provider and model support vision'
- name: debug_content_truncate_words
env: 'DEBUG_CONTENT_TRUNCATE_WORDS'
type: int
default: '10'
description: 'Number of words to truncate per content section in debug logs (development mode only)'
- name: debug_max_messages
env: 'DEBUG_MAX_MESSAGES'
type: int
default: '100'
description: 'Maximum number of messages to show in debug logs (development mode only)'
- telemetry:
title: 'Telemetry'
settings:
- name: telemetry_enable
env: 'TELEMETRY_ENABLE'
type: bool
default: 'false'
description: 'Enable telemetry'
- name: telemetry_metrics_push_enable
env: 'TELEMETRY_METRICS_PUSH_ENABLE'
type: bool
default: 'false'
description: 'Enable the OTLP metrics push endpoint (POST /v1/metrics)'
- name: telemetry_metrics_port
env: 'TELEMETRY_METRICS_PORT'
type: string
default: '9464'
description: 'Port for telemetry metrics server'
- mcp:
title: 'Model Context Protocol (MCP)'
settings:
- name: mcp_enable
env: 'MCP_ENABLE'
type: bool
default: 'false'
description: 'Enable MCP'
- name: mcp_expose
env: 'MCP_EXPOSE'
type: bool
default: 'false'
description: 'Expose MCP tools endpoint'
- name: mcp_servers
env: 'MCP_SERVERS'
type: string
description: 'List of MCP servers'
- name: mcp_include_tools
env: 'MCP_INCLUDE_TOOLS'
type: string
description: 'Comma-separated list of MCP tool names to inject. If empty, all tools are injected. Takes precedence over MCP_EXCLUDE_TOOLS'
- name: mcp_exclude_tools
env: 'MCP_EXCLUDE_TOOLS'
type: string
description: 'Comma-separated list of MCP tool names to skip injecting. If empty, no tools are excluded. Takes lower precedence than MCP_INCLUDE_TOOLS'
- name: mcp_client_timeout
env: 'MCP_CLIENT_TIMEOUT'
type: time.Duration
default: '5s'
description: 'MCP client HTTP timeout'
- name: mcp_dial_timeout
env: 'MCP_DIAL_TIMEOUT'
type: time.Duration
default: '3s'
description: 'MCP client dial timeout'
- name: mcp_tls_handshake_timeout
env: 'MCP_TLS_HANDSHAKE_TIMEOUT'
type: time.Duration
default: '3s'
description: 'MCP client TLS handshake timeout'
- name: mcp_response_header_timeout
env: 'MCP_RESPONSE_HEADER_TIMEOUT'
type: time.Duration
default: '3s'
description: 'MCP client response header timeout'
- name: mcp_expect_continue_timeout
env: 'MCP_EXPECT_CONTINUE_TIMEOUT'
type: time.Duration
default: '1s'
description: 'MCP client expect continue timeout'
- name: mcp_request_timeout
env: 'MCP_REQUEST_TIMEOUT'
type: time.Duration
default: '5s'
description: 'MCP client request timeout for initialize and tool calls'
- name: mcp_max_retries
env: 'MCP_MAX_RETRIES'
type: int
default: '3'
description: 'Maximum number of connection retry attempts'
- name: mcp_retry_interval
env: 'MCP_RETRY_INTERVAL'
type: time.Duration
default: '5s'
description: 'Interval between connection retry attempts'
- name: mcp_initial_backoff
env: 'MCP_INITIAL_BACKOFF'
type: time.Duration
default: '1s'
description: 'Initial backoff duration for exponential backoff retry'
- name: mcp_enable_reconnect
env: 'MCP_ENABLE_RECONNECT'
type: bool
default: 'true'
description: 'Enable automatic reconnection for failed servers'
- name: mcp_reconnect_interval
env: 'MCP_RECONNECT_INTERVAL'
type: time.Duration
default: '30s'
description: 'Interval between reconnection attempts'
- name: mcp_polling_enable
env: 'MCP_POLLING_ENABLE'
type: bool
default: 'true'
description: 'Enable health check polling'
- name: mcp_polling_interval
env: 'MCP_POLLING_INTERVAL'
type: time.Duration
default: '30s'
description: 'Interval between health check polling requests'
- name: mcp_polling_timeout
env: 'MCP_POLLING_TIMEOUT'
type: time.Duration
default: '5s'
description: 'Timeout for individual health check requests'
- name: mcp_disable_healthcheck_logs
env: 'MCP_DISABLE_HEALTHCHECK_LOGS'
type: bool
default: 'true'
description: 'Disable health check log messages to reduce noise'
- auth:
title: 'Authentication'
settings:
- name: auth_enable
env: 'AUTH_ENABLE'
type: bool
default: 'false'
description: 'Enable authentication'
- name: auth_oidc_issuer
env: 'AUTH_OIDC_ISSUER'
type: string
default: 'http://keycloak:8080/realms/inference-gateway-realm'
description: 'OIDC issuer URL'
- name: auth_oidc_client_id
env: 'AUTH_OIDC_CLIENT_ID'
type: string
default: 'inference-gateway-client'
description: 'OIDC client ID'
secret: true
- name: auth_oidc_client_secret
env: 'AUTH_OIDC_CLIENT_SECRET'
type: string
description: 'OIDC client secret'
secret: true
- server:
title: 'Server settings'
settings:
- name: host
env: 'SERVER_HOST'
type: string
default: '0.0.0.0'
description: 'Server host'
- name: port
env: 'SERVER_PORT'
type: string
default: '8080'
description: 'Server port'
- name: read_timeout
env: 'SERVER_READ_TIMEOUT'
type: time.Duration
default: '30s'
description: 'Read timeout'
- name: write_timeout
env: 'SERVER_WRITE_TIMEOUT'
type: time.Duration
default: '30s'
description: 'Write timeout'
- name: idle_timeout
env: 'SERVER_IDLE_TIMEOUT'
type: time.Duration
default: '120s'
description: 'Idle timeout'
- name: tls_cert_path
env: 'SERVER_TLS_CERT_PATH'
type: string
description: 'TLS certificate path'
- name: tls_key_path
env: 'SERVER_TLS_KEY_PATH'
type: string
description: 'TLS key path'
- client:
title: 'Client settings'
settings:
- name: timeout
env: 'CLIENT_TIMEOUT'
type: time.Duration
default: '30s'
description: 'Client timeout'
- name: max_idle_conns
env: 'CLIENT_MAX_IDLE_CONNS'
type: int
default: '20'
description: 'Maximum idle connections'
- name: max_idle_conns_per_host
env: 'CLIENT_MAX_IDLE_CONNS_PER_HOST'
type: int
default: '20'
description: 'Maximum idle connections per host'
- name: idle_conn_timeout
env: 'CLIENT_IDLE_CONN_TIMEOUT'
type: time.Duration
default: '30s'
description: 'Idle connection timeout'
- name: tls_min_version
env: 'CLIENT_TLS_MIN_VERSION'
type: string
default: 'TLS12'
description: 'Minimum TLS version'
- name: disable_compression
env: 'CLIENT_DISABLE_COMPRESSION'
type: bool
default: 'true'
description: 'Disable compression for faster streaming'
- name: response_header_timeout
env: 'CLIENT_RESPONSE_HEADER_TIMEOUT'
type: time.Duration
default: '10s'
description: 'Response header timeout'
- name: expect_continue_timeout
env: 'CLIENT_EXPECT_CONTINUE_TIMEOUT'
type: time.Duration
default: '1s'
description: 'Expect continue timeout'
- providers:
title: 'Providers'
settings:
- name: anthropic_api_url
env: 'ANTHROPIC_API_URL'
type: string
default: 'https://api.anthropic.com/v1'
description: 'Anthropic API URL'
- name: anthropic_api_key
env: 'ANTHROPIC_API_KEY'
type: string
description: 'Anthropic API Key'
secret: true
- name: cloudflare_api_url
env: 'CLOUDFLARE_API_URL'
type: string
default: 'https://api.cloudflare.com/client/v4/accounts/{ACCOUNT_ID}/ai'
description: 'Cloudflare API URL'
- name: cloudflare_api_key
env: 'CLOUDFLARE_API_KEY'
type: string
description: 'Cloudflare API Key'
secret: true
- name: cohere_api_url
env: 'COHERE_API_URL'
type: string
default: 'https://api.cohere.ai'
description: 'Cohere API URL'
- name: cohere_api_key
env: 'COHERE_API_KEY'
type: string
description: 'Cohere API Key'
secret: true
- name: groq_api_url
env: 'GROQ_API_URL'
type: string
default: 'https://api.groq.com/openai/v1'
description: 'Groq API URL'
- name: groq_api_key
env: 'GROQ_API_KEY'
type: string
description: 'Groq API Key'
secret: true
- name: llamacpp_api_url
env: 'LLAMACPP_API_URL'
type: string
default: 'http://llamacpp:8080/v1'
description: 'llama.cpp API URL'
- name: llamacpp_api_key
env: 'LLAMACPP_API_KEY'
type: string
description: 'llama.cpp API Key'
secret: true
- name: ollama_api_url
env: 'OLLAMA_API_URL'
type: string
default: 'http://ollama:8080/v1'
description: 'Ollama API URL'
- name: ollama_api_key
env: 'OLLAMA_API_KEY'
type: string
description: 'Ollama API Key'
secret: true
- name: ollama_cloud_api_url
env: 'OLLAMA_CLOUD_API_URL'
type: string
default: 'https://ollama.com/v1'
description: 'Ollama Cloud API URL'
- name: ollama_cloud_api_key
env: 'OLLAMA_CLOUD_API_KEY'
type: string
description: 'Ollama Cloud API Key'
secret: true
- name: openai_api_url
env: 'OPENAI_API_URL'
type: string
default: 'https://api.openai.com/v1'
description: 'OpenAI API URL'
- name: openai_api_key
env: 'OPENAI_API_KEY'
type: string
description: 'OpenAI API Key'
secret: true
- name: deepseek_api_url
env: 'DEEPSEEK_API_URL'
type: string
default: 'https://api.deepseek.com'
description: 'DeepSeek API URL'
- name: deepseek_api_key
env: 'DEEPSEEK_API_KEY'
type: string
description: 'DeepSeek API Key'
secret: true
- name: google_api_url
env: 'GOOGLE_API_URL'
type: string
default: 'https://generativelanguage.googleapis.com/v1beta/openai'
description: 'Google API URL'
- name: google_api_key
env: 'GOOGLE_API_KEY'
type: string
description: 'Google API Key'
secret: true
- name: mistral_api_url
env: 'MISTRAL_API_URL'
type: string
default: 'https://api.mistral.ai/v1'
description: 'Mistral API URL'
- name: mistral_api_key
env: 'MISTRAL_API_KEY'
type: string
description: 'Mistral API Key'
secret: true
- name: minimax_api_url
env: 'MINIMAX_API_URL'
type: string
default: 'https://api.minimax.io/v1'
description: 'MiniMax API URL'
- name: minimax_api_key
env: 'MINIMAX_API_KEY'
type: string
description: 'MiniMax API Key'
secret: true
- name: moonshot_api_url
env: 'MOONSHOT_API_URL'
type: string
default: 'https://api.moonshot.ai/v1'
description: 'Moonshot API URL'
- name: moonshot_api_key
env: 'MOONSHOT_API_KEY'
type: string
description: 'Moonshot API Key'
secret: true
- name: nvidia_api_url
env: 'NVIDIA_API_URL'
type: string
default: 'https://integrate.api.nvidia.com/v1'
description: 'NVIDIA API URL'
- name: nvidia_api_key
env: 'NVIDIA_API_KEY'
type: string
description: 'NVIDIA API Key'
secret: true
- name: zai_api_url
env: 'ZAI_API_URL'
type: string
default: 'https://api.z.ai/v1'
description: 'ZAI API URL'
- name: zai_api_key
env: 'ZAI_API_KEY'
type: string
description: 'ZAI API Key'
secret: true