--- name: omni-inference description: "The core OpenAI-compatible inference endpoints: chat completions, embeddings, images, audio (TTS/STT), moderations, rerank, and the Responses API. The primary integration surface for AI agents." --- ## Overview The core OpenAI-compatible inference endpoints: chat completions, embeddings, images, audio (TTS/STT), moderations, rerank, and the Responses API. The primary integration surface for AI agents. ## Authentication All requests require a valid Bearer token or session cookie. Obtain a token via `POST /api/auth/login` or configure `REQUIRE_API_KEY=false` for local development. ## Endpoints - [`POST /api/v1/session-leases`](references/endpoints.md#post-apiv1session-leases) - [`GET /api/v1/search`](references/endpoints.md#get-apiv1search) - [`POST /api/v1/search`](references/endpoints.md#post-apiv1search) - [`POST /api/v1/chat/completions`](references/endpoints.md#post-apiv1chatcompletions) - [`GET /api/v1/ws`](references/endpoints.md#get-apiv1ws) - [`POST /api/v1/providers/{provider}/chat/completions`](references/endpoints.md#post-apiv1providersproviderchatcompletions) - [`POST /api/v1/api/chat`](references/endpoints.md#post-apiv1apichat) - [`POST /api/v1/messages`](references/endpoints.md#post-apiv1messages) - [`POST /api/v1/messages/count_tokens`](references/endpoints.md#post-apiv1messagescounttokens) - [`POST /api/v1/responses`](references/endpoints.md#post-apiv1responses) - [`POST /api/v1/embeddings`](references/endpoints.md#post-apiv1embeddings) - [`GET /api/v1/multimodal-embeddings`](references/endpoints.md#get-apiv1multimodal-embeddings) - [`POST /api/v1/multimodal-embeddings`](references/endpoints.md#post-apiv1multimodal-embeddings) - [`POST /api/v1/providers/{provider}/embeddings`](references/endpoints.md#post-apiv1providersproviderembeddings) - [`POST /api/v1/images/generations`](references/endpoints.md#post-apiv1imagesgenerations) - [`POST /api/v1/providers/{provider}/images/generations`](references/endpoints.md#post-apiv1providersproviderimagesgenerations) - [`POST /api/v1/audio/speech`](references/endpoints.md#post-apiv1audiospeech) - [`POST /api/v1/audio/transcriptions`](references/endpoints.md#post-apiv1audiotranscriptions) - [`POST /api/v1/moderations`](references/endpoints.md#post-apiv1moderations) - [`POST /api/v1/rerank`](references/endpoints.md#post-apiv1rerank) - [`GET /api/v1`](references/endpoints.md#get-apiv1) - [`GET /api/v1/providers/{provider}/models`](references/endpoints.md#get-apiv1providersprovidermodels) - [`GET /api/v1/management/proxy-subscriptions`](references/endpoints.md#get-apiv1managementproxy-subscriptions) - [`POST /api/v1/management/proxy-subscriptions`](references/endpoints.md#post-apiv1managementproxy-subscriptions) - [`GET /api/v1/management/proxy-subscriptions/{id}`](references/endpoints.md#get-apiv1managementproxy-subscriptionsid) - [`PATCH /api/v1/management/proxy-subscriptions/{id}`](references/endpoints.md#patch-apiv1managementproxy-subscriptionsid) - [`DELETE /api/v1/management/proxy-subscriptions/{id}`](references/endpoints.md#delete-apiv1managementproxy-subscriptionsid) - [`GET /api/v1/management/proxy-subscriptions/{id}/nodes`](references/endpoints.md#get-apiv1managementproxy-subscriptionsidnodes) - [`POST /api/v1/management/proxy-subscriptions/{id}/refresh`](references/endpoints.md#post-apiv1managementproxy-subscriptionsidrefresh) - [`POST /api/v1/ocr`](references/endpoints.md#post-apiv1ocr) - [`POST /api/v1/audio/translations`](references/endpoints.md#post-apiv1audiotranslations) - [`GET /api/v1/voices`](references/endpoints.md#get-apiv1voices) - [`POST /api/v1/speech-to-text`](references/endpoints.md#post-apiv1speech-to-text) - [`POST /api/v1/text-to-speech/{voiceId}`](references/endpoints.md#post-apiv1text-to-speechvoiceid) - [`GET /api/v1/explain/routing`](references/endpoints.md#get-apiv1explainrouting) - [`GET /api/v1/providers/suggested-models`](references/endpoints.md#get-apiv1providerssuggested-models) - [`GET /api/v1/provider-plugin-manifest`](references/endpoints.md#get-apiv1provider-plugin-manifest) - [`GET /api/v1/{omnirouteCatchAll}`](references/endpoints.md#get-apiv1omniroutecatchall) - [`POST /api/v1/{omnirouteCatchAll}`](references/endpoints.md#post-apiv1omniroutecatchall) - [`PUT /api/v1/{omnirouteCatchAll}`](references/endpoints.md#put-apiv1omniroutecatchall) - [`PATCH /api/v1/{omnirouteCatchAll}`](references/endpoints.md#patch-apiv1omniroutecatchall) - [`DELETE /api/v1/{omnirouteCatchAll}`](references/endpoints.md#delete-apiv1omniroutecatchall) - [`GET /api/v1/accounts/{id}/limits`](references/endpoints.md#get-apiv1accountsidlimits) - [`PUT /api/v1/accounts/{id}/limits`](references/endpoints.md#put-apiv1accountsidlimits) - [`GET /api/v1/agents/credentials`](references/endpoints.md#get-apiv1agentscredentials) - [`POST /api/v1/agents/credentials`](references/endpoints.md#post-apiv1agentscredentials) - [`GET /api/v1/agents/health`](references/endpoints.md#get-apiv1agentshealth) - [`GET /api/v1/agents/tasks`](references/endpoints.md#get-apiv1agentstasks) - [`POST /api/v1/agents/tasks`](references/endpoints.md#post-apiv1agentstasks) - [`DELETE /api/v1/agents/tasks`](references/endpoints.md#delete-apiv1agentstasks) - [`GET /api/v1/agents/tasks/{id}`](references/endpoints.md#get-apiv1agentstasksid) - [`POST /api/v1/agents/tasks/{id}`](references/endpoints.md#post-apiv1agentstasksid) - [`DELETE /api/v1/agents/tasks/{id}`](references/endpoints.md#delete-apiv1agentstasksid) - [`POST /api/v1/antigravity`](references/endpoints.md#post-apiv1antigravity) - [`GET /api/v1/auto-combo/{channel}/candidates`](references/endpoints.md#get-apiv1auto-combochannelcandidates) - [`GET /api/v1/batches`](references/endpoints.md#get-apiv1batches) - [`POST /api/v1/batches`](references/endpoints.md#post-apiv1batches) - [`GET /api/v1/batches/{id}`](references/endpoints.md#get-apiv1batchesid) - [`DELETE /api/v1/batches/{id}`](references/endpoints.md#delete-apiv1batchesid) - [`POST /api/v1/batches/{id}/cancel`](references/endpoints.md#post-apiv1batchesidcancel) - [`DELETE /api/v1/batches/delete-completed`](references/endpoints.md#delete-apiv1batchesdelete-completed) - [`POST /api/v1/classify`](references/endpoints.md#post-apiv1classify) - [`GET /api/v1/combos`](references/endpoints.md#get-apiv1combos) - [`POST /api/v1/completions`](references/endpoints.md#post-apiv1completions) - [`GET /api/v1/files`](references/endpoints.md#get-apiv1files) - [`POST /api/v1/files`](references/endpoints.md#post-apiv1files) - [`GET /api/v1/files/{id}`](references/endpoints.md#get-apiv1filesid) - [`DELETE /api/v1/files/{id}`](references/endpoints.md#delete-apiv1filesid) - [`GET /api/v1/files/{id}/content`](references/endpoints.md#get-apiv1filesidcontent) - [`POST /api/v1/images/edits`](references/endpoints.md#post-apiv1imagesedits) - [`GET /api/v1/images/upscale`](references/endpoints.md#get-apiv1imagesupscale) - [`POST /api/v1/images/upscale`](references/endpoints.md#post-apiv1imagesupscale) - [`POST /api/v1/issues/report`](references/endpoints.md#post-apiv1issuesreport) - [`GET /api/v1/management/proxies`](references/endpoints.md#get-apiv1managementproxies) - [`POST /api/v1/management/proxies`](references/endpoints.md#post-apiv1managementproxies) - [`PATCH /api/v1/management/proxies`](references/endpoints.md#patch-apiv1managementproxies) - [`DELETE /api/v1/management/proxies`](references/endpoints.md#delete-apiv1managementproxies) - [`GET /api/v1/management/proxies/assignments`](references/endpoints.md#get-apiv1managementproxiesassignments) - [`PUT /api/v1/management/proxies/assignments`](references/endpoints.md#put-apiv1managementproxiesassignments) - [`PUT /api/v1/management/proxies/bulk-assign`](references/endpoints.md#put-apiv1managementproxiesbulk-assign) - [`GET /api/v1/management/proxies/health`](references/endpoints.md#get-apiv1managementproxieshealth) - [`GET /api/v1/me/status`](references/endpoints.md#get-apiv1mestatus) - [`GET /api/v1/muse-code/models`](references/endpoints.md#get-apiv1muse-codemodels) - [`GET /api/v1/music/generations`](references/endpoints.md#get-apiv1musicgenerations) - [`POST /api/v1/music/generations`](references/endpoints.md#post-apiv1musicgenerations) - [`GET /api/v1/providers/{provider}/limits`](references/endpoints.md#get-apiv1providersproviderlimits) - [`PUT /api/v1/providers/{provider}/limits`](references/endpoints.md#put-apiv1providersproviderlimits) - [`GET /api/v1/quotas/check`](references/endpoints.md#get-apiv1quotascheck) - [`GET /api/v1/registered-keys`](references/endpoints.md#get-apiv1registered-keys) - [`POST /api/v1/registered-keys`](references/endpoints.md#post-apiv1registered-keys) - [`GET /api/v1/registered-keys/{id}`](references/endpoints.md#get-apiv1registered-keysid) - [`DELETE /api/v1/registered-keys/{id}`](references/endpoints.md#delete-apiv1registered-keysid) - [`POST /api/v1/registered-keys/{id}/revoke`](references/endpoints.md#post-apiv1registered-keysidrevoke) - [`POST /api/v1/relay/chat/completions`](references/endpoints.md#post-apiv1relaychatcompletions) - [`POST /api/v1/relay/chat/completions/bifrost`](references/endpoints.md#post-apiv1relaychatcompletionsbifrost) - [`POST /api/v1/responses/{path}`](references/endpoints.md#post-apiv1responsespath) - [`GET /api/v1/search/analytics`](references/endpoints.md#get-apiv1searchanalytics) - [`POST /api/v1/segment`](references/endpoints.md#post-apiv1segment) - [`GET /api/v1/video-bridge/drilldown`](references/endpoints.md#get-apiv1video-bridgedrilldown) - [`DELETE /api/v1/video-bridge/drilldown`](references/endpoints.md#delete-apiv1video-bridgedrilldown) - [`GET /api/v1/videos/generations`](references/endpoints.md#get-apiv1videosgenerations) - [`POST /api/v1/videos/generations`](references/endpoints.md#post-apiv1videosgenerations) - [`GET /api/v1/vscode/{token}`](references/endpoints.md#get-apiv1vscodetoken) - [`POST /api/v1/vscode/{token}/api/chat`](references/endpoints.md#post-apiv1vscodetokenapichat) - [`POST /api/v1/vscode/{token}/api/show`](references/endpoints.md#post-apiv1vscodetokenapishow) - [`GET /api/v1/vscode/{token}/api/tags`](references/endpoints.md#get-apiv1vscodetokenapitags) - [`GET /api/v1/vscode/{token}/api/version`](references/endpoints.md#get-apiv1vscodetokenapiversion) - [`POST /api/v1/vscode/{token}/chat/completions`](references/endpoints.md#post-apiv1vscodetokenchatcompletions) - [`GET /api/v1/vscode/{token}/combos`](references/endpoints.md#get-apiv1vscodetokencombos) - [`GET /api/v1/vscode/{token}/models`](references/endpoints.md#get-apiv1vscodetokenmodels) - [`POST /api/v1/vscode/{token}/responses`](references/endpoints.md#post-apiv1vscodetokenresponses) - [`POST /api/v1/vscode/{token}/v1/chat/completions`](references/endpoints.md#post-apiv1vscodetokenv1chatcompletions) - [`GET /api/v1/vscode/{token}/v1/models`](references/endpoints.md#get-apiv1vscodetokenv1models) - [`GET /api/v1/vscode/combos/{token}/{{slug}}`](references/endpoints.md#get-apiv1vscodecombostokenslug) - [`POST /api/v1/vscode/combos/{token}/{{slug}}`](references/endpoints.md#post-apiv1vscodecombostokenslug) - [`GET /api/v1/vscode/raw/{token}`](references/endpoints.md#get-apiv1vscoderawtoken) - [`POST /api/v1/vscode/raw/{token}/api/chat`](references/endpoints.md#post-apiv1vscoderawtokenapichat) - [`POST /api/v1/vscode/raw/{token}/api/show`](references/endpoints.md#post-apiv1vscoderawtokenapishow) - [`GET /api/v1/vscode/raw/{token}/api/tags`](references/endpoints.md#get-apiv1vscoderawtokenapitags) - [`GET /api/v1/vscode/raw/{token}/api/version`](references/endpoints.md#get-apiv1vscoderawtokenapiversion) - [`POST /api/v1/vscode/raw/{token}/chat/completions`](references/endpoints.md#post-apiv1vscoderawtokenchatcompletions) - [`GET /api/v1/vscode/raw/{token}/combos`](references/endpoints.md#get-apiv1vscoderawtokencombos) - [`GET /api/v1/vscode/raw/{token}/models`](references/endpoints.md#get-apiv1vscoderawtokenmodels) - [`POST /api/v1/vscode/raw/{token}/responses`](references/endpoints.md#post-apiv1vscoderawtokenresponses) - [`POST /api/v1/vscode/raw/{token}/v1/chat/completions`](references/endpoints.md#post-apiv1vscoderawtokenv1chatcompletions) - [`GET /api/v1/vscode/raw/{token}/v1/models`](references/endpoints.md#get-apiv1vscoderawtokenv1models) - [`POST /api/v1/web/fetch`](references/endpoints.md#post-apiv1webfetch) ## Payloads See the full OpenAPI specification at `GET /api/openapi/spec` or `docs/openapi.yaml` for detailed request/response schemas. ## Chat completions Requires `OMNIROUTE_URL` and `OMNIROUTE_KEY`. See [entry-point SKILL](https://raw.githubusercontent.com/diegosouzapw/OmniRoute/main/skills/omniroute/SKILL.md) for setup. ### Endpoints - `POST $OMNIROUTE_URL/v1/chat/completions` — OpenAI format - `POST $OMNIROUTE_URL/v1/messages` — Anthropic Messages format - `POST $OMNIROUTE_URL/v1/responses` — OpenAI Responses API ### Discover ```bash curl $OMNIROUTE_URL/v1/models | jq '.data[].id' ``` Combos (e.g. `auto`, `cost-optimized`, `subscription`) auto-fallback through multiple providers. ### OpenAI format example ```bash curl -X POST $OMNIROUTE_URL/v1/chat/completions \ -H "Authorization: Bearer $OMNIROUTE_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-opus-4-7", "messages": [{"role": "user", "content": "Refactor this function"}], "stream": true }' ``` ### Anthropic format example ```bash curl -X POST $OMNIROUTE_URL/v1/messages \ -H "Authorization: Bearer $OMNIROUTE_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-opus-4-7", "max_tokens": 4096, "messages": [{"role": "user", "content": "Hi"}] }' ``` ### Tool use Supports OpenAI `tools` array and Anthropic `tools` block. Tool results auto-compressed via RTK (47 filters: git-diff, grep, test-jest, terraform-plan, docker-logs, etc.) — 20-40% token savings. Disable per-request with `X-Omniroute-Rtk: off` header. ### Reasoning / thinking Anthropic extended thinking and OpenAI Responses reasoning blocks are forwarded verbatim. Cached automatically via reasoning cache. ### Errors - `401` → invalid API key - `400 invalid_model` → model not in registry; check `/v1/models` - `503 circuit_open` → provider circuit breaker tripped; retry later or use combo - `429 rate_limited` → honor `Retry-After`; consider using a combo for auto-fallback ## Image generation Requires `OMNIROUTE_URL` and `OMNIROUTE_KEY`. See [entry-point SKILL](https://raw.githubusercontent.com/diegosouzapw/OmniRoute/main/skills/omniroute/SKILL.md) for setup. ### Endpoints - `POST $OMNIROUTE_URL/v1/images/generations` — Text-to-image - `POST $OMNIROUTE_URL/v1/images/edits` — Image edit (mask) - `POST $OMNIROUTE_URL/v1/images/variations` — Variations ### Discover ```bash curl $OMNIROUTE_URL/v1/models/image | jq '.data[]' ``` Returns `{ id, owned_by, sizes:[...], capabilities:[...] }` per model. ### Generate example ```bash curl -X POST $OMNIROUTE_URL/v1/images/generations \ -H "Authorization: Bearer $OMNIROUTE_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "dall-e-3", "prompt": "a red bicycle on a wet street, photoreal", "n": 1, "size": "1024x1024", "response_format": "b64_json" }' ``` Response: `{ created, data: [{ url? or b64_json, revised_prompt }] }` ### Errors - `400 invalid_size` → not supported by this model; check `/v1/models/image` - `400 content_policy_violation` → blocked by provider safety - `503` → provider unavailable; try another model in `/v1/models/image` ## Text-to-speech Requires `OMNIROUTE_URL` and `OMNIROUTE_KEY`. See [entry-point SKILL](https://raw.githubusercontent.com/diegosouzapw/OmniRoute/main/skills/omniroute/SKILL.md) for setup. ### Endpoint - `POST $OMNIROUTE_URL/v1/audio/speech` — returns binary audio (mp3/opus/wav/flac) ### Discover ```bash curl $OMNIROUTE_URL/v1/models/tts | jq '.data[]' ``` Each entry includes `voices:[...]` for the available voice names per provider. ### Example ```bash curl -X POST $OMNIROUTE_URL/v1/audio/speech \ -H "Authorization: Bearer $OMNIROUTE_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "tts-1", "input": "Hello from OmniRoute.", "voice": "alloy", "response_format": "mp3" }' --output speech.mp3 ``` ### Voices Voice names vary by provider. Check `/v1/models/tts` — each entry has `voices:[...]`. Common OpenAI voices: `alloy`, `echo`, `fable`, `onyx`, `nova`, `shimmer`. ### Errors - `400 invalid_voice` → voice not supported by this model - `400 input_too_long` → input exceeds model character limit - `503` → provider unavailable; try another model in `/v1/models/tts` ## Speech-to-text Requires `OMNIROUTE_URL` and `OMNIROUTE_KEY`. See [entry-point SKILL](https://raw.githubusercontent.com/diegosouzapw/OmniRoute/main/skills/omniroute/SKILL.md) for setup. ### Endpoints - `POST $OMNIROUTE_URL/v1/audio/transcriptions` — multipart upload, returns text - `POST $OMNIROUTE_URL/v1/audio/translations` — transcribe + translate to English ### Discover ```bash curl $OMNIROUTE_URL/v1/models/stt | jq '.data[]' ``` ### Example ```bash curl -X POST $OMNIROUTE_URL/v1/audio/transcriptions \ -H "Authorization: Bearer $OMNIROUTE_KEY" \ -F "file=@audio.mp3" \ -F "model=whisper-1" \ -F "response_format=verbose_json" ``` Response: `{ text, language, duration, segments?:[{ start, end, text }] }` ### Supported formats Audio: `mp3`, `mp4`, `mpeg`, `mpga`, `m4a`, `wav`, `webm`. Response formats: `json`, `text`, `srt`, `verbose_json`, `vtt`. ### Errors - `400 invalid_file_format` → unsupported audio format - `400 file_too_large` → exceeds provider limit (usually 25MB) - `503` → provider unavailable; try another model in `/v1/models/stt` ## Embeddings Requires `OMNIROUTE_URL` and `OMNIROUTE_KEY`. See [entry-point SKILL](https://raw.githubusercontent.com/diegosouzapw/OmniRoute/main/skills/omniroute/SKILL.md) for setup. ### Endpoint - `POST $OMNIROUTE_URL/v1/embeddings` ### Discover ```bash curl $OMNIROUTE_URL/v1/models/embedding | jq '.data[]' ``` Each entry: `{ id, owned_by, dimensions, max_input_tokens }`. ### Example ```bash curl -X POST $OMNIROUTE_URL/v1/embeddings \ -H "Authorization: Bearer $OMNIROUTE_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "text-embedding-3-large", "input": ["first text", "second text"], "encoding_format": "float" }' ``` Response: `{ data:[{ embedding:[...], index }], usage:{ prompt_tokens, total_tokens } }` ### Batch input `input` accepts a string or array of strings (up to provider batch limit, typically 2048 items). ### Errors - `400 input_too_long` → input exceeds `max_input_tokens` for this model - `400 invalid_encoding_format` → use `float` or `base64` - `503` → provider unavailable; try another model in `/v1/models/embedding` ## Web search Requires `OMNIROUTE_URL` and `OMNIROUTE_KEY`. See [entry-point SKILL](https://raw.githubusercontent.com/diegosouzapw/OmniRoute/main/skills/omniroute/SKILL.md) for setup. ### Endpoint - `POST $OMNIROUTE_URL/v1/web/search` — unified search format ### Discover ```bash curl $OMNIROUTE_URL/v1/models/web | jq '.data[] | select(.kind == "webSearch")' ``` ### Example ```bash curl -X POST $OMNIROUTE_URL/v1/web/search \ -H "Authorization: Bearer $OMNIROUTE_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "tavily/search", "query": "OmniRoute github latest release", "max_results": 5, "include_answer": true }' ``` Response: `{ answer?, results:[{ url, title, content, score }] }` ### Parameters | Field | Type | Description | | ---------------- | ------- | ------------------------------------ | | `model` | string | Provider model from `/v1/models/web` | | `query` | string | Search query | | `max_results` | number | Max results (default: 5) | | `include_answer` | boolean | Include AI-synthesized answer | | `search_depth` | string | `basic` or `advanced` (Tavily) | ### Errors - `400 query_too_long` → shorten the search query - `503` → provider unavailable; try another model in `/v1/models/web` ## Web fetch Requires `OMNIROUTE_URL` and `OMNIROUTE_KEY`. See [entry-point SKILL](https://raw.githubusercontent.com/diegosouzapw/OmniRoute/main/skills/omniroute/SKILL.md) for setup. ### Endpoint - `POST $OMNIROUTE_URL/v1/web/fetch` ### Discover ```bash curl $OMNIROUTE_URL/v1/models/web | jq '.data[] | select(.kind == "webFetch")' ``` ### Example ```bash curl -X POST $OMNIROUTE_URL/v1/web/fetch \ -H "Authorization: Bearer $OMNIROUTE_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "jina/reader", "url": "https://anthropic.com", "format": "markdown" }' ``` Response: `{ url, title, markdown, links?:[...], images?:[...] }` ### Parameters | Field | Type | Description | | -------- | ------ | ----------------------------------------------------------------------- | | `model` | string | Provider from `/v1/models/web` (e.g. `jina/reader`, `firecrawl/scrape`) | | `url` | string | URL to fetch | | `format` | string | `markdown` (default), `html`, `text` | ### Errors - `400 invalid_url` → URL must be http/https - `403 blocked` → provider blocked by target site; try a different model - `503` → provider unavailable; try another model in `/v1/models/web`