# Static model catalog for the openai connector — flagship reasoning models # plus the everyday chat workhorse, in the vendor's own model ids. Pricing is # cents per million tokens. # Prompt-cache prices (`cacheRead` = cached input, `cacheWrite` = cache writes) # and the gpt-5.6-sol pair ($4/$20) read from developers.openai.com/api/docs/pricing # on 2026-09-20 (Standard, short context). gpt-5.5's cached-input figure is the # OpenRouter listing's (the vendor page no longer lists the model); # gpt-5.5-pro and gpt-5.3-chat-latest publish no cache price. # Sol 6.1: https://developers.openai.com/api/docs/models/gpt-6.1-sol # (2026-10-03). Tools require Responses; Chat Completions accepts no tools. # Reasoning cannot be disabled: neither none nor minimal is supported. # Codex subscription availability: https://learn.chatgpt.com/docs/models # The GPT-6 family (Astra, Sol, Luna) read from the same pages and from the # vendor's latest-model guide on 2026-10-03: "GPT-6 Astra and GPT-6.1 Sol # support Chat Completions, but tool calling requires Responses", and "GPT-6 # Sol and GPT-6 Luna support function calling in Chat Completions only with # reasoning_effort: none". So Astra declares the Responses tool API (direct # chat calls it there), while Sol and Luna declare off + toolsRequireOff like # the 5.6 trio. Astra and Sol bill prompts over 272K input tokens at 2x input # and cache rates and 1.5x output, which this per-token schema cannot express; # the short-context prices are listed. Luna's cache writes are 1.25x its input # rate, as the vendor states it. - id: gpt-6.1-sol provider: openai tags: - chat supportsTools: true toolCallingApi: responses supportsVision: true reasoning: knob: effort contextWindow: 1050000 maxOutputTokens: 128000 pricing: inputCentsPerMillion: 200 outputCentsPerMillion: 1000 cacheReadCentsPerMillion: 10 cacheWriteCentsPerMillion: 250 - id: gpt-6-astra provider: openai tags: - chat supportsTools: true toolCallingApi: responses supportsVision: true reasoning: knob: effort contextWindow: 1050000 maxOutputTokens: 128000 pricing: inputCentsPerMillion: 1000 outputCentsPerMillion: 5000 cacheReadCentsPerMillion: 100 cacheWriteCentsPerMillion: 1250 - id: gpt-6-sol provider: openai tags: - chat supportsTools: true supportsVision: true reasoning: knob: effort off: none toolsRequireOff: true contextWindow: 1050000 maxOutputTokens: 128000 pricing: inputCentsPerMillion: 200 outputCentsPerMillion: 1000 cacheReadCentsPerMillion: 20 cacheWriteCentsPerMillion: 250 - id: gpt-6-luna provider: openai tags: - chat supportsTools: true supportsVision: true reasoning: knob: effort off: none toolsRequireOff: true contextWindow: 1050000 maxOutputTokens: 128000 pricing: inputCentsPerMillion: 10 outputCentsPerMillion: 50 cacheReadCentsPerMillion: 1 cacheWriteCentsPerMillion: 12.5 - id: gpt-5.5 provider: openai tags: - chat supportsTools: true supportsVision: true reasoning: knob: effort # Probed 2026-08-14 via the provider's own 400: "Function tools with # reasoning_effort are not supported for gpt-5.5 in /v1/chat/completions. # To use function tools, use /v1/responses or set reasoning_effort to # 'none'." — chat always carries tools, so every turn (Default step AND # sticky effort picks) must send none; the picker offers no levels. off: none toolsRequireOff: true # Vendor model page: 1,050,000-token context, 128,000-token output. contextWindow: 1050000 maxOutputTokens: 128000 pricing: inputCentsPerMillion: 500 outputCentsPerMillion: 3000 cacheReadCentsPerMillion: 50 - id: gpt-5.5-pro provider: openai tags: - chat supportsTools: true supportsVision: true reasoning: knob: effort # Family behavior (see gpt-5.5 above); the pro tier may turn out to be # served on /v1/responses only — a 404 here means repoint, not retune. off: none toolsRequireOff: true contextWindow: 400000 maxOutputTokens: 128000 pricing: inputCentsPerMillion: 3000 outputCentsPerMillion: 18000 # The chat-tuned variant ships under the vendor's rolling "-latest" alias — # the bare id 404s ("The model `gpt-5.3-chat` does not exist…", probed # 2026-08-14). Non-reasoning: takes max_completion_tokens AND a custom # temperature. - id: gpt-5.3-chat-latest provider: openai tags: - chat supportsTools: true supportsVision: true contextWindow: 128000 maxOutputTokens: 16384 pricing: inputCentsPerMillion: 175 outputCentsPerMillion: 1400 # The 5.6 frontier trio — ids, sizes, and prices from the vendor's model # catalog page (sol $5/$30, terra $2/$12, luna $0.20/$1.20 per million; # 1.05M context / 128K output each). The page frames them as Responses-API # models, but /v1/chat/completions serves them too — its parameter 400 # ("Function tools with reasoning_effort are not supported for gpt-5.6-sol…") # proves the route and names the restriction, so all three declare # off + toolsRequireOff. Codex variants are Responses-only and NOT listed; # gpt-5.6-luna-pro exists only in the openrouter catalog, not here. - id: gpt-5.6-sol provider: openai tags: - chat supportsTools: true supportsVision: true reasoning: knob: effort off: none toolsRequireOff: true contextWindow: 1050000 maxOutputTokens: 128000 pricing: inputCentsPerMillion: 400 outputCentsPerMillion: 2000 cacheReadCentsPerMillion: 40 cacheWriteCentsPerMillion: 500 - id: gpt-5.6-terra provider: openai tags: - chat supportsTools: true supportsVision: true reasoning: knob: effort off: none toolsRequireOff: true contextWindow: 1050000 maxOutputTokens: 128000 pricing: inputCentsPerMillion: 200 outputCentsPerMillion: 1200 cacheReadCentsPerMillion: 20 cacheWriteCentsPerMillion: 250 - id: gpt-5.6-luna provider: openai tags: - chat supportsTools: true supportsVision: true reasoning: knob: effort off: none toolsRequireOff: true contextWindow: 1050000 maxOutputTokens: 128000 pricing: inputCentsPerMillion: 20 outputCentsPerMillion: 120 cacheReadCentsPerMillion: 2 cacheWriteCentsPerMillion: 25 # Text-to-speech: powers "Read replies aloud" / "Speak out loud". The context # window is nominal — synthesis takes character-capped requests, not context. - id: gpt-4o-mini-tts provider: openai tags: - text-to-speech supportsTools: false supportsVision: false contextWindow: 2000 tts: defaultVoice: alloy voicesByLocale: en: alloy de: onyx fr: nova audioFormat: mp3 centsPerMillionCharacters: 1200 # Transcription: powers browser dictation's server fallback and audio/video # attachment transcription, both through `/audio/transcriptions`. Without an # entry carrying this tag `resolveTranscriptionModel` finds nothing and every # upload fails `NO_TRANSCRIPTION_MODEL` — the capability exists only because # it is declared here, exactly as with the embedding widths below. # # whisper-1 rather than the gpt-4o-transcribe family: the request asks for # `response_format=verbose_json` and the paragraphizer needs the per-segment # timestamps it returns, which the newer transcribe models do not serve. # Pricing is per audio MINUTE, which this catalog's per-token schema cannot # express — so it is omitted and the ledger records minutes at a zero cost # estimate. The context window is nominal: transcription takes a file. - id: whisper-1 provider: openai tags: - transcription supportsTools: false supportsVision: false contextWindow: 2000 # Image generation: the models agents call through the generate_image tool, # over /v1/images/generations (and /v1/images/edits when the agent passes # reference images). They write images, never text, so they carry no chat # tag and every chat list skips them. Prices read from # developers.openai.com/api/docs/pricing on 2026-09-29 (Standard): `input` is # the text-input price, `imageInput` the reference-image input price and # `output` the image-output price, per million tokens — the ledger costs a # call from the token counts the response reports. The context window is # nominal: generation takes a prompt of at most 32,000 characters. - id: gpt-image-1 provider: openai tags: - image-generation supportsTools: false supportsVision: true outputsMedia: true contextWindow: 32000 pricing: inputCentsPerMillion: 500 outputCentsPerMillion: 4000 imageInputCentsPerMillion: 1000 - id: gpt-image-1-mini provider: openai tags: - image-generation supportsTools: false supportsVision: true outputsMedia: true contextWindow: 32000 pricing: inputCentsPerMillion: 200 outputCentsPerMillion: 800 imageInputCentsPerMillion: 250 # Embedding model — the everyday knowledge-indexing workhorse. 1536 native # dimensions, under pgvector's HNSW index limit. - id: text-embedding-3-small provider: openai tags: - embedding supportsTools: false supportsVision: false contextWindow: 8191 pricing: inputCentsPerMillion: 2 outputCentsPerMillion: 0 embedding: dimensions: 1536 recommended: true