{ "version": 2, "license": "CC-BY-4.0", "models": [ { "provider": "Anthropic", "provider_url": "https://www.anthropic.com", "model_id": "claude-opus-4-7", "display_name": "Claude Opus 4.7", "model_family": "Claude 4", "knowledge_cutoff": "2026-01-31", "aliases": ["anthropic.claude-opus-4-7"], "context_window": 1000000, "max_output_tokens": 128000, "input_per_mtok_usd": "5.0", "output_per_mtok_usd": "25.0", "cache_read_per_mtok_usd": "0.5", "cache_write_per_mtok_usd": "6.25", "batch_input_per_mtok_usd": "2.5", "batch_output_per_mtok_usd": "12.5", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "bedrock", "vertex"], "last_verified": "2026-09-02", "last_changed_at": "2026-05-05", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.anthropic.com/pricing", "notes": "Cache hit pricing is 0.1x base input ($0.50/MTok); 5-minute cache write is 1.25x base ($6.25/MTok); 1-hour cache write is 2x ($10/MTok). Batch API discounts both input and output by 50%. Knowledge cutoff Jan 2026. Uses a new tokenizer vs prior Claude models (may use up to 35% more tokens for identical text). Supports adaptive thinking (no extended-thinking toggle); thinking output tokens are billed at the output rate. Legacy listing (superseded by Claude Opus 4.8 as of 2026-05-28) but still fully active with unchanged pricing." }, { "provider": "Anthropic", "provider_url": "https://www.anthropic.com", "model_id": "claude-sonnet-4-6", "display_name": "Claude Sonnet 4.6", "model_family": "Claude 4", "knowledge_cutoff": "2025-08-31", "aliases": ["anthropic.claude-sonnet-4-6"], "context_window": 1000000, "max_output_tokens": 128000, "input_per_mtok_usd": "3.0", "output_per_mtok_usd": "15.0", "cache_read_per_mtok_usd": "0.3", "cache_write_per_mtok_usd": "3.75", "batch_input_per_mtok_usd": "1.5", "batch_output_per_mtok_usd": "7.5", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "bedrock", "vertex"], "last_verified": "2026-09-02", "last_changed_at": "2026-05-05", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.anthropic.com/pricing", "notes": "Cache hit pricing is 0.1x base input ($0.30/MTok); 5-minute cache write is 1.25x base ($3.75/MTok); 1-hour cache write is 2x ($6/MTok). Batch API discounts both input and output by 50%. Knowledge cutoff Aug 2025. Supports extended thinking and adaptive thinking; thinking output tokens are billed at the output rate. Corrected max_output_tokens to 128000 (was incorrectly 64000) per platform.claude.com/docs/en/docs/about-claude/models; not a price field, so last_changed_at is not bumped. Still active; legacy listing alongside Claude Sonnet 5 (launched 2026-06-30)." }, { "provider": "Anthropic", "provider_url": "https://www.anthropic.com", "model_id": "claude-fable-5", "display_name": "Claude Fable 5", "model_family": "Claude 5", "knowledge_cutoff": "2026-01-31", "aliases": ["anthropic.claude-fable-5"], "context_window": 1000000, "max_output_tokens": 128000, "input_per_mtok_usd": "10.0", "output_per_mtok_usd": "50.0", "cache_read_per_mtok_usd": "1.0", "cache_write_per_mtok_usd": "12.5", "batch_input_per_mtok_usd": "5.0", "batch_output_per_mtok_usd": "25.0", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "bedrock", "vertex"], "last_verified": "2026-09-02", "last_changed_at": "2026-06-09", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://platform.claude.com/docs/en/about-claude/pricing", "notes": "Anthropic's most capable widely released model as of its GA on 2026-06-09; now a legacy listing (superseded by Claude Fable 5.1 as of 2026-09-01) but still fully active with unchanged pricing. GA on the Claude API, Claude Platform on AWS, Amazon Bedrock, Google Cloud, and Microsoft Foundry (Foundry has no deployment_options enum value, so omitted). Cache hit is 0.1x base input ($1.00/MTok); 5-minute cache write is 1.25x ($12.50/MTok); 1-hour cache write is 2x ($20/MTok). Batch API discounts both input and output by 50%. Adaptive thinking is always on (thinking:{type:'disabled'} returns 400); raw chain of thought is never returned (thinking.display defaults to 'omitted'). Safety classifiers may decline requests (stop_reason:'refusal'). Requires 30-day data retention; not available under zero data retention. Uses the same tokenizer as Opus 4.7/4.8 (~30% more tokens than pre-4.7 models for identical text). Access to Fable 5 was briefly suspended and restored on 2026-07-01 (no price change). Claude Mythos 5 (claude-mythos-5) shares identical specs and pricing but is invitation-only via Project Glasswing (not self-serve) — deferred, not added as a separate row." }, { "provider": "Anthropic", "provider_url": "https://www.anthropic.com", "model_id": "claude-fable-5-1", "display_name": "Claude Fable 5.1", "model_family": "Claude 5", "knowledge_cutoff": "2026-06-30", "aliases": ["anthropic.claude-fable-5-1"], "context_window": 1000000, "max_output_tokens": 128000, "input_per_mtok_usd": "10.0", "output_per_mtok_usd": "50.0", "cache_read_per_mtok_usd": "0.25", "cache_write_per_mtok_usd": "12.5", "batch_input_per_mtok_usd": "5.0", "batch_output_per_mtok_usd": "25.0", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "bedrock", "vertex"], "last_verified": "2026-09-02", "last_changed_at": "2026-09-01", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://platform.claude.com/docs/en/about-claude/pricing", "notes": "New row this refresh. Released 2026-09-01 as the latest, recommended model for demanding reasoning and long-horizon agentic work (Anthropic recommends starting with Opus 5 for most workloads; use Fable 5.1 when Opus 5 at higher effort falls short). Extends Claude Fable 5 at the same base input/output, 5-minute cache write, 1-hour cache write, and batch prices, but cache reads (hits/refreshes) are priced at 0.025x base input ($0.25/MTok) instead of the standard 0.1x multiplier used by every other current Claude model — this is the one structured price field that differs from Fable 5, hence the new row rather than an in-place update. 1M context window, 128k max output. Adaptive thinking always on; forced tool use (tool_choice any/tool) returns an error (breaking change vs Fable 5); thinking blocks are tied to the model that produced them and are invalidated by editing earlier turns. Available on Claude API, Claude Platform on AWS, Amazon Bedrock, Google Cloud, and Microsoft Foundry. Uses the same tokenizer as Fable 5/Opus 4.7+ (~30% more tokens than pre-4.7 models for identical text). Reliable knowledge cutoff and training data cutoff both Jun 2026. Retirement commitment: not sooner than 2027-09-01. Claude Mythos 5.1 (claude-mythos-5-1) shares identical specs and pricing but is invitation-only via Project Glasswing (not self-serve) — deferred, not added as a separate row." }, { "provider": "Anthropic", "provider_url": "https://www.anthropic.com", "model_id": "claude-opus-4-8", "display_name": "Claude Opus 4.8", "model_family": "Claude 4", "knowledge_cutoff": "2026-01-31", "aliases": ["anthropic.claude-opus-4-8"], "context_window": 1000000, "max_output_tokens": 128000, "input_per_mtok_usd": "5.0", "output_per_mtok_usd": "25.0", "cache_read_per_mtok_usd": "0.5", "cache_write_per_mtok_usd": "6.25", "batch_input_per_mtok_usd": "2.5", "batch_output_per_mtok_usd": "12.5", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "bedrock", "vertex"], "last_verified": "2026-09-02", "last_changed_at": "2026-05-28", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://platform.claude.com/docs/en/about-claude/pricing", "notes": "Anthropic's most capable generally available model, launched 2026-05-28 as the successor to Opus 4.7; still fully active with unchanged pricing. Same base pricing as Opus 4.7. Cache hit is 0.1x base input ($0.50/MTok); 5-minute cache write is 1.25x ($6.25/MTok); 1-hour cache write is 2x ($10/MTok). Batch API discounts both input and output by 50%. 1M context window at standard pricing (no long-context premium), 128k max output. Adaptive thinking supported; effort parameter defaults to 'high' on all surfaces. Fast mode available as a research preview (Claude API only) at $10/$50 per MTok input/output. Knowledge cutoff Jan 2026." }, { "provider": "Anthropic", "provider_url": "https://www.anthropic.com", "model_id": "claude-sonnet-5", "display_name": "Claude Sonnet 5", "featured": true, "model_family": "Claude 5", "knowledge_cutoff": "2026-01-31", "aliases": ["anthropic.claude-sonnet-5"], "context_window": 1000000, "max_output_tokens": 128000, "input_per_mtok_usd": "2.0", "output_per_mtok_usd": "10.0", "cache_read_per_mtok_usd": "0.2", "cache_write_per_mtok_usd": "2.5", "batch_input_per_mtok_usd": "1.0", "batch_output_per_mtok_usd": "5.0", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "bedrock", "vertex"], "last_verified": "2026-09-02", "last_changed_at": "2026-06-30", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://platform.claude.com/docs/en/about-claude/pricing", "notes": "Launched 2026-06-30, next-generation Sonnet. The $2/$10 per MTok input/output pricing (cache hit $0.20, 5-min cache write $2.50, 1-hour cache write $4, batch $1.00/$5.00) was announced at launch as introductory pricing through 2026-08-31; Anthropic's pricing page now states this is the standard price and the previously scheduled increase to $3/$15 on 2026-09-01 will not occur. No structured price field changed this refresh, so last_changed_at is not bumped. 1M context window, 128k max output. Adaptive thinking on by default (omitting thinking runs adaptive); manual extended thinking (budget_tokens) removed and returns 400; non-default temperature/top_p/top_k also return 400. New tokenizer produces ~30% more tokens than Sonnet 4.6 for identical text. Not available with Priority Tier. Knowledge cutoff Jan 2026." }, { "provider": "Anthropic", "provider_url": "https://www.anthropic.com", "model_id": "claude-opus-5", "display_name": "Claude Opus 5", "featured": true, "model_family": "Claude 5", "knowledge_cutoff": "2026-05-31", "aliases": ["anthropic.claude-opus-5"], "context_window": 1000000, "max_output_tokens": 128000, "input_per_mtok_usd": "5.0", "output_per_mtok_usd": "25.0", "cache_read_per_mtok_usd": "0.5", "cache_write_per_mtok_usd": "6.25", "batch_input_per_mtok_usd": "2.5", "batch_output_per_mtok_usd": "12.5", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "bedrock", "vertex"], "last_verified": "2026-09-02", "last_changed_at": "2026-06-09", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://platform.claude.com/docs/en/about-claude/models", "notes": "Launched 2026-06-09, recommended for complex agentic coding and enterprise work; still active with unchanged pricing. 1M context window, 128k max output. Pricing: $5/$25 per MTok base input/output. Cache hit is 0.1x base input ($0.50/MTok); 5-minute cache write is 1.25x ($6.25/MTok); 1-hour cache write is 2x ($10/MTok). Batch API discounts both input and output by 50% ($2.50/$12.50 per MTok). Adaptive thinking on by default (omitting thinking runs adaptive); manual extended thinking not supported (returns 400); non-default temperature/top_p/top_k also return 400. Uses the new tokenizer as of Opus 4.7/4.8 (~30% more tokens than pre-4.7 models for identical text). Available on Claude API, Amazon Bedrock, Google Cloud Vertex AI, Claude Platform on AWS, and Microsoft Foundry. Knowledge cutoff May 2026 (training data cutoff May 2026)." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-5", "display_name": "GPT-5", "model_family": "GPT-5", "knowledge_cutoff": "2024-09-30", "context_window": 400000, "max_output_tokens": 128000, "input_per_mtok_usd": "1.25", "output_per_mtok_usd": "10.0", "cache_read_per_mtok_usd": "0.125", "batch_input_per_mtok_usd": "0.625", "batch_output_per_mtok_usd": "5.0", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "azure"], "deprecated_at": "2026-06-11", "replaced_by_model_id": "gpt-5.6-sol", "last_verified": "2026-09-02", "last_changed_at": "2026-05-05", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-5", "notes": "Reasoning model with adjustable reasoning_effort; reasoning tokens are billed at the output rate. Context window 400k, max output 128k confirmed against developers.openai.com/api/docs/models/gpt-5. Cached input at 10% of base ($0.125/MTok). Batch API at flat 50% off input and output. PDF input via the Files API; image input native; audio is NOT supported on this model_id. Knowledge cutoff Sept 2024. Deprecation announced 2026-06-11: dated snapshot gpt-5-2025-08-07 scheduled for API removal 2026-12-11, recommended replacement gpt-5.5 (per developers.openai.com/api/docs/deprecations); bare 'gpt-5' still serving as of 2026-07-19, prices unchanged. 2026-07-13 re-verify: model card prose now points to the newer GPT-5.6 family ('We recommend using the latest GPT-5.6') as of GPT-5.6's release, but the structured deprecations table still redirects the dated gpt-5-2025-08-07 snapshot to gpt-5.5; kept replaced_by_model_id at gpt-5.5 pending an update to the formal deprecation entry. 2026-07-19 re-verify: model card live with prices/specs unchanged and deprecations table unchanged; note this model no longer appears on the main pricing page (prices verified against the model card). 2026-08-11 re-verify: prices/specs unchanged (back on the main pricing table this pass, matching the model card); the deprecations table now redirects gpt-5-2025-08-07 to gpt-5.6-sol (GPT-5.6 family's flagship, launched since the last pass) instead of gpt-5.5 — replaced_by_model_id updated to match." }, { "provider": "Google", "provider_url": "https://deepmind.google", "model_id": "gemini-2.5-pro", "display_name": "Gemini 2.5 Pro", "featured": true, "model_family": "Gemini 2.5", "knowledge_cutoff": "2025-01-31", "context_window": 1048576, "max_output_tokens": 65536, "input_per_mtok_usd": "1.25", "output_per_mtok_usd": "10.0", "cache_read_per_mtok_usd": "0.125", "cache_storage_per_mtok_per_hour_usd": "4.50", "batch_input_per_mtok_usd": "0.625", "batch_output_per_mtok_usd": "5.0", "pricing_tiers": [ { "threshold_tokens": 200000, "input_per_mtok_usd": "2.5", "output_per_mtok_usd": "15.0", "cache_read_per_mtok_usd": "0.25" } ], "modalities": { "input": ["text", "image", "audio", "video"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": true, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "vertex"], "confidence": "high", "last_verified": "2026-09-02", "last_changed_at": "2026-05-05", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://ai.google.dev/pricing", "notes": "Input pricing shown is for prompts <=200k tokens; prompts >200k tokens are billed per pricing_tiers; cached input also tiers at 200k (the >200k rate is captured in pricing_tiers[0].cache_read_per_mtok_usd). Context window and max_output_tokens confirmed via ai.google.dev/gemini-api/docs/models/gemini-2.5-pro. Thinking is always on and cannot be disabled; thinking tokens are billed at the output rate. Knowledge cutoff January 2025. Batch Mode discount is a flat 50% off input/output. Audio input is billed at the standard input rate of $1.25/MTok (no separate audio premium, unlike 2.5 Flash/Flash-Lite/2.0 Flash); audio_input_per_mtok_usd omitted. Explicit context caching storage is captured in cache_storage_per_mtok_per_hour_usd. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. PDF input via the Files API; image, audio, and video native. Re-verified 2026-07-02: prices, context window, and modalities unchanged; no new pricing tiers. Re-verified 2026-07-13: prices, tiers, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-07-19: prices, tiers, and modalities unchanged against ai.google.dev/pricing; still listed stable on ai.google.dev/gemini-api/docs/models. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models." }, { "provider": "DeepSeek", "provider_url": "https://www.deepseek.com", "model_id": "deepseek-v4-flash", "display_name": "DeepSeek V4 Flash", "model_family": "DeepSeek V4", "context_window": 1048576, "max_output_tokens": 384000, "input_per_mtok_usd": "0.44", "output_per_mtok_usd": "1.32", "cache_read_per_mtok_usd": "0.014", "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-08-16", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://api-docs.deepseek.com/quick_start/pricing", "notes": "MoE architecture: 284B total parameters, 13B activated. Replaces deepseek-chat (V3-era alias). Supports both non-thinking and thinking (default) modes; thinking output is billed at the same output rate, so reasoning_tokens_billed: true. deepseek-chat and deepseek-reasoner legacy aliases were discontinued 2026-07-24. Re-verified 2026-07-02 through 2026-08-11: prices unchanged; snapshot updated to DeepSeek-V4-Flash-0731 on 2026-08-11 check (no pricing impact then). Re-verified 2026-09-02: DeepSeek introduced peak/off-peak pricing effective 2026-08-16 (announced api-docs.deepseek.com/updates changelog entry dated 2026-08-13, filed under the V4 Pro GA release), realizing the price increase flagged as forthcoming in the 2026-08-11 refresh. Values recorded here are PEAK rates (list price): input (cache miss) $0.44, output $1.32, cache-hit input $0.014 per MTok — confirmed against the raw pricing table at api-docs.deepseek.com/quick_start/pricing. Off-peak rates are exactly half: cache hit $0.007, cache miss $0.22, output $0.66 per MTok, applying during peak-labeled hours' complement — i.e. off-peak is all hours EXCEPT 01:00-04:00 and 06:00-10:00 UTC, Monday-Friday (peak hours are the pricier window despite the off-peak/peak naming referring to demand, not price; peak hours are ~21% of the week). Schema has no time-of-day pricing axis, so off-peak is not separately encoded; use peak here as the conservative list price. Cache-hit input is now ~1/31 of cache-miss input ($0.014 / $0.44), not the previous 1/50 ratio. Cross-checked api-docs.deepseek.com/news/news260821 and api-docs.deepseek.com/updates: same-day (2026-08-21) launch of deepseek-v4-flash-vision-exp (new sibling row, vision-only variant); no other changes to this row's model_id, snapshot, context window, or concurrency limit (2500)." }, { "provider": "DeepSeek", "provider_url": "https://www.deepseek.com", "model_id": "deepseek-v4-pro", "display_name": "DeepSeek V4 Pro", "featured": true, "model_family": "DeepSeek V4", "context_window": 1048576, "max_output_tokens": 384000, "input_per_mtok_usd": "1.32", "output_per_mtok_usd": "3.96", "cache_read_per_mtok_usd": "0.044", "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-08-16", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://api-docs.deepseek.com/quick_start/pricing", "notes": "Pro-tier sibling to deepseek-v4-flash, positioned for higher-quality responses at lower concurrency (500 vs Flash's 2500). The permanent post-promo rate ($0.435 / $0.87 / $0.003625 cache hit) held from 2026-05-22 through the 2026-08-11 refresh. Re-verified 2026-09-02: DeepSeek-V4-Pro reached GA on 2026-08-13 (snapshot DeepSeek-V4-Pro-0813, previously unrecorded) and introduced peak/off-peak pricing effective 2026-08-16, per the changelog at api-docs.deepseek.com/updates. Values recorded here are PEAK rates (list price): input (cache miss) $1.32, output $3.96, cache-hit input $0.044 per MTok — confirmed against the raw pricing table at api-docs.deepseek.com/quick_start/pricing; exactly 3x the prior permanent rate. Off-peak rates are half: cache hit $0.022, cache miss $0.66, output $1.98 per MTok, applying all hours EXCEPT 01:00-04:00 and 06:00-10:00 UTC Monday-Friday (peak hours, ~21% of the week, are the pricier window). Schema has no time-of-day pricing axis, so off-peak is not separately encoded; peak used here as the conservative list price. Cache-hit input is now 1/30 of cache-miss input, versus 1/120 previously. GA also brought native Responses API support (previously flagged as coming early August, now live) and three thinking-effort levels (low/high/max), no further pricing impact. Cache miss continues to be billed at the base input rate; thinking output billed at the output rate, so reasoning_tokens_billed: true. Cross-checked api-docs.deepseek.com/news/news260821, no changes specific to this row (that entry covers the separate deepseek-v4-flash-vision-exp launch)." }, { "provider": "DeepSeek", "provider_url": "https://www.deepseek.com", "model_id": "deepseek-chat", "display_name": "DeepSeek V3", "model_family": "DeepSeek V3", "context_window": 1048576, "max_output_tokens": 384000, "input_per_mtok_usd": "0.14", "output_per_mtok_usd": "0.28", "cache_read_per_mtok_usd": "0.0028", "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native"], "deprecated_at": "2026-04-24", "replaced_by_model_id": "deepseek-v4-flash", "last_verified": "2026-09-02", "last_changed_at": "2026-04-24", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://api-docs.deepseek.com/quick_start/pricing", "notes": "Canonical API model_id for the DeepSeek V3 lineage (V3 launched 2024-12-26 as deepseek-chat; upgraded through V3-0324, V3.1, V3.1-Terminus, V3.2 by 2025-12-01). Deprecated 2026-04-24 when V4 launched; still callable until scheduled discontinuation 2026-07-24, currently routing to deepseek-v4-flash non-thinking mode (prices captured here reflect that routing, pre-2026-08-16 peak/off-peak rate change). DeepSeek's pricing page no longer publishes V3-era historical rates; standalone deepseek-v3 model_id was never exposed by the API. Re-verified 2026-07-02: discontinuation date (2026-07-24 15:59 UTC) and routing/pricing unchanged. Re-verified 2026-07-13: discontinuation date (2026-07-24 15:59 UTC) and routing/pricing unchanged; now 11 days out, expect this row to flip from deprecated to fully retired at the next refresh. Re-verified 2026-07-19: discontinuation date (2026-07-24 15:59 UTC) and routing/pricing unchanged against api-docs.deepseek.com/quick_start/pricing; now 5 days out. Re-verified 2026-07-27: model past discontinuation date (2026-07-24 15:59 UTC); no longer available via API. Kept for backward compatibility and referential resolution (replaced_by_model_id: deepseek-v4-flash). Re-verified 2026-08-11: still absent from api-docs.deepseek.com/quick_start/pricing (fully retired, confirms 2026-07-24 discontinuation held); no reversal. Re-verified 2026-09-02: still absent from api-docs.deepseek.com/quick_start/pricing and api-docs.deepseek.com/news/news260821; no reactivation. Prices frozen here at the pre-retirement snapshot, not restated for the 2026-08-16 peak/off-peak change since the model_id is no longer callable." }, { "provider": "DeepSeek", "provider_url": "https://www.deepseek.com", "model_id": "deepseek-reasoner", "display_name": "DeepSeek R1", "model_family": "DeepSeek R1", "context_window": 1048576, "max_output_tokens": 384000, "input_per_mtok_usd": "0.14", "output_per_mtok_usd": "0.28", "cache_read_per_mtok_usd": "0.0028", "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native"], "deprecated_at": "2026-04-24", "replaced_by_model_id": "deepseek-v4-flash", "last_verified": "2026-09-02", "last_changed_at": "2026-04-24", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://api-docs.deepseek.com/quick_start/pricing", "notes": "Canonical API model_id for the DeepSeek R1 reasoning lineage (R1 launched 2025-01-20 as deepseek-reasoner; R1-0528 update 2025-05-28). Deprecated 2026-04-24 when V4 launched; still callable until scheduled discontinuation 2026-07-24, currently routing to deepseek-v4-flash thinking mode (prices captured here reflect that routing, pre-2026-08-16 peak/off-peak rate change). Reasoning output is billed at the standard output rate (reasoning_tokens_billed: true). DeepSeek's pricing page no longer publishes R1-era historical rates; standalone deepseek-r1 model_id was never exposed by the API. Re-verified 2026-07-02: discontinuation date (2026-07-24 15:59 UTC) and routing/pricing unchanged. Re-verified 2026-07-13: discontinuation date (2026-07-24 15:59 UTC) and routing/pricing unchanged; now 11 days out, expect this row to flip from deprecated to fully retired at the next refresh. Re-verified 2026-07-19: discontinuation date (2026-07-24 15:59 UTC) and routing/pricing unchanged against api-docs.deepseek.com/quick_start/pricing; now 5 days out. Re-verified 2026-07-27: model past discontinuation date (2026-07-24 15:59 UTC); no longer available via API. Kept for backward compatibility and referential resolution (replaced_by_model_id: deepseek-v4-flash). Re-verified 2026-08-11: still absent from api-docs.deepseek.com/quick_start/pricing (fully retired, confirms 2026-07-24 discontinuation held); no reversal. Re-verified 2026-09-02: still absent from api-docs.deepseek.com/quick_start/pricing and api-docs.deepseek.com/news/news260821; no reactivation. Prices frozen here at the pre-retirement snapshot, not restated for the 2026-08-16 peak/off-peak change since the model_id is no longer callable." }, { "provider": "DeepSeek", "provider_url": "https://www.deepseek.com", "model_id": "deepseek-v4-flash-vision-exp", "display_name": "DeepSeek V4 Flash Vision (Experimental)", "model_family": "DeepSeek V4", "context_window": 1048576, "max_output_tokens": 384000, "input_per_mtok_usd": "0.44", "output_per_mtok_usd": "1.32", "cache_read_per_mtok_usd": "0.014", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-08-21", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://api-docs.deepseek.com/quick_start/pricing", "notes": "New row, added at the 2026-09-02 refresh. Launched 2026-08-21 (model version DeepSeek-V4-Flash-Vision-Exp) as an experimental vision-capable sibling to deepseek-v4-flash; matches V4 Flash on text capabilities (agents, reasoning, world knowledge) while adding multimodal input, targeted at agent benchmarks requiring visual understanding. Same peak-rate pricing as deepseek-v4-flash (identical MoE architecture): input (cache miss) $0.44, output $1.32, cache-hit input $0.014 per MTok; off-peak rates are half (cache hit $0.007, cache miss $0.22, output $0.66), applying all hours except 01:00-04:00 and 06:00-10:00 UTC Monday-Friday (peak hours, ~21% of the week) — schema has no time-of-day pricing axis so only the peak/list rate is encoded. Images are tokenized for billing at up to 384 tokens each per the launch announcement, using the standard V4 Flash input rate (no separate image-pricing field in schema; folded into input_per_mtok_usd). Supports both non-thinking and thinking (default) modes; thinking output billed at the output rate, so reasoning_tokens_billed: true. FIM completion is not supported (unlike plain deepseek-v4-flash, which supports it non-thinking-mode only). Concurrency limit 2500, matching deepseek-v4-flash. Works with Chat Completions, Anthropic-format, and Responses APIs; accepts base64, URL, or Files-API image references. Two-source verification: primary pricing table at api-docs.deepseek.com/quick_start/pricing plus the launch announcement at api-docs.deepseek.com/news/news260821 and the changelog at api-docs.deepseek.com/updates (both dated 2026-08-21)." }, { "provider": "Anthropic", "provider_url": "https://www.anthropic.com", "model_id": "claude-haiku-4-5-20251001", "display_name": "Claude Haiku 4.5", "featured": true, "model_family": "Claude 4", "knowledge_cutoff": "2025-02-28", "aliases": [ "claude-haiku-4-5", "anthropic.claude-haiku-4-5-20251001-v1:0", "claude-haiku-4-5@20251001" ], "context_window": 200000, "max_output_tokens": 64000, "input_per_mtok_usd": "1.0", "output_per_mtok_usd": "5.0", "cache_read_per_mtok_usd": "0.1", "cache_write_per_mtok_usd": "1.25", "batch_input_per_mtok_usd": "0.5", "batch_output_per_mtok_usd": "2.5", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "bedrock", "vertex"], "last_verified": "2026-09-02", "last_changed_at": "2025-10-01", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://platform.claude.com/docs/en/about-claude/pricing", "notes": "Anthropic's fastest model with near-frontier intelligence; positioned for high-volume agentic workloads. Cache hit is 0.1x base input ($0.10/MTok); 5-minute cache write is 1.25x ($1.25/MTok); 1-hour cache write is 2x ($2/MTok). Batch API discounts both input and output by 50%. Supports extended thinking; thinking output tokens are billed at the output rate. Reliable knowledge cutoff Feb 2025; training data cutoff Jul 2025." }, { "provider": "Anthropic", "provider_url": "https://www.anthropic.com", "model_id": "claude-sonnet-4-5-20250929", "display_name": "Claude Sonnet 4.5", "model_family": "Claude 4", "knowledge_cutoff": "2025-01-31", "aliases": [ "claude-sonnet-4-5", "anthropic.claude-sonnet-4-5-20250929-v1:0", "claude-sonnet-4-5@20250929" ], "context_window": 200000, "max_output_tokens": 64000, "input_per_mtok_usd": "3.0", "output_per_mtok_usd": "15.0", "cache_read_per_mtok_usd": "0.3", "cache_write_per_mtok_usd": "3.75", "batch_input_per_mtok_usd": "1.5", "batch_output_per_mtok_usd": "7.5", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "bedrock", "vertex"], "last_verified": "2026-09-02", "last_changed_at": "2025-09-29", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://platform.claude.com/docs/en/about-claude/pricing", "notes": "Legacy listing in Anthropic's models overview but still active. Pricing identical to Sonnet 4.6, but 200k context window (vs 1M on Sonnet 4.6). Cache hit is 0.1x base input ($0.30/MTok); 5-minute cache write is 1.25x ($3.75/MTok); 1-hour cache write is 2x ($6/MTok). Batch API discounts both input and output by 50%. Supports extended thinking; thinking output tokens are billed at the output rate. Reliable knowledge cutoff Jan 2025." }, { "provider": "Anthropic", "provider_url": "https://www.anthropic.com", "model_id": "claude-opus-4-5-20251101", "display_name": "Claude Opus 4.5", "model_family": "Claude 4", "knowledge_cutoff": "2025-05-31", "aliases": [ "claude-opus-4-5", "anthropic.claude-opus-4-5-20251101-v1:0", "claude-opus-4-5@20251101" ], "context_window": 200000, "max_output_tokens": 64000, "input_per_mtok_usd": "5.0", "output_per_mtok_usd": "25.0", "cache_read_per_mtok_usd": "0.5", "cache_write_per_mtok_usd": "6.25", "batch_input_per_mtok_usd": "2.5", "batch_output_per_mtok_usd": "12.5", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "bedrock", "vertex"], "last_verified": "2026-09-02", "last_changed_at": "2025-11-01", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://platform.claude.com/docs/en/about-claude/pricing", "notes": "Legacy listing in Anthropic's models overview but still active. Pricing identical to Opus 4.6/4.7, but 200k context window (vs 1M on 4.6/4.7) and 64k max output (vs 128k). Cache hit is 0.1x base input ($0.50/MTok); 5-minute cache write is 1.25x ($6.25/MTok); 1-hour cache write is 2x ($10/MTok). Batch API discounts both input and output by 50%. Supports extended thinking; thinking output tokens are billed at the output rate. Reliable knowledge cutoff May 2025." }, { "provider": "Anthropic", "provider_url": "https://www.anthropic.com", "model_id": "claude-3-7-sonnet-20250219", "display_name": "Claude Sonnet 3.7", "model_family": "Claude 3.7", "knowledge_cutoff": "2024-10-31", "aliases": [ "claude-3-7-sonnet-latest", "anthropic.claude-3-7-sonnet-20250219-v1:0", "claude-3-7-sonnet@20250219" ], "context_window": 200000, "max_output_tokens": 64000, "input_per_mtok_usd": "3.0", "output_per_mtok_usd": "15.0", "cache_read_per_mtok_usd": "0.3", "cache_write_per_mtok_usd": "3.75", "batch_input_per_mtok_usd": "1.5", "batch_output_per_mtok_usd": "7.5", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["bedrock", "vertex"], "deprecated_at": "2025-10-28", "replaced_by_model_id": "claude-sonnet-4-6", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2025-10-28", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://platform.claude.com/docs/en/about-claude/model-deprecations", "notes": "RETIRED on the Claude API on 2026-02-19; still available on Amazon Bedrock and Google Vertex AI under partner retirement schedules. Anthropic's first reasoning model with extended thinking; can output up to 64k tokens in thinking mode (128k with the output-128k-2025-02-19 beta header). Prices are no longer listed on Anthropic's current pricing page; values sourced from OpenRouter and pricepertoken.com (confidence: medium). Cache and batch pricing inferred from Anthropic's standard multipliers (1.25x 5-min write, 0.1x cache read, 0.5x batch)." }, { "provider": "Anthropic", "provider_url": "https://www.anthropic.com", "model_id": "claude-3-5-haiku-20241022", "display_name": "Claude Haiku 3.5", "model_family": "Claude 3.5", "knowledge_cutoff": "2024-07-31", "aliases": [ "claude-3-5-haiku-latest", "anthropic.claude-3-5-haiku-20241022-v1:0", "claude-3-5-haiku@20241022" ], "context_window": 200000, "max_output_tokens": 8192, "input_per_mtok_usd": "0.8", "output_per_mtok_usd": "4.0", "cache_read_per_mtok_usd": "0.08", "cache_write_per_mtok_usd": "1.0", "batch_input_per_mtok_usd": "0.4", "batch_output_per_mtok_usd": "2.0", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": false, "deployment_options": ["bedrock", "vertex"], "deprecated_at": "2025-12-19", "replaced_by_model_id": "claude-haiku-4-5-20251001", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2025-12-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://platform.claude.com/docs/en/about-claude/pricing", "notes": "RETIRED on the Claude API on 2026-02-19; still listed on Anthropic's pricing page as available on Amazon Bedrock and Google Vertex AI only. No extended-thinking support. max_output_tokens=8192 sourced from Anthropic legacy model card and OpenRouter (confidence: medium — not present in current docs). Cache hit is 0.1x base input ($0.08/MTok); 5-minute cache write is 1.25x ($1/MTok); 1-hour cache write is 2x ($1.60/MTok). Batch API discounts both input and output by 50%." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-5-mini", "display_name": "GPT-5 mini", "model_family": "GPT-5", "knowledge_cutoff": "2024-05-31", "context_window": 400000, "max_output_tokens": 128000, "input_per_mtok_usd": "0.25", "output_per_mtok_usd": "2.0", "cache_read_per_mtok_usd": "0.025", "batch_input_per_mtok_usd": "0.125", "batch_output_per_mtok_usd": "1.0", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "azure"], "deprecated_at": "2026-06-11", "replaced_by_model_id": "gpt-5.6-terra", "last_verified": "2026-09-02", "last_changed_at": "2026-05-17", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-5-mini", "notes": "Faster, more cost-efficient GPT-5 variant for low-latency, high-volume workloads. Cached input at 10% of base ($0.025/MTok). Batch API at flat 50% off input and output. Reasoning model with adjustable reasoning_effort; reasoning tokens are billed at the output rate. PDF input via the Files API; image input native. Knowledge cutoff May 2024. Deprecation announced 2026-06-11: dated snapshot gpt-5-mini-2025-08-07 scheduled for API removal 2026-12-11, recommended replacement gpt-5.4-mini (per developers.openai.com/api/docs/deprecations); bare 'gpt-5-mini' still serving as of 2026-07-19, prices unchanged. 2026-07-13 re-verify: model card now also points new low-latency workloads to GPT-5.6 Terra, but bare gpt-5-mini remains active. 2026-07-19 re-verify: model card live with prices/specs unchanged; no longer listed on the main pricing page (prices verified against the model card). 2026-08-11 re-verify: prices/specs unchanged against the model card; the deprecations table now redirects gpt-5-mini-2025-08-07 to gpt-5.6-terra (GPT-5.6 family's mid tier) instead of gpt-5.4-mini — replaced_by_model_id updated to match." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-5-nano", "display_name": "GPT-5 nano", "model_family": "GPT-5", "knowledge_cutoff": "2024-05-31", "context_window": 400000, "max_output_tokens": 128000, "input_per_mtok_usd": "0.05", "output_per_mtok_usd": "0.4", "cache_read_per_mtok_usd": "0.005", "batch_input_per_mtok_usd": "0.025", "batch_output_per_mtok_usd": "0.2", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": false, "deployment_options": ["native", "azure"], "deprecated_at": "2026-06-11", "replaced_by_model_id": "gpt-5.6-luna", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-05-17", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-5-nano", "notes": "Smallest GPT-5 variant; OpenAI model card lists 'Reasoning model: No' with 'Average' reasoning capability — reasoning_tokens_billed set to false on that basis (confidence: medium because other GPT-5 family members are reasoning models). Cached input at 10% of base ($0.005/MTok). Batch API at flat 50% off input and output. Knowledge cutoff May 2024. Deprecation announced 2026-06-11: dated snapshot gpt-5-nano-2025-08-07 scheduled for API removal 2026-12-11, recommended replacement gpt-5.4-nano (per developers.openai.com/api/docs/deprecations); bare 'gpt-5-nano' still serving as of 2026-07-19, prices unchanged. 2026-07-13 re-verify: model card now also points new speed/cost-sensitive workloads to GPT-5.6 Luna, but bare gpt-5-nano remains active. 2026-07-19 re-verify: model card live with prices/specs unchanged; no longer listed on the main pricing page (prices verified against the model card). 2026-08-11 re-verify: prices/specs unchanged against the model card; the deprecations table now redirects gpt-5-nano-2025-08-07 to gpt-5.6-luna (GPT-5.6 family's cheapest tier) instead of gpt-5.4-nano — replaced_by_model_id updated to match." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-4.1", "display_name": "GPT-4.1", "model_family": "GPT-4.1", "knowledge_cutoff": "2024-06-01", "context_window": 1047576, "max_output_tokens": 32768, "input_per_mtok_usd": "2.0", "output_per_mtok_usd": "8.0", "cache_read_per_mtok_usd": "0.5", "batch_input_per_mtok_usd": "1.0", "batch_output_per_mtok_usd": "4.0", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": false, "deployment_options": ["native", "azure"], "last_verified": "2026-09-02", "last_changed_at": "2026-05-17", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-4.1", "notes": "Non-reasoning flagship with a ~1M-token context window (1,047,576). Cached input at 25% of base ($0.50/MTok). Batch API at flat 50% off input and output. PDF input via the Files API; image input native. Knowledge cutoff June 2024. Not on the June 2026 deprecation announcement; remains active alongside the new GPT-5.4/5.5 family. Re-verified 2026-07-13: prices unchanged, still not on the deprecations table; remains active alongside the new GPT-5.6 family. Re-verified 2026-07-19: model card live with prices/specs unchanged, still not on the deprecations table; no longer listed on the main pricing page (prices verified against the model card). Re-verified 2026-08-11: prices unchanged against the main pricing table ($2.00/$0.50/$8.00, batch $1.00/$4.00), still not on the deprecations table." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-4o", "display_name": "GPT-4o", "model_family": "GPT-4o", "knowledge_cutoff": "2023-10-01", "context_window": 128000, "max_output_tokens": 16384, "input_per_mtok_usd": "2.5", "output_per_mtok_usd": "10.0", "cache_read_per_mtok_usd": "1.25", "batch_input_per_mtok_usd": "1.25", "batch_output_per_mtok_usd": "5.0", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": false, "deployment_options": ["native", "azure"], "deprecated_at": "2026-04-22", "replaced_by_model_id": "gpt-5.6-sol", "last_verified": "2026-09-02", "last_changed_at": "2026-04-22", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-4o", "notes": "Deprecated 2026-04-22 (dated alias gpt-4o-2024-05-13 scheduled for shutdown 2026-10-23 per OpenAI deprecations page); still serving on the API as of 2026-07-19. 2026-07-19 re-verify: model card live with prices/specs unchanged, deprecations table unchanged (gpt-4o-2024-05-13 -> gpt-5.5); no longer listed on the main pricing page (prices verified against the model card). Audio input/output are NOT supported on this model_id — they live on a sibling gpt-4o-audio-preview model card with separate pricing (audio input $40/MTok, audio output $80/MTok); text-mode prices captured here. Cached input at 50% of base ($1.25/MTok). Batch API at flat 50% off input and output. Knowledge cutoff Oct 2023. 2026-07-13 re-verify: the deprecations table now redirects the dated gpt-4o-2024-05-13 snapshot to gpt-5.5 (previously gpt-4.1); replaced_by_model_id updated to match the current table verbatim ('gpt-4o-2024-05-13 | gpt-5.5'). 2026-08-11 re-verify: prices/specs unchanged against the main pricing table; the deprecations table now redirects gpt-4o-2024-05-13 to gpt-5.6-sol instead of gpt-5.5 — replaced_by_model_id updated to match." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-4o-mini", "display_name": "GPT-4o mini", "model_family": "GPT-4o", "knowledge_cutoff": "2023-10-01", "context_window": 128000, "max_output_tokens": 16384, "input_per_mtok_usd": "0.15", "output_per_mtok_usd": "0.6", "cache_read_per_mtok_usd": "0.075", "batch_input_per_mtok_usd": "0.075", "batch_output_per_mtok_usd": "0.3", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": false, "deployment_options": ["native", "azure"], "last_verified": "2026-09-02", "last_changed_at": "2026-05-17", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-4o-mini", "notes": "Not on the April 2026 deprecation list; remains active. Re-verified 2026-07-19: model card live with prices/specs unchanged, still not on the deprecations table; no longer listed on the main pricing page (prices verified against the model card). Audio input/output are NOT supported on this model_id — they live on a sibling gpt-4o-mini-audio-preview model card; text-mode prices captured here. Cached input at 50% of base ($0.075/MTok). Batch API at flat 50% off input and output. Knowledge cutoff Oct 2023. Not on the June 2026 deprecation announcement either; remains active alongside the new GPT-5.4/5.5 family. Re-verified 2026-07-13: prices unchanged, listed as 'Default' tier, still not on the deprecations table. Re-verified 2026-08-11: prices unchanged against the main pricing table ($0.15/$0.075/$0.60, batch $0.075/$0.30), still not on the deprecations table." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "o3", "display_name": "OpenAI o3", "model_family": "o-series", "knowledge_cutoff": "2024-06-01", "context_window": 200000, "max_output_tokens": 100000, "input_per_mtok_usd": "2.0", "output_per_mtok_usd": "8.0", "cache_read_per_mtok_usd": "0.5", "batch_input_per_mtok_usd": "1.0", "batch_output_per_mtok_usd": "4.0", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "azure"], "deprecated_at": "2026-06-11", "replaced_by_model_id": "gpt-5.6-sol", "last_verified": "2026-09-02", "last_changed_at": "2026-05-17", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/o3", "notes": "Reasoning model for complex tasks; reasoning tokens are billed at the output rate. Cached input at 25% of base ($0.50/MTok). Batch API at flat 50% off input and output. Knowledge cutoff June 2024. Deprecation announced 2026-06-11: dated snapshot o3-2025-04-16 scheduled for API removal 2026-12-11, recommended replacement gpt-5.5 (per developers.openai.com/api/docs/deprecations); bare 'o3' still serving as of 2026-07-19, prices unchanged. 2026-07-19 re-verify: model card live with prices (incl. batch $1.00/$4.00) and specs unchanged, deprecations table unchanged; no longer listed on the main pricing page (prices verified against the model card). 2026-08-11 re-verify: prices unchanged against the main pricing table; the deprecations table now redirects o3-2025-04-16 to gpt-5.6-sol instead of gpt-5.5 — replaced_by_model_id updated to match." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "o4-mini", "display_name": "OpenAI o4-mini", "model_family": "o-series", "knowledge_cutoff": "2024-06-01", "context_window": 200000, "max_output_tokens": 100000, "input_per_mtok_usd": "1.1", "output_per_mtok_usd": "4.4", "cache_read_per_mtok_usd": "0.275", "batch_input_per_mtok_usd": "0.55", "batch_output_per_mtok_usd": "2.2", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "azure"], "deprecated_at": "2026-04-22", "replaced_by_model_id": "gpt-5.6-terra", "last_verified": "2026-09-02", "last_changed_at": "2026-04-22", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/o4-mini", "notes": "Fast, cost-efficient reasoning model; reasoning tokens are billed at the output rate. Deprecated 2026-04-22 (dated alias o4-mini-2025-04-16 scheduled for shutdown 2026-10-23 per OpenAI deprecations page); still serving as of 2026-07-19. 2026-07-19 re-verify: model card live with prices/specs unchanged, deprecations table unchanged (o4-mini-2025-04-16 -> gpt-5.4-mini); no longer listed on the main pricing page (prices verified against the model card). OpenAI's model card prose notes 'succeeded by GPT-5 mini'. Cached input at 25% of base ($0.275/MTok). Batch API at flat 50% off input and output. Knowledge cutoff June 2024. 2026-07-13 re-verify: the structured deprecations table lists the dated o4-mini-2025-04-16 snapshot's recommended replacement as gpt-5.4-mini (not gpt-5-mini); replaced_by_model_id updated to match that table verbatim. 2026-08-11 re-verify: prices unchanged; the deprecations table now redirects o4-mini-2025-04-16 to gpt-5.6-terra instead of gpt-5.4-mini — replaced_by_model_id updated to match." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "o1", "display_name": "OpenAI o1", "model_family": "o-series", "knowledge_cutoff": "2023-10-01", "context_window": 200000, "max_output_tokens": 100000, "input_per_mtok_usd": "15.0", "output_per_mtok_usd": "60.0", "cache_read_per_mtok_usd": "7.5", "batch_input_per_mtok_usd": "7.5", "batch_output_per_mtok_usd": "30.0", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "azure"], "deprecated_at": "2026-04-22", "replaced_by_model_id": "gpt-5.6-sol", "last_verified": "2026-09-02", "last_changed_at": "2026-04-22", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/o1", "notes": "First-generation reasoning model; reasoning tokens are billed at the output rate. Deprecated 2026-04-22 (dated alias o1-2024-12-17 scheduled for shutdown 2026-10-23 per OpenAI deprecations page); still serving as of 2026-07-19. 2026-07-19 re-verify: model card live with prices/specs unchanged, deprecations table unchanged (o1-2024-12-17 -> gpt-5.5); no longer listed on the main pricing page (prices verified against the model card). OpenAI's recommended replacement is gpt-5.5, which is now in-file; replaced_by_model_id updated from the interim 'o3' successor (o3 itself was deprecated 2026-06-11). Cached input at 50% of base ($7.50/MTok). Batch API at flat 50% off input and output. Knowledge cutoff Oct 2023. Re-verified 2026-07-13: deprecations table confirms shutdown date 2026-10-23 and replacement gpt-5.5 unchanged. 2026-08-11 re-verify: shutdown date 2026-10-23 unchanged, prices unchanged; deprecations table now redirects o1-2024-12-17 to gpt-5.6-sol instead of gpt-5.5 — replaced_by_model_id updated to match." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-5.5", "display_name": "GPT-5.5", "featured": true, "model_family": "GPT-5.5", "knowledge_cutoff": "2025-12-01", "context_window": 1050000, "max_output_tokens": 128000, "input_per_mtok_usd": "5.0", "output_per_mtok_usd": "30.0", "cache_read_per_mtok_usd": "0.5", "batch_input_per_mtok_usd": "2.5", "batch_output_per_mtok_usd": "15.0", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native", "azure"], "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-5.5", "notes": "Re-verified 2026-07-19 against the main pricing page: standard ($5.00/$0.50/$30.00) and batch ($2.50/$15.00) prices unchanged. OpenAI frontier reasoning model, released 2025-12-01; replaces gpt-5 and o3 as the recommended default per the 2026-06-11 deprecations announcement. Reasoning model with adjustable reasoning_effort (none/low/medium/high/xhigh); reasoning tokens billed at the output rate. Context window 1,050,000 (Azure Foundry lists 922k input / 128k output split within that total). Cached input at 10% of base ($0.50/MTok). Batch API at flat 50% off input and output. Prompts over 272k input tokens incur a 2x input / 1.5x output surcharge (not modeled as a separate pricing tier here). Azure availability confirmed via learn.microsoft.com/azure/ai-foundry model list (context/output match exactly). supports_pdf set to false: unlike gpt-5's page, the modality table on this model's docs page lists only Text/Image/Audio/Video rows with no PDF/file callout — treating conservatively pending explicit confirmation. Re-verified 2026-08-11: standard ($5.00/$0.50/$30.00) and batch ($2.50/$15.00) prices unchanged against the main pricing table; no longer OpenAI's top recommendation on the models overview page (superseded there by the GPT-5.6 family) but still active with no deprecation entry." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-5.4", "display_name": "GPT-5.4", "model_family": "GPT-5.4", "knowledge_cutoff": "2025-08-31", "context_window": 1050000, "max_output_tokens": 128000, "input_per_mtok_usd": "2.5", "output_per_mtok_usd": "15.0", "cache_read_per_mtok_usd": "0.25", "batch_input_per_mtok_usd": "1.25", "batch_output_per_mtok_usd": "7.5", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native", "azure"], "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-5.4", "notes": "Re-verified 2026-07-19 against the main pricing page: standard ($2.50/$0.25/$15.00) and batch ($1.25/$7.50) prices unchanged. Default GPT-5.4-class frontier model for professional work, snapshot gpt-5.4-2026-03-05. Reasoning model with adjustable reasoning_effort (none/low/medium/high/xhigh); reasoning tokens billed at the output rate. Cached input at 10% of base ($0.25/MTok). Batch API at flat 50% off input and output. Azure availability confirmed via learn.microsoft.com/azure/ai-foundry model list (context/output match exactly). supports_pdf set to false: this model's modality table explicitly lists Audio and Video as not supported with no PDF/file row — treating conservatively pending explicit confirmation. Re-verified 2026-08-11: standard ($2.50/$0.25/$15.00) and batch ($1.25/$7.50) prices unchanged against the main pricing table; remains active alongside GPT-5.6, no deprecation entry." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-5.4-mini", "display_name": "GPT-5.4 mini", "featured": true, "model_family": "GPT-5.4", "knowledge_cutoff": "2025-08-31", "context_window": 400000, "max_output_tokens": 128000, "input_per_mtok_usd": "0.75", "output_per_mtok_usd": "4.5", "cache_read_per_mtok_usd": "0.075", "batch_input_per_mtok_usd": "0.375", "batch_output_per_mtok_usd": "2.25", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native", "azure"], "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-5.4-mini", "notes": "Re-verified 2026-07-19 against the main pricing page: standard ($0.75/$0.075/$4.50) and batch ($0.375/$2.25) prices unchanged. Default mini model for low-latency, high-volume workloads; replaces gpt-5-mini per the 2026-06-11 deprecations announcement. Reasoning model with adjustable reasoning_effort (none/low/medium/high/xhigh); reasoning tokens billed at the output rate. Cached input at 10% of base ($0.075/MTok). Batch API at flat 50% off input and output. Azure availability confirmed via learn.microsoft.com/azure/ai-foundry model list (context/output match exactly). supports_pdf set to false: modality table lists no PDF/file row — treating conservatively pending explicit confirmation. Re-verified 2026-08-11: standard ($0.75/$0.075/$4.50) and batch ($0.375/$2.25) prices unchanged against the main pricing table; remains active alongside GPT-5.6, no deprecation entry." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-5.4-nano", "display_name": "GPT-5.4 nano", "model_family": "GPT-5.4", "knowledge_cutoff": "2025-08-31", "context_window": 400000, "max_output_tokens": 128000, "input_per_mtok_usd": "0.2", "output_per_mtok_usd": "1.25", "cache_read_per_mtok_usd": "0.02", "batch_input_per_mtok_usd": "0.1", "batch_output_per_mtok_usd": "0.625", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native", "azure"], "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-5.4-nano", "notes": "Re-verified 2026-07-19 against the main pricing page: standard ($0.20/$0.02/$1.25) and batch ($0.10/$0.625) prices unchanged. Cheapest GPT-5.4-class model for simple, high-volume tasks (classification, extraction, ranking, sub-agents); replaces gpt-5-nano per the 2026-06-11 deprecations announcement. Reasoning model with adjustable reasoning_effort (none/low/medium/high/xhigh); reasoning tokens billed at the output rate. Cached input at 10% of base ($0.02/MTok). Batch API at flat 50% off input and output. Azure availability confirmed via learn.microsoft.com/azure/ai-foundry model list (context/output match exactly). supports_pdf set to false: modality table lists no PDF/file row — treating conservatively pending explicit confirmation. Re-verified 2026-08-11: standard ($0.20/$0.02/$1.25) and batch ($0.10/$0.625) prices unchanged against the main pricing table; remains active alongside GPT-5.6, no deprecation entry." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-5.6-sol", "display_name": "GPT-5.6 Sol", "model_family": "GPT-5.6", "knowledge_cutoff": "2026-02-16", "aliases": ["gpt-5.6"], "context_window": 1050000, "max_output_tokens": 128000, "input_per_mtok_usd": "4.0", "output_per_mtok_usd": "20.0", "cache_read_per_mtok_usd": "0.4", "batch_input_per_mtok_usd": "2.0", "batch_output_per_mtok_usd": "10.0", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-09-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-5.6-sol", "notes": "Frontier-tier model of the GPT-5.6 family; model_id and price cross-checked across the OpenAI pricing page, the API models overview page, and this model's own docs page. Reasoning model with adjustable reasoning_effort; model card now explicitly lists reasoning token support with billing (confirmed 2026-07-19). Context window 1,050,000, max output 128k, knowledge cutoff 2026-02-16. 2026-07-19 re-verify: batch pricing now published on the main pricing page, flat 50% off. Alias 'gpt-5.6' added — now confirmed by two sources (models overview page and this model card: 'gpt-5.6 routes to GPT-5.6 Sol'). supports_pdf false and deployment_options limited to native (no Azure availability confirmed) — still treated conservatively. Re-verified 2026-08-11: standard ($5.00/$0.50/$30.00) and batch ($2.50/$15.00) prices unchanged; confirmed as the top-recommended flagship on the models overview page, and the recommended replacement target for most older deprecated OpenAI models per the deprecations table. PRICE CHANGE 2026-09-02: standard price cut from $5.00/$0.50/$30.00 to $4.00/$0.40/$20.00 (cache still 10% of base); batch cut correspondingly from $2.50/$15.00 to $2.00/$10.00 (still flat 50% off). Confirmed by two sources: the main pricing page and this model's own docs page, which now states promotional pricing ('20% input and 33% output reductions versus prior generations') available at least through 2026-11-21. Batch rate still not restated verbatim on the model card itself, so confidence stays medium. Deprecations table redirects unchanged (gpt-5-2025-08-07, gpt-4o-2024-05-13, o3-2025-04-16, o1-2024-12-17 all still -> gpt-5.6-sol)." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-5.6-terra", "display_name": "GPT-5.6 Terra", "model_family": "GPT-5.6", "knowledge_cutoff": "2026-02-16", "context_window": 1050000, "max_output_tokens": 128000, "input_per_mtok_usd": "2.0", "output_per_mtok_usd": "12.0", "cache_read_per_mtok_usd": "0.2", "batch_input_per_mtok_usd": "1.0", "batch_output_per_mtok_usd": "6.0", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-08-11", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-5.6-terra", "notes": "Mid-tier model of the GPT-5.6 family, described as balancing intelligence and cost; model_id and price cross-checked across the OpenAI pricing page, the API models overview page, and this model's own docs page. Reasoning model with adjustable reasoning_effort; model card lists reasoning token support (confirmed 2026-07-19). Context window 1,050,000, max output 128k, knowledge cutoff 2026-02-16. Cached input at 10% of base ($0.25/MTok). 2026-07-19 re-verify: base prices unchanged; batch pricing now published on the main pricing page (input $1.25 / output $7.50, flat 50% off) — added; batch rates appear on the pricing page only (model card quotes no batch rate), so confidence stays medium. supports_pdf false and deployment_options limited to native — still treated conservatively. PRICE CHANGE 2026-08-11: standard price cut from $2.50/$0.25/$15.00 to $2.00/$0.20/$12.00 (cache still 10% of base); batch cut correspondingly from $1.25/$7.50 to $1.00/$6.00 (still flat 50% off standard). Confirmed by two sources: the main pricing page and this model's own docs page both show the new $2/$0.2/$12 base rate. Batch rate still not stated on the model card itself (only implied by the flat-50%-off pattern and the pricing page's explicit $1.00/$6.00 batch column), so confidence stays medium." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-5.6-luna", "display_name": "GPT-5.6 Luna", "model_family": "GPT-5.6", "knowledge_cutoff": "2026-02-16", "context_window": 1050000, "max_output_tokens": 128000, "input_per_mtok_usd": "0.2", "output_per_mtok_usd": "1.2", "cache_read_per_mtok_usd": "0.02", "batch_input_per_mtok_usd": "0.1", "batch_output_per_mtok_usd": "0.6", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-08-11", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-5.6-luna", "notes": "Cheapest tier of the GPT-5.6 family, designed for cost-sensitive, high-volume workloads; model_id and price cross-checked across the OpenAI pricing page, the API models overview page, and this model's own docs page. Reasoning model with adjustable reasoning_effort; model card lists reasoning token support (confirmed 2026-07-19). Context window 1,050,000, max output 128k, knowledge cutoff 2026-02-16. Cached input at 10% of base ($0.10/MTok). 2026-07-19 re-verify: base prices unchanged; batch pricing now published on the main pricing page (input $0.50 / output $3.00, flat 50% off) — added; batch rates appear on the pricing page only (model card quotes no batch rate), so confidence stays medium. supports_pdf false and deployment_options limited to native — still treated conservatively. PRICE CHANGE 2026-08-11: standard price cut from $1.00/$0.10/$6.00 to $0.20/$0.02/$1.20 (cache still 10% of base, an 80% reduction); batch cut correspondingly from $0.50/$3.00 to $0.10/$0.60 (still flat 50% off standard). Confirmed by two sources: the main pricing page and this model's own docs page both show the new $0.2/$0.02/$1.2 base rate — this now undercuts even gpt-5.4-nano's $0.20/$0.02/$1.25. Batch rate still not stated on the model card itself (only implied by the flat-50%-off pattern and the pricing page's explicit $0.10/$0.60 batch column), so confidence stays medium." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-5.5-pro", "display_name": "GPT-5.5 Pro", "model_family": "GPT-5.5", "knowledge_cutoff": "2025-12-01", "context_window": 1050000, "max_output_tokens": 128000, "input_per_mtok_usd": "30.0", "output_per_mtok_usd": "180.0", "batch_input_per_mtok_usd": "30.0", "batch_output_per_mtok_usd": "180.0", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-13", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-5.5-pro", "notes": "2026-07-19 re-verify: base prices unchanged; SOURCE CONFLICT on batch — the main pricing page now shows Batch input $15.00 / output $90.00 (50% off), but this model's docs page still states Batch runs at the standard $30/$180 rate; per the two-source rule the batch fields are left at $30/$180 (matching the model card and the prior verification) pending agreement between sources — re-check next refresh. Highest-effort reasoning tier of GPT-5.5, 'uses more compute to think harder'; model_id and price ($30/$180 per MTok) cross-checked across the OpenAI pricing page and this model's own docs page. No cache_read_per_mtok_usd field: the docs page states this model 'does not offer a cached input discount'. Batch API documented as the same rate as standard (no batch discount for this tier) — unusual versus sibling rows, recorded as given rather than assumed. Prompts over 272k input tokens incur a surcharge per the shared GPT-5.5 pricing note (not modeled as a separate pricing tier here, consistent with the gpt-5.5 row). supports_pdf false and deployment_options limited to native — treated conservatively pending explicit confirmation. confidence: medium because tool/structured-output flags were read from an AI-summarized fetch rather than raw page content. 2026-08-11 re-verify: base prices ($30/$180) unchanged, confirmed on both sources; SOURCE CONFLICT UNRESOLVED — pricing page still shows Batch $15.00/$90.00 while the model card still says batch runs at the standard $30/$180 rate; batch fields left unchanged at $30/$180 per the two-source rule, matching the model card — re-check next refresh." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-5.4-pro", "display_name": "GPT-5.4 Pro", "model_family": "GPT-5.4", "knowledge_cutoff": "2025-08-31", "context_window": 1050000, "max_output_tokens": 128000, "input_per_mtok_usd": "30.0", "output_per_mtok_usd": "180.0", "batch_input_per_mtok_usd": "30.0", "batch_output_per_mtok_usd": "180.0", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": false, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-13", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-5.4-pro", "notes": "2026-07-19 re-verify: base prices unchanged; SOURCE CONFLICT on batch — the main pricing page now shows Batch input $15.00 / output $90.00 (50% off), but this model's docs page still quotes Batch at the standard $30/$180 rate; per the two-source rule the batch fields are left at $30/$180 (matching the model card and the prior verification) pending agreement between sources — re-check next refresh. Highest-effort reasoning tier of GPT-5.4, 'uses more compute to think harder'; model_id and price ($30/$180 per MTok) cross-checked across the OpenAI pricing page and this model's own docs page. structured_output recorded as false: re-confirmed 2026-07-19 — the docs page still explicitly states structured outputs are 'Not supported'; genuine outlier, no longer treated as a fetch artifact. No cache_read_per_mtok_usd field: no cache discount documented for this tier. Batch API documented as the same rate as standard (no batch discount). Prompts over 272k input tokens incur a surcharge per the shared GPT-5.4 pricing note (not modeled as a separate pricing tier here, consistent with the gpt-5.4 row). supports_pdf false and deployment_options limited to native — treated conservatively pending explicit confirmation. confidence: medium. 2026-08-11 re-verify: base prices ($30/$180) unchanged, confirmed on both sources; SOURCE CONFLICT UNRESOLVED — pricing page still shows Batch $15.00/$90.00 while the model card still quotes the standard $30/$180 rate; batch fields left unchanged at $30/$180 per the two-source rule — re-check next refresh." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-5.3-codex", "display_name": "GPT-5.3-Codex", "model_family": "Codex", "knowledge_cutoff": "2025-08-31", "context_window": 400000, "max_output_tokens": 128000, "input_per_mtok_usd": "1.75", "output_per_mtok_usd": "14.0", "cache_read_per_mtok_usd": "0.175", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-07-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/models/gpt-5.3-codex", "notes": "New row this refresh — agentic coding model ('the most capable agentic coding model to date'), optimized for Codex environments; model_id and price ($1.75 input / $0.175 cached / $14.00 output per MTok) cross-checked across the OpenAI main pricing page and this model's own docs page (two-source rule met). Reasoning model with low/medium/high/xhigh reasoning_effort; reasoning tokens billed at the output rate. Context window 400,000, max output 128k, knowledge cutoff 2025-08-31. Cached input at 10% of base. Batch endpoint supported per the model card but no batch rate published on either source — batch fields omitted rather than assumed. Fine-tuning and predicted outputs not supported. supports_pdf false and deployment_options limited to native — treated conservatively pending explicit confirmation, matching how sibling rows were first added. Re-verified 2026-08-11: still active with prices ($1.75/$0.175/$14.00) unchanged, confirmed directly via its own model card (not present in the main pricing table's flagship-tier rows this pass, so the model card served as primary confirmation); no deprecation notice; model card notes GPT-5.3-Codex and GPT-5.2-Codex share identical pricing, no successor guidance given." }, { "provider": "Google", "provider_url": "https://deepmind.google", "model_id": "gemini-2.5-flash", "display_name": "Gemini 2.5 Flash", "model_family": "Gemini 2.5", "knowledge_cutoff": "2025-01-31", "context_window": 1048576, "max_output_tokens": 65536, "input_per_mtok_usd": "0.30", "output_per_mtok_usd": "2.50", "audio_input_per_mtok_usd": "1.00", "cache_read_per_mtok_usd": "0.03", "cache_storage_per_mtok_per_hour_usd": "1.00", "batch_input_per_mtok_usd": "0.15", "batch_output_per_mtok_usd": "1.25", "modalities": { "input": ["text", "image", "audio", "video"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": true, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "vertex"], "confidence": "high", "last_verified": "2026-09-02", "last_changed_at": "2026-05-18", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://ai.google.dev/gemini-api/docs/pricing", "notes": "Hybrid reasoning model with dynamic thinking on by default; thinking can be disabled via thinkingBudget=0. When thinking is on, response pricing is the sum of output and thinking tokens (both billed at the output rate). Audio input is priced separately (captured in audio_input_per_mtok_usd); cached audio input is $0.10/MTok vs $0.03/MTok for text/image/video. Batch Mode at flat 50% off; audio batch input is $0.50/MTok. No long-context tier. Context window and max_output_tokens confirmed via ai.google.dev/gemini-api/docs/models/gemini-2.5-flash. Knowledge cutoff January 2025. Explicit context caching storage is captured in cache_storage_per_mtok_per_hour_usd. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. Cross-verified prompt/completion price against openrouter.ai/google/gemini-2.5-flash. Re-verified 2026-07-02: prices unchanged. Re-verified 2026-07-13: prices unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-07-19: prices unchanged against ai.google.dev/pricing; still listed stable on ai.google.dev/gemini-api/docs/models. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models." }, { "provider": "Google", "provider_url": "https://deepmind.google", "model_id": "gemini-2.5-flash-lite", "display_name": "Gemini 2.5 Flash-Lite", "model_family": "Gemini 2.5", "knowledge_cutoff": "2025-01-31", "context_window": 1048576, "max_output_tokens": 65536, "input_per_mtok_usd": "0.10", "output_per_mtok_usd": "0.40", "audio_input_per_mtok_usd": "0.30", "cache_read_per_mtok_usd": "0.01", "cache_storage_per_mtok_per_hour_usd": "1.00", "batch_input_per_mtok_usd": "0.05", "batch_output_per_mtok_usd": "0.20", "modalities": { "input": ["text", "image", "audio", "video"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": true, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "vertex"], "confidence": "high", "last_verified": "2026-09-02", "last_changed_at": "2026-05-18", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://ai.google.dev/gemini-api/docs/pricing", "notes": "Hybrid reasoning model; thinking is OFF by default (unlike 2.5 Flash/Pro) but can be enabled by setting thinkingBudget. When thinking is enabled, response pricing is the sum of output and thinking tokens at the output rate, so reasoning_tokens_billed is true. Audio input is priced separately (captured in audio_input_per_mtok_usd); cached audio input is $0.03/MTok vs $0.01/MTok for text/image/video. Batch Mode at flat 50% off; audio batch input is $0.15/MTok. No long-context tier. Context window and max_output_tokens confirmed via ai.google.dev/gemini-api/docs/models/gemini-2.5-flash-lite. Knowledge cutoff January 2025. Explicit context caching storage is captured in cache_storage_per_mtok_per_hour_usd. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. Cross-verified prompt/completion price against openrouter.ai/google/gemini-2.5-flash-lite. Re-verified 2026-07-02: prices unchanged. Re-verified 2026-07-13: prices unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-07-19: prices unchanged against ai.google.dev/pricing; still listed stable on ai.google.dev/gemini-api/docs/models. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models." }, { "provider": "Google", "provider_url": "https://deepmind.google", "model_id": "gemini-2.0-flash", "display_name": "Gemini 2.0 Flash", "model_family": "Gemini 2.0", "knowledge_cutoff": "2024-08-31", "context_window": 1048576, "max_output_tokens": 8192, "input_per_mtok_usd": "0.10", "output_per_mtok_usd": "0.40", "audio_input_per_mtok_usd": "0.70", "cache_read_per_mtok_usd": "0.025", "cache_storage_per_mtok_per_hour_usd": "1.00", "batch_input_per_mtok_usd": "0.05", "batch_output_per_mtok_usd": "0.20", "modalities": { "input": ["text", "image", "audio", "video"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": true, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native", "vertex"], "deprecated_at": "2026-06-01", "replaced_by_model_id": "gemini-3.5-flash", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-05-18", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://ai.google.dev/gemini-api/docs/pricing", "notes": "Deprecated; shut down 2026-06-01 per the model card's verbatim notice: \"Gemini 2.0 Flash is deprecated and has been shut down June 1, 2026. Migrate to Gemini 3.5 Flash to avoid service disruption.\" replaced_by_model_id updated 2026-07-02 to gemini-3.5-flash (Google's documented migration target), now that it is in-dataset; previously pointed at gemini-2.5-flash as a placeholder in-file successor. Standard production 2.0 Flash does not support thinking (thinking exists only on Gemini 2.5+ and 3 series per ai.google.dev/gemini-api/docs/thinking); reasoning_tokens_billed=false. Audio input is priced separately (captured in audio_input_per_mtok_usd); cached audio input is $0.175/MTok vs $0.025/MTok for text/image/video. Batch Mode at flat 50% off. No long-context tier. supports_pdf=false since the 2.0 Flash model card lists supported inputs as audio/images/video/text (PDF not enumerated). Confidence medium because Vertex AI's published pricing for the same model name differs ($0.15 input / $0.60 output) from AI Studio's $0.10/$0.40; AI Studio primary value retained per spec, and OpenRouter (openrouter.ai/google/gemini-2.0-flash-001) cross-confirms $0.10/$0.40. Explicit context caching storage is captured in cache_storage_per_mtok_per_hour_usd. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. Prices unchanged as of 2026-07-02 re-verification. Re-verified 2026-07-13: deprecation notice and shutdown date still confirmed on ai.google.dev/pricing; no new information. Re-verified 2026-07-19: deprecation notice (shut down June 1, 2026) still shown on ai.google.dev/pricing; model marked shut down on ai.google.dev/gemini-api/docs/models. Re-verified 2026-08-11: deprecation notice (shut down June 1, 2026) still shown on ai.google.dev/pricing; model remains shut down per ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models." }, { "provider": "Google", "provider_url": "https://deepmind.google", "model_id": "gemini-3.5-flash", "display_name": "Gemini 3.5 Flash", "model_family": "Gemini 3.5", "knowledge_cutoff": "2025-01-31", "context_window": 1048576, "max_output_tokens": 65536, "input_per_mtok_usd": "1.50", "output_per_mtok_usd": "9.00", "cache_read_per_mtok_usd": "0.15", "cache_storage_per_mtok_per_hour_usd": "1.00", "batch_input_per_mtok_usd": "0.75", "batch_output_per_mtok_usd": "4.50", "modalities": { "input": ["text", "image", "audio", "video"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": true, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native"], "confidence": "high", "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://ai.google.dev/pricing", "notes": "New GA/stable row added 2026-07-02; this is Google's documented migration target for the retired Gemini 2.0 Flash (see that row's notes). Thinking is on by default (medium level) and configurable (minimal/low/medium/high); when on, response pricing is the sum of output and thinking tokens billed at the output rate, so reasoning_tokens_billed=true. No >200k-token pricing tier is published for this model (unlike Gemini 3.1 Pro Preview). Context window and max_output_tokens confirmed via ai.google.dev/gemini-api/docs/models/gemini-3.5-flash. Knowledge cutoff January 2025, unchanged from the Gemini 2.5 family. Explicit context caching storage is captured in cache_storage_per_mtok_per_hour_usd. Google Search grounding is billed separately (5,000 prompts/month free, shared across Gemini 3 models, then $14/1,000 queries); not captured as a structured field since sibling 2.5-series rows in this file omit the same grounding fee for consistency. Vertex AI availability not confirmed as of this pass; deployment_options limited to native pending verification. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. Re-verified 2026-07-13: prices unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-07-19: prices unchanged against ai.google.dev/pricing; still listed stable on ai.google.dev/gemini-api/docs/models. Pricing page now also lists Flex (same rates as Batch) and Priority ($2.70/$16.20) service tiers; no schema fields for those, standard/batch rates unchanged. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models." }, { "provider": "Google", "provider_url": "https://deepmind.google", "model_id": "gemini-3.1-pro-preview", "display_name": "Gemini 3.1 Pro Preview", "model_family": "Gemini 3.1", "knowledge_cutoff": "2025-01-31", "context_window": 1048576, "max_output_tokens": 65536, "input_per_mtok_usd": "2.00", "output_per_mtok_usd": "12.00", "cache_read_per_mtok_usd": "0.20", "cache_storage_per_mtok_per_hour_usd": "4.50", "batch_input_per_mtok_usd": "1.00", "batch_output_per_mtok_usd": "6.00", "pricing_tiers": [ { "threshold_tokens": 200000, "input_per_mtok_usd": "4.00", "output_per_mtok_usd": "18.00", "cache_read_per_mtok_usd": "0.40" } ], "modalities": { "input": ["text", "image", "audio", "video"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": true, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native"], "confidence": "high", "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://ai.google.dev/pricing", "notes": "New preview row added 2026-07-02; not GA. Optimized for software engineering and agentic workflows; a specialized gemini-3.1-pro-preview-customtools variant exists but is not tracked as a separate row. Input/output/cache-read pricing tiers at >200k tokens (>200k rate captured in pricing_tiers[0]); batch pricing also tiers ($2.00/$9.00 input/output above 200k) but batch has no tiered field in this schema, so only the <=200k batch rate is captured. Thinking is on by default at high level (configurable low/medium/high); thinking tokens billed at the output rate, so reasoning_tokens_billed=true. Context window and max_output_tokens confirmed via ai.google.dev/gemini-api/docs/models/gemini-3.1-pro-preview. Knowledge cutoff January 2025, unchanged from the Gemini 2.5 family. Google Search grounding billed separately (5,000 prompts/month free, shared across Gemini 3 models, then $14/1,000 queries); omitted as a structured field for consistency with sibling rows. Vertex AI availability not confirmed as of this pass; deployment_options limited to native pending verification. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. Re-verified 2026-07-13: prices, tiers, and status (still preview, not GA) unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-07-19: prices, tiers, and status (still preview) unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models." }, { "provider": "Google", "provider_url": "https://deepmind.google", "model_id": "gemini-3-flash-preview", "display_name": "Gemini 3 Flash Preview", "model_family": "Gemini 3", "knowledge_cutoff": "2025-01-31", "context_window": 1048576, "max_output_tokens": 65536, "input_per_mtok_usd": "0.50", "output_per_mtok_usd": "3.00", "audio_input_per_mtok_usd": "1.00", "cache_read_per_mtok_usd": "0.05", "cache_storage_per_mtok_per_hour_usd": "1.00", "batch_input_per_mtok_usd": "0.25", "batch_output_per_mtok_usd": "1.50", "modalities": { "input": ["text", "image", "audio", "video"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": true, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native"], "confidence": "high", "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://ai.google.dev/pricing", "notes": "New preview row added 2026-07-02; not GA. Audio input is priced separately (captured in audio_input_per_mtok_usd); cached audio input is $0.10/MTok vs $0.05/MTok for text/image/video. No >200k-token pricing tier published. Thinking is on by default at high level (configurable minimal/low/medium/high); thinking tokens billed at the output rate, so reasoning_tokens_billed=true. Context window and max_output_tokens confirmed via ai.google.dev/gemini-api/docs/models/gemini-3-flash-preview. Knowledge cutoff January 2025, unchanged from the Gemini 2.5 family. Google Search grounding billed separately (5,000 prompts/month free, shared across Gemini 3 models, then $14/1,000 queries); omitted as a structured field for consistency with sibling rows. Vertex AI availability not confirmed as of this pass; deployment_options limited to native pending verification. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. Re-verified 2026-07-13: prices and status (still preview, not GA) unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-07-19: prices and status (still preview) unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models." }, { "provider": "Google", "provider_url": "https://deepmind.google", "model_id": "gemini-3.1-flash-lite", "display_name": "Gemini 3.1 Flash-Lite", "model_family": "Gemini 3.1", "knowledge_cutoff": "2025-01-31", "context_window": 1048576, "max_output_tokens": 65536, "input_per_mtok_usd": "0.25", "output_per_mtok_usd": "1.50", "audio_input_per_mtok_usd": "0.50", "cache_read_per_mtok_usd": "0.025", "cache_storage_per_mtok_per_hour_usd": "1.00", "batch_input_per_mtok_usd": "0.125", "batch_output_per_mtok_usd": "0.75", "modalities": { "input": ["text", "image", "audio", "video"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": true, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native"], "confidence": "high", "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://ai.google.dev/pricing", "notes": "New GA/stable row added 2026-07-02. Audio input is priced separately (captured in audio_input_per_mtok_usd); cached audio input is $0.05/MTok vs $0.025/MTok for text/image/video. No >200k-token pricing tier published. Thinking is minimal by default (options: minimal or high); thinking tokens billed at the output rate, so reasoning_tokens_billed=true. Context window and max_output_tokens confirmed via ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite. Knowledge cutoff January 2025, unchanged from the Gemini 2.5 family. Google Search grounding billed separately (5,000 prompts/month free, shared across Gemini 3 models, then $14/1,000 queries); omitted as a structured field for consistency with sibling rows. Vertex AI availability not confirmed as of this pass; deployment_options limited to native pending verification. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. Re-verified 2026-07-13: prices unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-07-19: prices unchanged against ai.google.dev/pricing; still listed stable on ai.google.dev/gemini-api/docs/models. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models." }, { "provider": "Meta", "provider_url": "https://www.llama.com", "model_id": "llama-4-maverick", "display_name": "Llama 4 Maverick", "featured": true, "model_family": "Llama 4", "knowledge_cutoff": "2024-08-31", "aliases": [ "meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8", "meta-llama/llama-4-maverick" ], "context_window": 1048576, "max_output_tokens": 8192, "input_per_mtok_usd": "0.27", "output_per_mtok_usd": "0.85", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["together"], "aggregators": ["openrouter"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-05-18", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct", "notes": "Multi-host pricing re-verified 2026-07-19: Together still $0.27/$0.85 per MTok (meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8) per its model page (together.ai/models/llama-4-maverick), the row's structured price as the sole confirmed direct-host rate; unchanged since 2026-05-18. Confidence lowered to medium this pass: Together's model page is the only live source for the serverless rate — together.ai/pricing and the docs.together.ai serverless catalog no longer list Maverick (fine-tuning tables only), and OpenRouter's endpoint list routes no Together endpoint. Groq formally deprecated Maverick (announced 2026-02-20, shutdown 2026-03-09 per console.groq.com/docs/deprecations); Fireworks still does not offer Maverick on serverless (on-demand deployments only). Bedrock pricing tables did not render on this pass; omitted per the single-primary-source rule for deployment_options[]. OpenRouter aggregator still routes at $0.20/$0.80 (informational only; structured input/output stay at the lowest direct-host price per the PR4 convention). Context window is Meta's published 1M (1048576 tokens); OpenRouter advertises 1.05M but the HuggingFace model card spec is 1M. max_output_tokens not published on the model card; defaulted to 8192. 17B activated / 400B total MoE with 128 experts. Re-verified 2026-08-11: Together's model page still the sole live source at $0.27/$0.85, unchanged; together.ai/pricing still omits Maverick from the serverless table (fine-tuning only). Fireworks' own model page confirms serverless still not supported (on-demand only). huggingface.co/meta-llama org page shows no new Llama 4 Maverick variant. Noted for the record: Meta Superintelligence Labs is shipping a separate closed-weights \"Muse\" line (Muse Spark et al.) via its own Meta Model API — unrelated to the open-weights Llama family this row tracks; out of scope here, would need its own provider row and primary source if ever added. Re-verified 2026-09-02: Together's model page (together.ai/models/llama-4-maverick) still the sole live source at $0.27/$0.85, unchanged; together.ai/pricing still omits Maverick from the serverless table. Fireworks' model page reconfirms serverless still not supported. console.groq.com/docs/deprecations still shows Maverick deprecated 2026-03-09 in favor of openai/gpt-oss-120b, no change. huggingface.co/meta-llama org page shows no new Llama 4 Maverick variant." }, { "provider": "Meta", "provider_url": "https://www.llama.com", "model_id": "llama-4-scout", "display_name": "Llama 4 Scout", "model_family": "Llama 4", "knowledge_cutoff": "2024-08-31", "aliases": [ "meta-llama/Llama-4-Scout-17B-16E-Instruct", "meta-llama/llama-4-scout-17b-16e-instruct", "accounts/fireworks/models/llama4-scout-instruct-basic", "meta-llama/llama-4-scout" ], "context_window": 10485760, "max_output_tokens": 8192, "input_per_mtok_usd": "0.18", "output_per_mtok_usd": "0.59", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["together"], "aggregators": ["openrouter"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct", "notes": "Price changed 2026-07-19: Groq retired Llama 4 Scout (announced 2026-06-17, shutdown 2026-07-17 per console.groq.com/docs/deprecations; removed from groq.com/pricing and console.groq.com/docs/models), so its $0.11/$0.34 rate no longer exists. Fireworks' model page now states serverless is not supported for accounts/fireworks/models/llama4-scout-instruct-basic, so Fireworks is also removed from deployment_options. Together is now the sole confirmed direct host at $0.18/$0.59 per MTok (meta-llama/Llama-4-Scout-17B-16E-Instruct, per together.ai/models/llama-4-scout; same Together rate as verified 2026-05-18 and 2026-07-13), which becomes the row's structured price. Confidence lowered to medium: Together's model page is the only live source for that rate — together.ai/pricing and the docs.together.ai serverless catalog no longer list Scout, and OpenRouter routes no Together endpoint (its Groq endpoint still listed at $0.11/$0.34 is stale post-shutdown). Bedrock pricing tables did not render on this pass; omitted per the single-primary-source rule for deployment_options[]. OpenRouter aggregator still routes at $0.10/$0.30 (informational only; structured input/output stay at the lowest direct-host price per the PR4 convention). Retired-host aliases (Groq, Fireworks) retained for lookup. Context window is Meta's published 10M (10485760 tokens); hosts cap below Meta's spec. max_output_tokens not published on the model card; defaulted to 8192. 17B activated / 109B total MoE with 16 experts. Re-verified 2026-08-11: Together's model page still the sole live source at $0.18/$0.59, unchanged; Fireworks' own model page reconfirms serverless still not supported for llama4-scout-instruct-basic (on-demand only). No new Scout variant on huggingface.co/meta-llama. Re-verified 2026-09-02: Together's model page still the sole live source at $0.18/$0.59, unchanged; Fireworks' model page reconfirms serverless still not supported. console.groq.com/docs/deprecations still shows Scout deprecated 2026-07-17 in favor of openai/gpt-oss-120b or qwen/qwen3.6-27b, no change. No new Scout variant on huggingface.co/meta-llama." }, { "provider": "Meta", "provider_url": "https://www.llama.com", "model_id": "llama-3.3-70b", "display_name": "Llama 3.3 70B Instruct", "model_family": "Llama 3.3", "knowledge_cutoff": "2023-12-31", "aliases": [ "meta-llama/Llama-3.3-70B-Instruct", "meta-llama/Llama-3.3-70B-Instruct-Turbo", "llama-3.3-70b-versatile", "meta-llama/llama-3.3-70b-instruct" ], "context_window": 131072, "max_output_tokens": 8192, "input_per_mtok_usd": "1.04", "output_per_mtok_usd": "1.04", "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["together"], "aggregators": ["openrouter"], "confidence": "high", "last_verified": "2026-09-02", "last_changed_at": "2026-09-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct", "notes": "Multi-host pricing re-verified 2026-07-19: Together still $1.04/$1.04 per MTok for Llama 3.3 70B (meta-llama/Llama-3.3-70B-Instruct-Turbo, confirmed on both together.ai/pricing and the docs.together.ai serverless catalog; unchanged since the 2026-07-02 jump from $0.88/$0.88); Groq still $0.59/$0.79 (llama-3.3-70b-versatile, confirmed on both groq.com/pricing and console.groq.com/docs/models) remains the lowest direct-host rate and stays the row's structured price. Fireworks still publishes $0.90 input on accounts/fireworks/models/llama-v3p3-70b-instruct and its own model page states serverless is not supported for this model, so Fireworks remains omitted from deployment_options. OpenRouter aggregator still routes at $0.10/$0.32 (informational only; structured input/output stay at the lowest direct-host price per the PR4 convention) and is structured as aggregators[\"openrouter\"]. Context window 128K per Meta's spec (131072 tokens). Text-only; no vision. max_output_tokens not published on the model card; defaulted to 8192. Knowledge cutoff December 2023 per model card. Re-verified 2026-08-11: Groq (console.groq.com/docs/model/llama-3.3-70b-versatile) still $0.59/$0.79 and still the cheapest direct host, no deprecation announced per console.groq.com/docs/deprecations; Together (docs.together.ai serverless catalog) still $1.04/$1.04; Fireworks' own model page reconfirms serverless still not supported for llama-v3p3-70b-instruct. No new 3.3-family variant found. Price changed 2026-09-02: Groq moved llama-3.3-70b-versatile to Enterprise-only. console.groq.com/docs/models now lists its pricing and rate limits as \"Contact Sales\" (no public PAYG rate); console.groq.com/docs/deprecations confirms deprecation effective 2026-08-16, stating it \"applies to free and developer-tier usage; enterprise customers with a committed-spend contract are not affected,\" and recommends migrating to openai/gpt-oss-120b or qwen/qwen3.6-27b. groq.com/pricing still redirects to the marketing homepage with no model pricing table. Groq removed from deployment_options accordingly. Together is now the sole confirmed direct host — $1.04/$1.04 per MTok, confirmed on both together.ai/pricing and together.ai/models/llama-3.3-70b, unchanged since 2026-07-02 — and becomes the row's structured price; the $0.59/$0.79 -> $1.04/$1.04 move is entirely the loss of the cheaper Groq host, not a Together repricing. Fireworks' model page (fireworks.ai/models/fireworks/llama-v3p3-70b-instruct) reconfirms serverless still not supported for llama-v3p3-70b-instruct (states \"Serverless: Not supported\" despite a $0.90 shared-endpoint figure elsewhere on the page). Retired-host alias llama-3.3-70b-versatile retained for lookup. No new 3.3-family variant found on huggingface.co/meta-llama." }, { "provider": "Mistral", "provider_url": "https://mistral.ai", "model_id": "mistral-large-2411", "display_name": "Mistral Large 2 (24.11)", "model_family": "Mistral Large", "aliases": ["mistral-large-2407", "mistral.mistral-large-2407-v1:0"], "context_window": 131072, "max_output_tokens": 8192, "input_per_mtok_usd": "2.00", "output_per_mtok_usd": "6.00", "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native", "bedrock", "vertex", "azure"], "deprecated_at": "2026-02-27", "replaced_by_model_id": "mistral-medium-3-5", "last_verified": "2026-09-02", "last_changed_at": "2024-11-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://mistral.ai/news/mistral-large-2407", "notes": "La Plateforme rates ($2.00 / $6.00 per MTok) are the row's structured price. Multi-host availability: Bedrock (mistral.mistral-large-2407-v1:0 in us-west-2), Vertex AI, Azure AI Foundry, IBM watsonx. Bedrock published the 24.07 build only, not 24.11. Deprecated on La Plateforme 2026-02-27; retirement 2026-05-31 per Mistral's legacy table has now passed (re-verified 2026-07-13 via docs.mistral.ai/getting-started/models/models_overview) so the model_id should be treated as fully retired/non-callable, not merely deprecated. `mistral-large-latest` moved to Mistral Large 3 (mistral-large-2512, added to this dataset) so the alias was removed from this row to avoid a duplicate. Mistral's own deprecation table lists Mistral Medium 3.5 (mistral-medium-3-5) as the recommended alternative, not Large 3; replaced_by_model_id follows the vendor's stated alternative. Text-only; no vision (Pixtral Large was the multimodal sibling, also retired). max_output_tokens not published on the model card; defaulted to 8192. Batch API is a 50% discount where available but per-model availability is not confirmed from a single primary source on this date, so batch fields are unset. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02 via mistral.ai/pricing/api and docs.mistral.ai/getting-started/models/models_overview: still deprecated 2026-02-27, retired 2026-05-31, alternative still Mistral Medium 3.5; no field changes." }, { "provider": "Mistral", "provider_url": "https://mistral.ai", "model_id": "mistral-medium-2505", "display_name": "Mistral Medium 3", "model_family": "Mistral Medium", "context_window": 131072, "max_output_tokens": 8192, "input_per_mtok_usd": "0.40", "output_per_mtok_usd": "2.00", "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native"], "deprecated_at": "2026-05-22", "replaced_by_model_id": "mistral-medium-3-5", "last_verified": "2026-09-02", "last_changed_at": "2025-05-07", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://mistral.ai/news/mistral-medium-3", "notes": "La Plateforme rates ($0.40 / $2.00 per MTok) are the row's structured price; unchanged this pass. Mistral's launch post (2025-05-07) lists La Plateforme and Amazon SageMaker at GA with IBM watsonx, NVIDIA NIM, Azure AI Foundry, and Google Cloud Vertex as forthcoming; SageMaker is not in the deployment_options enum and Bedrock has not been confirmed, so deployment_options is restricted to native. Optimized for agentic and coding use cases. Deprecated on La Plateforme 2026-05-22 per docs.mistral.ai's legacy table; retirement 2026-08-31 has now passed, so the model_id should be treated as fully retired/non-callable, not merely deprecated. Alternative is Mistral Medium 3.5 (mistral-medium-3-5). `mistral-medium-latest` alias moved to the new Medium 3.5 row and was removed here to avoid a duplicate. max_output_tokens not published on the model card; defaulted to 8192. Batch API discount per-model availability not confirmed on this date, so batch fields are unset. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02 via mistral.ai/pricing/api and docs.mistral.ai/getting-started/models/models_overview: still deprecated 2026-05-22, retirement 2026-08-31 has now passed (model no longer listed on mistral.ai/pricing/api), alternative still Mistral Medium 3.5; no price field changes." }, { "provider": "Mistral", "provider_url": "https://mistral.ai", "model_id": "mistral-small-2501", "display_name": "Mistral Small 3", "model_family": "Mistral Small", "aliases": ["mistralai/Mistral-Small-24B-Instruct-2501"], "context_window": 32768, "max_output_tokens": 8192, "input_per_mtok_usd": "0.10", "output_per_mtok_usd": "0.30", "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native"], "deprecated_at": "2025-11-06", "replaced_by_model_id": "mistral-small-2603", "last_verified": "2026-09-02", "last_changed_at": "2025-01-30", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://mistral.ai/news/mistral-small-3", "notes": "La Plateforme rates ($0.10 / $0.30 per MTok) are the row's structured price; per Mistral's launch post, half the price of the previous mistral-small ($0.20 / $0.60). 24B-parameter latency-optimized model under Apache 2.0; text-only. Context window 32K per Mistral's spec (33000 tokens rounded; 32768 used here). Deprecated on La Plateforme 2025-11-06 and retired 2025-11-30 per Mistral's legacy table (both dates now well in the past; model_id should be treated as fully retired). Chain of intermediate successors (mistral-small-2503 / 3.1, mistral-small-2506 / 3.2, both also since deprecated) led to Mistral Small 4 (mistral-small-2603, added to this dataset this pass); replaced_by_model_id now resolves in-file. max_output_tokens not published on the model card; defaulted to 8192. Batch API discount per-model availability not confirmed on this date, so batch fields are unset. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02 via mistral.ai/pricing/api and docs.mistral.ai/getting-started/models/models_overview: still deprecated 2025-11-06, retired 2025-11-30, alternative still Mistral Small 4; no field changes." }, { "provider": "Mistral", "provider_url": "https://mistral.ai", "model_id": "codestral-2508", "display_name": "Codestral 25.08", "model_family": "Codestral", "aliases": ["codestral-latest", "codestral-2"], "context_window": 131072, "max_output_tokens": 8192, "input_per_mtok_usd": "0.30", "output_per_mtok_usd": "0.90", "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native", "vertex"], "last_verified": "2026-09-02", "last_changed_at": "2025-07-31", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://mistral.ai/news/codestral-25-08", "notes": "La Plateforme rates ($0.30 / $0.90 per MTok) are the row's structured price; unchanged this pass, cross-checked against mistral.ai/pricing/api. Code-specialized model optimized for fill-in-the-middle (FIM), code completion, code correction, and test generation; supports tool use and structured output per the 25.08 release. Not on Mistral's legacy/deprecation table on this date; still current. Context window corrected to 128K (131072 tokens) per the live docs.mistral.ai/models/model-cards/codestral-25-08 spec card, which lists 128k context; the original launch blog's 256K figure (262144, previously recorded here) does not match the current model card and is superseded by it. Also available on Google Cloud Vertex AI Model Garden as `codestral-2` under the `mistralai` publisher (Mistral Docs: Vertex AI cloud deployments page). max_output_tokens not published on the model card; defaulted to 8192. Batch API discount per-model availability not confirmed on this date, so batch fields are unset. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02: still $0.30/$0.90 on mistral.ai/pricing/api, still absent from docs.mistral.ai's deprecation table (still current); no field changes." }, { "provider": "Mistral", "provider_url": "https://mistral.ai", "model_id": "pixtral-large-2411", "display_name": "Pixtral Large", "model_family": "Pixtral", "aliases": ["pixtral-large-latest", "mistral.pixtral-large-2502-v1:0"], "context_window": 131072, "max_output_tokens": 8192, "input_per_mtok_usd": "2.00", "output_per_mtok_usd": "6.00", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native", "bedrock"], "deprecated_at": "2026-02-27", "replaced_by_model_id": "mistral-medium-3-5", "last_verified": "2026-09-02", "last_changed_at": "2024-11-18", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://mistral.ai/news/pixtral-large", "notes": "La Plateforme rates ($2.00 / $6.00 per MTok) are the row's structured price; pricing parity with Mistral Large 2 since Pixtral Large is the multimodal 124B-parameter open-weight model built on top of Mistral Large 2. Vision-capable: handles documents, charts, and natural images alongside text. Context window 128K (131072 tokens). Bedrock publishes the 25.02 refresh (`mistral.pixtral-large-2502-v1:0`, also routed via `us.mistral.pixtral-large-2502-v1:0`), not the 24.11 build. Deprecated on La Plateforme 2026-02-27; retirement 2026-05-31 per Mistral's legacy table has now passed (re-verified 2026-07-13) so the model_id should be treated as fully retired/non-callable. Mistral's stated alternative is Mistral Medium 3.5 (mistral-medium-3-5, added to this dataset this pass); Pixtral as a standalone product line has been discontinued, its vision capability absorbed into Large 3 / Medium 3.5. `pixtral-large-latest` alias is likely non-functional post-retirement but is left on this row since nothing else claims it. max_output_tokens not published on the model card; defaulted to 8192. Vertex/Azure availability not confirmed for Pixtral Large on this date. Batch API discount per-model availability not confirmed on this date, so batch fields are unset. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02 via mistral.ai/pricing/api and docs.mistral.ai/getting-started/models/models_overview: still deprecated 2026-02-27, retired 2026-05-31, alternative still Mistral Medium 3.5; no field changes." }, { "provider": "Mistral", "provider_url": "https://mistral.ai", "model_id": "mistral-large-2512", "display_name": "Mistral Large 3", "featured": true, "model_family": "Mistral Large", "aliases": ["mistral-large-latest"], "context_window": 262144, "max_output_tokens": 8192, "input_per_mtok_usd": "0.50", "output_per_mtok_usd": "1.50", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2025-12-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://mistral.ai/news/mistral-3", "notes": "La Plateforme rates ($0.50 / $1.50 per MTok) confirmed on mistral.ai/pricing/api under alias `mistral-large-latest`; canonical dated model_id `mistral-large-2512` and 256K (262144) context confirmed on docs.mistral.ai/models/model-cards/mistral-large-3-25-12. Open-weight, general-purpose multimodal model (text + image input) with a Mixture-of-Experts architecture (41B active / 675B total parameters), released 2025-12-02 alongside the Ministral 3 family via the same announcement post. Successor to mistral-large-2411 per Mistral's own alternative-model recommendation is actually Mistral Medium 3.5, not this row, per the legacy table; Large 3 is nonetheless the direct version-number successor and is tracked here as a new, independently-priced row. Bedrock/Vertex/Azure availability not confirmed for this build on this date, so deployment_options is restricted to native. max_output_tokens not published on the model card; defaulted to 8192. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02: still $0.50/$1.50 on mistral.ai/pricing/api under `mistral-large-latest`, still absent from docs.mistral.ai's deprecation table (still current); no field changes." }, { "provider": "Mistral", "provider_url": "https://mistral.ai", "model_id": "mistral-medium-3-5", "display_name": "Mistral Medium 3.5", "model_family": "Mistral Medium", "aliases": ["mistral-medium-latest"], "context_window": 262144, "max_output_tokens": 8192, "input_per_mtok_usd": "1.50", "output_per_mtok_usd": "7.50", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-04-28", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://mistral.ai/news/vibe-remote-agents-mistral-medium-3-5", "notes": "La Plateforme rates ($1.50 / $7.50 per MTok) confirmed on mistral.ai/pricing/api under alias `mistral-medium-latest`; canonical model_id `mistral-medium-3-5` and 256K (262144) context confirmed on docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04, released 2026-04-28. Frontier-class multimodal model (text + image input) optimized for agentic and coding use cases; released as open weights under a Modified MIT license. This is the model Mistral's own deprecation table names as the current alternative for mistral-large-2411, pixtral-large-2411, and mistral-medium-2505 (all now deprecated in this dataset), so replaced_by_model_id on those rows points here. Bedrock/Vertex/Azure availability not confirmed on this date, so deployment_options is restricted to native. max_output_tokens not published on the model card; defaulted to 8192. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02: still $1.50/$7.50 on mistral.ai/pricing/api under `mistral-medium-latest`, still absent from docs.mistral.ai's deprecation table (still current); no field changes." }, { "provider": "Mistral", "provider_url": "https://mistral.ai", "model_id": "mistral-small-2603", "display_name": "Mistral Small 4", "model_family": "Mistral Small", "aliases": ["mistral-small-latest"], "context_window": 262144, "max_output_tokens": 8192, "input_per_mtok_usd": "0.15", "output_per_mtok_usd": "0.60", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-03-16", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03", "notes": "La Plateforme rates ($0.15 / $0.60 per MTok) confirmed on mistral.ai/pricing/api under alias `mistral-small-latest`; canonical model_id `mistral-small-2603` and 256K (262144) context confirmed on the same docs.mistral.ai model card, released 2026-03-16. No dedicated mistral.ai/news announcement post was found for this release on this date, so the docs model card is used as source_url; the pricing page (mistral.ai/pricing/api) is the second confirming source, satisfying the two-source rule. Hybrid model unifying instruct, reasoning, and coding capabilities (119B parameters, 6.5B active); multimodal (text + image input). Named as the alternative to mistral-small-2501 (chain via 2503/2506) and to mistral-small-2506 in Mistral's legacy table. Bedrock/Vertex/Azure availability not confirmed on this date, so deployment_options is restricted to native. max_output_tokens not published on the model card; defaulted to 8192. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02: still $0.15/$0.60 on mistral.ai/pricing/api under `mistral-small-latest`, still absent from docs.mistral.ai's deprecation table (still current; mistral-small-2506 is the deprecated row alternative-pointing here, not this row); no field changes." }, { "provider": "Mistral", "provider_url": "https://mistral.ai", "model_id": "ministral-3b-2512", "display_name": "Ministral 3 3B", "model_family": "Ministral", "aliases": ["ministral-3b-latest"], "context_window": 262144, "max_output_tokens": 8192, "input_per_mtok_usd": "0.10", "output_per_mtok_usd": "0.10", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2025-12-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://mistral.ai/news/mistral-3", "notes": "First Ministral-family row in this dataset. La Plateforme rates ($0.10 / $0.10 per MTok) confirmed on mistral.ai/pricing/api under alias `ministral-3b-latest`; canonical model_id `ministral-3b-2512` and 256K (262144) context confirmed on docs.mistral.ai/models/model-cards/ministral-3-3b-25-12, released 2025-12-02. Smallest/most efficient model in the Ministral 3 family; edge-deployment focused; multimodal (text + image input per the model card's 'robust language and vision capabilities'). Bedrock/Vertex/Azure availability not confirmed on this date, so deployment_options is restricted to native. max_output_tokens not published on the model card; defaulted to 8192. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02: still $0.10/$0.10 on mistral.ai/pricing/api under `ministral-3b-latest`, still absent from docs.mistral.ai's deprecation table (still current); no field changes." }, { "provider": "Mistral", "provider_url": "https://mistral.ai", "model_id": "ministral-8b-2512", "display_name": "Ministral 3 8B", "model_family": "Ministral", "aliases": ["ministral-8b-latest"], "context_window": 262144, "max_output_tokens": 8192, "input_per_mtok_usd": "0.15", "output_per_mtok_usd": "0.15", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2025-12-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://mistral.ai/news/mistral-3", "notes": "La Plateforme rates ($0.15 / $0.15 per MTok) confirmed on mistral.ai/pricing/api under alias `ministral-8b-latest`; canonical model_id `ministral-8b-2512` and 256K (262144) context confirmed on docs.mistral.ai/models/model-cards/ministral-3-8b-25-12, released 2025-12-02. Best-in-class text and vision capabilities for edge deployment; multimodal (text + image input). Bedrock/Vertex/Azure availability not confirmed on this date, so deployment_options is restricted to native. max_output_tokens not published on the model card; defaulted to 8192. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02: still $0.15/$0.15 on mistral.ai/pricing/api under `ministral-8b-latest`, still absent from docs.mistral.ai's deprecation table (still current); no field changes." }, { "provider": "Mistral", "provider_url": "https://mistral.ai", "model_id": "ministral-14b-2512", "display_name": "Ministral 3 14B", "model_family": "Ministral", "aliases": ["ministral-14b-latest"], "context_window": 262144, "max_output_tokens": 8192, "input_per_mtok_usd": "0.20", "output_per_mtok_usd": "0.20", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2025-12-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://mistral.ai/news/mistral-3", "notes": "La Plateforme rates ($0.20 / $0.20 per MTok) confirmed on mistral.ai/pricing/api under alias `ministral-14b-latest`; canonical model_id `ministral-14b-2512` and 256K (262144) context confirmed on docs.mistral.ai/models/model-cards/ministral-3-14b-25-12, released 2025-12-02. Largest model in the Ministral 3 family, performance comparable to the larger Mistral Small 3.2 24B; multimodal (text + image input), optimized for local deployment. Bedrock/Vertex/Azure availability not confirmed on this date, so deployment_options is restricted to native. max_output_tokens not published on the model card; defaulted to 8192. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02: still $0.20/$0.20 on mistral.ai/pricing/api under `ministral-14b-latest`, still absent from docs.mistral.ai's deprecation table (still current); no field changes." }, { "provider": "Cohere", "provider_url": "https://cohere.com", "model_id": "command-a-03-2025", "display_name": "Command A", "model_family": "Command A", "aliases": ["cohere.command-a-03-2025"], "context_window": 256000, "max_output_tokens": 8000, "input_per_mtok_usd": "2.50", "output_per_mtok_usd": "10.00", "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native", "azure"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2025-03-01", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://docs.cohere.com/docs/command-a", "notes": "Cohere's flagship 111B-parameter model: 256K context, text-only, optimized for tool use, RAG, agents, and 23-language multilingual workloads. Price ($2.50 / $10.00 per MTok) per artificialanalysis.ai citing Cohere's API (re-confirmed on its Command A model page 2026-07-19); Command A is still not listed on cohere.com/pricing as of 2026-09-02 (page only publishes a \"Legacy Model Pricing\" FAQ table for Command/Command-light/Command R/Command R+ 04-2024/Command R+ 08-2024, plus Model Vault dedicated-instance rates; newer generative models including Command A still route to \"Get in touch for custom enterprise pricing\"), so confidence remains medium. docs.cohere.com/docs/models confirms `command-a-03-2025` is still Live (256K context, 8K max output, text-only) and Cohere's deprecations page (docs.cohere.com/docs/deprecations, re-checked 2026-09-02) does not list it, so status and price are unchanged since the 2026-08-11 pass. Cohere's model catalog still lists newer generative models (command-a-reasoning-08-2025, command-a-vision-07-2025, command-a-translate-08-2025, command-r-08-2024, and command-a-plus-05-2026 — Live on docs.cohere.com/docs/models with 128K context / 64K max output, text+image input) but none of the generative-model additions publish per-token input/output pricing on cohere.com/pricing or a corroborating secondary source, so they remain deferred as of 2026-09-02. Cohere docs still don't surface AWS Bedrock availability for this model (so `bedrock` remains omitted from deployment_options); Azure AI Foundry availability is published but uses per-deployment IDs, so no Azure alias is encoded. Oracle OCI exposes it as `cohere.command-a-03-2025` (kept as alias). Cache and batch pricing not published by Cohere. Knowledge cutoff not published on the Cohere model card." }, { "provider": "Cohere", "provider_url": "https://cohere.com", "model_id": "command-r-plus-08-2024", "display_name": "Command R+", "model_family": "Command R", "knowledge_cutoff": "2024-03-31", "aliases": ["command-r-plus", "cohere.command-r-plus-v1:0"], "context_window": 128000, "max_output_tokens": 4000, "input_per_mtok_usd": "2.50", "output_per_mtok_usd": "10.00", "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native", "bedrock", "azure"], "last_verified": "2026-09-02", "last_changed_at": "2024-08-30", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://cohere.com/pricing", "notes": "Cohere Platform rates ($2.50 / $10.00 per MTok) are the row's structured price, listed on cohere.com/pricing as \"Command R+ 08-2024\" (page now files it under a \"Legacy Model Pricing\" FAQ table alongside Command/Command-light/Command R/Command R+ 04-2024, meaning it's held for existing customers rather than a headline SKU, but the price is unchanged and it remains orderable, re-confirmed 2026-09-02). 128K context, text-only, optimized for complex RAG and multi-step tool use. Cohere's deprecations page (docs.cohere.com/docs/deprecations, re-checked 2026-09-02) still sunsets only the predecessor `command-r-plus-04-2024` on 2025-09-15 and names this 08-2024 build as the recommended replacement, so it is active on Cohere Platform; docs.cohere.com/docs/models also lists `command-r-plus-08-2024` as Live. Bedrock SKU `cohere.command-r-plus-v1:0` launched Aug 2024 with a Mar 2024 knowledge cutoff (per Bedrock model card), matching this row; as of the prior pass Bedrock had marked the model \"Legacy\" with an EOL of 2026-08-19, which has now passed — this refresh did not re-check the Bedrock console directly, so confirm the Bedrock-SKU status next pass (this is a Bedrock-side lifecycle marker, not a Cohere platform deprecation, so the row itself stays active on Cohere Platform regardless). Azure AI Foundry availability published by Cohere; per-deployment IDs there, so no Azure alias is encoded. Cache and batch pricing not published by Cohere. Cohere's docs/models catalog also now lists a plain `command-r-08-2024` (non-plus) SKU, but it has no published per-token pricing on cohere.com/pricing or a corroborating secondary source, so it is deferred rather than added this pass." }, { "provider": "Cohere", "provider_url": "https://cohere.com", "model_id": "c4ai-aya-expanse-32b", "display_name": "Aya Expanse 32B", "model_family": "Aya Expanse", "context_window": 128000, "max_output_tokens": 4000, "input_per_mtok_usd": "0.50", "output_per_mtok_usd": "1.50", "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": false, "structured_output": false, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2024-10-24", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://cohere.com/pricing", "notes": "32B-parameter multilingual research/generative model (23 languages), 128K context, text-only. Cohere's \"Legacy Model Pricing\" FAQ table on cohere.com/pricing still lists the same single bundled rate for \"Aya Expanse (8B and 32B): $0.50 / $1.50 per MTok\" (not broken out per size) as of 2026-09-02, so confidence remains medium. docs.cohere.com/docs/models lists `c4ai-aya-expanse-32b` as Live, 128K context, 4K max output, text-only, unchanged, and it still does not appear on Cohere's deprecations page (re-checked 2026-09-02). The 8B sibling `c4ai-aya-expanse-8b` reached its scheduled retirement/shutdown date of 2026-04-04 per docs.cohere.com/docs/deprecations (replacement: `command-r7b-12-2024`, `command-a-03-2025`, or `command-a-reasoning-08-2025`) and remains intentionally excluded as a row (never tracked, so no deprecated_at entry needed here). Tool use, structured output, and knowledge cutoff not published by Cohere for this model; booleans defaulted to false pending confirmation. Cache and batch pricing not published." }, { "provider": "xAI", "provider_url": "https://x.ai", "model_id": "grok-4-0709", "display_name": "Grok 4", "model_family": "Grok 4", "knowledge_cutoff": "2024-11-30", "aliases": ["grok-4"], "context_window": 256000, "max_output_tokens": 256000, "input_per_mtok_usd": "3.00", "output_per_mtok_usd": "15.00", "cache_read_per_mtok_usd": "0.75", "pricing_tiers": [ { "threshold_tokens": 128000, "input_per_mtok_usd": "6.00", "output_per_mtok_usd": "30.00" } ], "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native"], "deprecated_at": "2026-05-15", "replaced_by_model_id": "grok-4.3", "last_verified": "2026-09-02", "last_changed_at": "2026-05-15", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://docs.x.ai/developers/migration/may-15-retirement", "notes": "xAI native rates ($3.00 / $15.00 per MTok, $0.75 cached input) are the row's structured price; prompts above 128K total tokens are billed at the higher pricing_tiers rate ($6.00 / $30.00) per xAI's documented long-context tiering. Grok 4 (snapshot `grok-4-0709`, released 2025-07-09) was xAI's flagship reasoning model: reasoning is always on (thinking tokens billed at the output rate, hence reasoning_tokens_billed: true), parallel tool calling and structured outputs supported, accepts text and image inputs. max_output_tokens of 256000 reflects xAI's documented \"up to 256K tokens of output\" within the shared 256K prompt+response context. Retired from the xAI API on 2026-05-15 12:00 PM PT alongside seven other legacy slugs; requests to `grok-4-0709` and `grok-4` continue to resolve but are now redirected to `grok-4.3` with `low` reasoning effort and billed at grok-4.3 rates. Successor is `grok-4.3`, captured in this dataset and referenced via `replaced_by_model_id`. xAI's API is native-only (not on Bedrock/Vertex/Azure). Batch API not published for this model. Re-verified 2026-07-02, 2026-07-13, 2026-07-19, and 2026-08-11: `grok-4-0709` absent from both docs.x.ai/docs/models and docs.x.ai/docs/pricing, and still listed on docs.x.ai/developers/migration/may-15-retirement with redirect target `grok-4.3`, confirming retirement status is unchanged." }, { "provider": "xAI", "provider_url": "https://x.ai", "model_id": "grok-3", "display_name": "Grok 3", "model_family": "Grok 3", "knowledge_cutoff": "2024-11-30", "context_window": 131072, "max_output_tokens": 131072, "input_per_mtok_usd": "3.00", "output_per_mtok_usd": "15.00", "cache_read_per_mtok_usd": "0.75", "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["native"], "deprecated_at": "2026-05-15", "replaced_by_model_id": "grok-4.3", "last_verified": "2026-09-02", "last_changed_at": "2026-05-15", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://docs.x.ai/developers/migration/may-15-retirement", "notes": "xAI native rates ($3.00 / $15.00 per MTok, $0.75 cached input) are the row's structured price (xAI's pricing page and mem0/pricepertoken aggregator both report $3/$15; artificialanalysis.ai reports a higher $4/$20 — choosing xAI-aligned figures). Grok 3 (released 2025-02-19) was xAI's flagship non-reasoning chat model; text-only inputs, function calling and structured outputs supported, 131,072-token combined prompt+response context window. Not a reasoning model (direct responses, no extended chain-of-thought; the reasoning sibling was `grok-3-mini`, not in this dataset). max_output_tokens defaulted to the documented context cap; xAI does not publish a separate max-output limit beyond the shared 131,072-token window. Retired from the xAI API on 2026-05-15 12:00 PM PT; requests to `grok-3` continue to resolve but are now redirected to `grok-4.3` with `none` reasoning effort and billed at grok-4.3 rates. Successor is `grok-4.3`, captured in this dataset and referenced via `replaced_by_model_id`. xAI's API is native-only. Batch API not published for this model. Re-verified 2026-07-02, 2026-07-13, 2026-07-19, and 2026-08-11: `grok-3` absent from both docs.x.ai/docs/models and docs.x.ai/docs/pricing, and still listed on docs.x.ai/developers/migration/may-15-retirement with redirect target `grok-4.3`, confirming retirement status is unchanged." }, { "provider": "xAI", "provider_url": "https://x.ai", "model_id": "grok-code-fast-1", "display_name": "Grok Code Fast 1", "model_family": "Grok Code", "context_window": 256000, "max_output_tokens": 256000, "input_per_mtok_usd": "0.20", "output_per_mtok_usd": "1.50", "cache_read_per_mtok_usd": "0.02", "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["native"], "deprecated_at": "2026-05-15", "replaced_by_model_id": "grok-build-0.1", "last_verified": "2026-09-02", "last_changed_at": "2026-05-15", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://docs.x.ai/developers/migration/may-15-retirement", "notes": "xAI native rates ($0.20 / $1.50 per MTok, $0.02 cached input) are the row's structured price. Grok Code Fast 1 (released 2025-08-26) was xAI's speedy, economical coding-specialized reasoning model: 314B-parameter MoE architecture, 256K combined prompt+response context, agentic coding focus, visible reasoning traces (`reasoning_content` field in streaming responses), function calling and structured outputs supported, text-only. Reasoning is enabled by default so reasoning tokens are billed at the output rate. max_output_tokens defaulted to the documented 256K context cap; xAI does not publish a separate max-output limit beyond the shared window. Retired from the xAI API on 2026-05-15 12:00 PM PT; requests to `grok-code-fast-1` continue to resolve. CORRECTION on 2026-07-13 re-verification: docs.x.ai/developers/migration/may-15-retirement now states the redirect target is `grok-build-0.1` (\"After May 15, requests to `grok-code-fast-1` are routed to `grok-build-0.1`\"), not `grok-4.3` as previously recorded on 2026-07-02 — `replaced_by_model_id` updated accordingly; this is a referential correction, not a price change to this row, so `last_changed_at` is not bumped. `grok-build-0.1` is now captured in this dataset. xAI's API is native-only. Batch API not published for this model. Knowledge cutoff not published by xAI. `grok-code-fast-1` remains absent from both docs.x.ai/docs/models and docs.x.ai/docs/pricing, confirming retirement status is unchanged. Re-verified 2026-07-19 and 2026-08-11: still absent from both vendor pages, and docs.x.ai/developers/migration/may-15-retirement still lists redirect target `grok-build-0.1`." }, { "provider": "xAI", "provider_url": "https://x.ai", "model_id": "grok-4.3", "display_name": "Grok 4.3", "model_family": "Grok 4", "context_window": 1000000, "max_output_tokens": 1000000, "input_per_mtok_usd": "1.25", "output_per_mtok_usd": "2.50", "cache_read_per_mtok_usd": "0.20", "pricing_tiers": [ { "threshold_tokens": 200000, "input_per_mtok_usd": "2.50", "output_per_mtok_usd": "5.00", "cache_read_per_mtok_usd": "0.40" } ], "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-07-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://docs.x.ai/docs/models", "notes": "xAI native rates ($1.25 / $2.50 per MTok, $0.20 cached input) are the row's structured price, listed on docs.x.ai/docs/models and docs.x.ai/docs/pricing; base $1.25/$2.50 unchanged since 2026-05-15. PRICE STRUCTURE UPDATE 2026-07-19: xAI now publishes a cached-input rate ($0.20/MTok) and long-context tiering for this model — prompts ≥200K tokens bill at $2.50 input / $5.00 output / $0.40 cached input — captured in cache_read_per_mtok_usd and pricing_tiers; cross-checked against OpenRouter's public models API (openrouter.ai/api/v1/models, x-ai/grok-4.3: prompt $0.00000125/token = $1.25/MTok, completion $0.0000025/token = $2.50/MTok, input_cache_read $0.0000002/token = $0.20/MTok) — satisfies the two-source rule; last_changed_at bumped. 1M-token combined prompt+response context window (a 4x expansion over the 256K window on Grok 4 / Grok Code Fast 1). Successor to `grok-4-0709`, `grok-3`, and (via corrected redirect) predecessor line of `grok-build-0.1`, retired 2026-05-15 12:00 PM PT. Thinking mode is exposed via the `reasoning_effort` parameter (`none` / `low` / `medium` / `high`); when reasoning is on, thinking tokens are billed at the output rate, hence reasoning_tokens_billed: true. Accepts text and image inputs, text output; parallel tool calling and structured outputs supported. supports_pdf flipped to true on 2026-07-19: OpenRouter's architecture for x-ai/grok-4.3 now lists a `file` input modality — the same derivation already used for the grok-4.5 and grok-build-0.1 rows; capability fix, not a price change. max_output_tokens reflects the documented 1M-token shared window — xAI does not publish a separate max-output limit. Knowledge cutoff and release date still not published by xAI; omitted rather than guessed. xAI's API is native-only (not on Bedrock / Vertex / Azure). Batch API not published for this model. Remains listed on both vendor pages with no deprecation banner as of 2026-07-19; xAI's top-line recommendation remains `grok-4.5`. The `grok-4.20-0309-*` beta variants deferred on 2026-07-02/2026-07-13 were added as rows on 2026-07-19 after sources converged — see those rows. Re-verified 2026-08-11: price, tiering, and context window unchanged on both docs.x.ai/docs/models and docs.x.ai/docs/pricing." }, { "provider": "xAI", "provider_url": "https://x.ai", "model_id": "grok-4.5", "display_name": "Grok 4.5", "featured": true, "model_family": "Grok 4", "knowledge_cutoff": "2026-02-01", "context_window": 500000, "max_output_tokens": 500000, "input_per_mtok_usd": "2.00", "output_per_mtok_usd": "6.00", "cache_read_per_mtok_usd": "0.30", "pricing_tiers": [ { "threshold_tokens": 200000, "input_per_mtok_usd": "4.00", "output_per_mtok_usd": "12.00", "cache_read_per_mtok_usd": "0.60" } ], "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-07-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://docs.x.ai/docs/models", "notes": "Grok 4.5 launched 2026-07-08 as xAI's top-line recommended model (\"the most intelligent and fastest model we've built\"; docs.x.ai/docs/models: \"For everything else, including code, use Grok 4.5\"), positioned for coding, agentic, and knowledge-work use. xAI native rates ($2.00 / $6.00 per MTok, $0.30 cached input) are the row's structured price, confirmed on both docs.x.ai/docs/models and docs.x.ai/docs/pricing, cross-checked against OpenRouter's public models API (openrouter.ai/api/v1/models, x-ai/grok-4.5: prompt $0.000002/token = $2.00/MTok, completion $0.000006/token = $6.00/MTok, input_cache_read $0.0000003/token = $0.30/MTok cache read) — satisfies the two-source rule. PRICE CHANGE 2026-07-19: cached-input rate dropped from $0.50 to $0.30/MTok (both vendor pages and OpenRouter agree; $0.50 was the verified rate on 2026-07-13), and xAI now publishes long-context tiering — prompts ≥200K tokens bill at $4.00 input / $12.00 output / $0.60 cached input — captured in pricing_tiers (OpenRouter does not model per-tier rates; tier figures are per docs.x.ai/docs/models and docs.x.ai/docs/pricing, which agree); last_changed_at bumped. Context window 500,000 tokens (smaller than grok-4.3's 1M) per docs.x.ai and OpenRouter agreement; knowledge cutoff \"February 1, 2026\" per docs.x.ai/docs/models verbatim. max_output_tokens defaulted to the 500K context cap; xAI does not publish a separate max-output limit (OpenRouter reports max_completion_tokens: null, consistent with no separate cap). Input modalities text + image (OpenRouter architecture also lists a `file` input modality, taken here as supports_pdf: true; images up to 20MiB, jpg/jpeg/png); output text-only. OpenRouter's supported_parameters list for this model includes `tools`/`tool_choice` (supports_tool_use: true), `response_format`/`structured_outputs` (structured_output: true), and `reasoning`/`reasoning_effort`/`include_reasoning` (reasoning_tokens_billed: true, consistent with the `reasoning_effort` parameter pattern on grok-4.3). Audio input/output not supported. Batch pricing and deployment options beyond native not published by xAI as of last_verified; xAI's API remains native-only (not on Bedrock/Vertex/Azure). Does not affect the existing `grok-4.3` row — see that row's notes. Re-verified 2026-08-11: price, tiering, and context window unchanged on both docs.x.ai/docs/models and docs.x.ai/docs/pricing." }, { "provider": "xAI", "provider_url": "https://x.ai", "model_id": "grok-4.6", "display_name": "Grok 4.6", "model_family": "Grok 4", "knowledge_cutoff": "2026-02-01", "context_window": 500000, "max_output_tokens": 500000, "input_per_mtok_usd": "2.00", "output_per_mtok_usd": "6.00", "cache_read_per_mtok_usd": "0.50", "pricing_tiers": [ { "threshold_tokens": 200000, "input_per_mtok_usd": "4.00", "output_per_mtok_usd": "12.00", "cache_read_per_mtok_usd": "1.00" } ], "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-09-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://docs.x.ai/docs/models", "notes": "New row, added 2026-09-02. Grok 4.6 launched as xAI's new top-line recommended model (docs.x.ai/docs/models: \"For everything else, including code, use Grok 4.6. It is the most intelligent and fastest model we've built.\" — the same positioning language previously used for grok-4.5), for coding, agentic, and knowledge-work use. xAI native rates ($2.00 / $6.00 per MTok, $0.50 cached input) confirmed on both docs.x.ai/docs/models and docs.x.ai/docs/pricing, cross-checked against OpenRouter's public models API (openrouter.ai/api/v1/models, x-ai/grok-4.6: prompt $0.000002/token = $2.00/MTok, completion $0.000006/token = $6.00/MTok, input_cache_read $0.0000005/token = $0.50/MTok) — satisfies the two-source rule. Long-context tiering: prompts ≥200K tokens bill at $4.00 input / $12.00 output / $1.00 cached input, per docs.x.ai/docs/pricing (OpenRouter does not model per-tier rates; tier figures are vendor-only). Context window 500,000 tokens, unchanged from grok-4.5; knowledge cutoff \"February 1, 2026\" per docs.x.ai/docs/models verbatim. max_output_tokens defaulted to the 500K context cap; xAI does not publish a separate max-output limit (OpenRouter reports no separate completion cap). Input modalities text + image (OpenRouter architecture also lists a `file` input modality, taken here as supports_pdf: true, same derivation used for sibling xAI rows); output text-only. OpenRouter's supported_parameters for x-ai/grok-4.6 include `tools`/`tool_choice` (supports_tool_use: true), `response_format`/`structured_outputs` (structured_output: true), and `reasoning`/`reasoning_effort`/`include_reasoning` (reasoning_tokens_billed: true). Audio input/output not supported. Batch pricing and deployment options beyond native not published; xAI's API remains native-only (not on Bedrock/Vertex/Azure). The predecessor `grok-4.5` row remains listed on both vendor pages with no deprecation banner as of 2026-09-02 and unchanged pricing ($2.00/$6.00/$0.30 cached, tiering $4.00/$12.00/$0.60); not marked deprecated in this dataset since xAI has only repositioned its top-line recommendation, not retired the model — grok-4.5's `featured: true` flag is preserved per the runbook's never-remove-featured rule. `featured: true` was NOT added to this new grok-4.6 row in this pass: that requires a `pnpm costs:featured --write` lockfile regen (`web/lib/featured.lock.json`), which is outside this dispatch's scope (costs/llm.json only, per workflow constraints). Flagging grok-4.6 as a featured candidate for the controller to evaluate in a follow-up pass." }, { "provider": "xAI", "provider_url": "https://x.ai", "model_id": "grok-build-0.1", "display_name": "Grok Build 0.1", "model_family": "Grok Code", "context_window": 256000, "max_output_tokens": 256000, "input_per_mtok_usd": "1.00", "output_per_mtok_usd": "2.00", "cache_read_per_mtok_usd": "0.20", "pricing_tiers": [ { "threshold_tokens": 200000, "input_per_mtok_usd": "2.00", "output_per_mtok_usd": "4.00", "cache_read_per_mtok_usd": "0.40" } ], "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-07-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://docs.x.ai/docs/models", "notes": "Grok Build 0.1 (public beta, first opened via API 2026-05-20) is xAI's fast, economical agentic-coding model that powers the Grok Build CLI; it is the documented redirect target for the retired `grok-code-fast-1` slug (docs.x.ai/developers/migration/may-15-retirement: \"After May 15, requests to `grok-code-fast-1` are routed to `grok-build-0.1`\") — see that row's `replaced_by_model_id`. xAI native rates ($1.00 / $2.00 per MTok) confirmed on docs.x.ai/docs/models and docs.x.ai/docs/pricing, cross-checked against OpenRouter's public models API (openrouter.ai/api/v1/models, x-ai/grok-build-0.1: prompt $0.000001/token = $1.00/MTok, completion $0.000002/token = $2.00/MTok, input_cache_read $0.0000002/token = $0.20/MTok cache read) — satisfies the two-source rule; base and cached rates re-confirmed unchanged on 2026-07-19. PRICE STRUCTURE UPDATE 2026-07-19: xAI now publishes long-context tiering for this model — prompts ≥200K tokens bill at $2.00 input / $4.00 output / $0.40 cached input — captured in pricing_tiers (tier figures per docs.x.ai/docs/models and docs.x.ai/docs/pricing, which agree; OpenRouter does not model per-tier rates); last_changed_at bumped. Context window 256,000 tokens per both sources; max_output_tokens defaulted to the 256K context cap (OpenRouter reports max_completion_tokens: null, no separate limit published). Input modalities text + image (OpenRouter architecture lists a `file` input modality too, taken here as supports_pdf: true); output text-only. OpenRouter's supported_parameters for this model include `tools`/`tool_choice` (supports_tool_use: true), `response_format`/`structured_outputs` (structured_output: true), and `reasoning`/`include_reasoning` (reasoning_tokens_billed: true — consistent with predecessor `grok-code-fast-1`'s visible reasoning traces). Audio input/output not supported. Knowledge cutoff and release date beyond the 2026-05-20 API-opening not published by xAI; omitted rather than guessed. Batch pricing not published. xAI's API is native-only (not on Bedrock/Vertex/Azure); limited regional availability noted by third-party docs (us-east-1, us-west-2) but not encoded as a separate deployment_option since it isn't a distinct hosted-API SKU distinction xAI itself publishes. Re-verified 2026-08-11: price, tiering, and context window unchanged on both docs.x.ai/docs/models and docs.x.ai/docs/pricing." }, { "provider": "xAI", "provider_url": "https://x.ai", "model_id": "grok-4.20-0309-reasoning", "display_name": "Grok 4.20 Reasoning", "model_family": "Grok 4.20", "context_window": 1000000, "max_output_tokens": 1000000, "input_per_mtok_usd": "1.25", "output_per_mtok_usd": "2.50", "cache_read_per_mtok_usd": "0.20", "pricing_tiers": [ { "threshold_tokens": 200000, "input_per_mtok_usd": "2.50", "output_per_mtok_usd": "5.00", "cache_read_per_mtok_usd": "0.40" } ], "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://docs.x.ai/docs/models", "notes": "New row, added on the 2026-07-19 refresh; deferred on 2026-07-02 and 2026-07-13 because a third-party writeup (buildfastwithai.com, 2026-03-14) contradicted xAI's published specs ($2.00/$6.00 per MTok, 2M context). Sources have now converged on price: docs.x.ai/docs/models and docs.x.ai/docs/pricing both list `grok-4.20-0309-reasoning` at $1.25 / $2.50 per MTok with $0.20 cached input and long-context tiering (prompts ≥200K tokens: $2.50 / $5.00, $0.40 cached input), and OpenRouter's public models API corroborates the family pricing (openrouter.ai/api/v1/models, x-ai/grok-4.20: prompt $0.00000125/token = $1.25/MTok, completion $0.0000025/token = $2.50/MTok, input_cache_read $0.0000002/token = $0.20/MTok). confidence: medium because the secondary-source match is at family level, not exact slug — OpenRouter lists the family as `x-ai/grok-4.20` (no `-0309-` snapshot suffix) and reports a 2M-token context, while xAI's own models and pricing pages state 1M for this API model_id; the vendor figure (1,000,000) is used as authoritative for the native API. Grok 4.20 beta reasoning variant (snapshot suffix 0309, consistent with a 2026-03-09 snapshot; xAI does not publish a release date). Reasoning tokens billed at the output rate (reasoning_tokens_billed: true). Capabilities per OpenRouter x-ai/grok-4.20 supported_parameters: `tools`/`tool_choice` (supports_tool_use: true), `response_format`/`structured_outputs` (structured_output: true), `reasoning`/`include_reasoning`. Input modalities text + image (OpenRouter architecture also lists `file`, taken as supports_pdf: true, same derivation as sibling xAI rows); output text-only; audio not supported. max_output_tokens defaulted to the 1M context cap — no separate max-output limit published. Knowledge cutoff not published; omitted rather than guessed. Batch pricing not published. xAI's API is native-only (not on Bedrock/Vertex/Azure). Priced identically to `grok-4.3`. Sibling rows: `grok-4.20-0309-non-reasoning`, `grok-4.20-multi-agent-0309`. Re-verified 2026-08-11: price, tiering, and context window unchanged on docs.x.ai/docs/models and docs.x.ai/docs/pricing; confidence left at medium pending re-check of the OpenRouter family-level slug/context discrepancy." }, { "provider": "xAI", "provider_url": "https://x.ai", "model_id": "grok-4.20-0309-non-reasoning", "display_name": "Grok 4.20 Non-Reasoning", "model_family": "Grok 4.20", "context_window": 1000000, "max_output_tokens": 1000000, "input_per_mtok_usd": "1.25", "output_per_mtok_usd": "2.50", "cache_read_per_mtok_usd": "0.20", "pricing_tiers": [ { "threshold_tokens": 200000, "input_per_mtok_usd": "2.50", "output_per_mtok_usd": "5.00", "cache_read_per_mtok_usd": "0.40" } ], "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": false, "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://docs.x.ai/docs/models", "notes": "New row, added on the 2026-07-19 refresh; deferred on 2026-07-02 and 2026-07-13 because a third-party writeup (buildfastwithai.com, 2026-03-14) contradicted xAI's published specs. Sources have now converged on price: docs.x.ai/docs/models and docs.x.ai/docs/pricing both list `grok-4.20-0309-non-reasoning` at $1.25 / $2.50 per MTok with $0.20 cached input and long-context tiering (prompts ≥200K tokens: $2.50 / $5.00, $0.40 cached input), and OpenRouter's public models API corroborates the family pricing under `x-ai/grok-4.20` (prompt $0.00000125/token, completion $0.0000025/token, input_cache_read $0.0000002/token). confidence: medium for the same reasons as the sibling `grok-4.20-0309-reasoning` row: family-level (not exact-slug) secondary corroboration, and a context-window discrepancy (OpenRouter reports 2M; xAI's models and pricing pages state 1M — vendor figure 1,000,000 used as authoritative for the native API). Grok 4.20 beta non-reasoning variant: direct responses without extended chain-of-thought, so reasoning_tokens_billed: false (analogous to the retired `grok-3` and to grok-4.3's `reasoning_effort: none` mode). Function calling and structured outputs supported per OpenRouter x-ai/grok-4.20 supported_parameters (`tools`/`tool_choice`, `response_format`/`structured_outputs`). Input modalities text + image (OpenRouter architecture also lists `file`, taken as supports_pdf: true, same derivation as sibling xAI rows); output text-only; audio not supported. max_output_tokens defaulted to the 1M context cap — no separate max-output limit published. Knowledge cutoff and release date not published; omitted rather than guessed. Batch pricing not published. xAI's API is native-only (not on Bedrock/Vertex/Azure). Priced identically to `grok-4.3`. Sibling rows: `grok-4.20-0309-reasoning`, `grok-4.20-multi-agent-0309`. Re-verified 2026-08-11: price, tiering, and context window unchanged on docs.x.ai/docs/models and docs.x.ai/docs/pricing; confidence left at medium pending re-check of the OpenRouter family-level slug/context discrepancy." }, { "provider": "xAI", "provider_url": "https://x.ai", "model_id": "grok-4.20-multi-agent-0309", "display_name": "Grok 4.20 Multi-Agent", "model_family": "Grok 4.20", "context_window": 1000000, "max_output_tokens": 1000000, "input_per_mtok_usd": "1.25", "output_per_mtok_usd": "2.50", "cache_read_per_mtok_usd": "0.20", "pricing_tiers": [ { "threshold_tokens": 200000, "input_per_mtok_usd": "2.50", "output_per_mtok_usd": "5.00", "cache_read_per_mtok_usd": "0.40" } ], "modalities": { "input": ["text", "image"], "output": ["text"] }, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://docs.x.ai/docs/models", "notes": "New row, added on the 2026-07-19 refresh; deferred on 2026-07-02 and 2026-07-13 (a third-party writeup then reported the multi-agent API as 'coming soon' with conflicting prices). Sources have now converged: docs.x.ai/docs/models and docs.x.ai/docs/pricing both list `grok-4.20-multi-agent-0309` at $1.25 / $2.50 per MTok with $0.20 cached input and long-context tiering (prompts ≥200K tokens: $2.50 / $5.00, $0.40 cached input), and OpenRouter carries a live listing `x-ai/grok-4.20-multi-agent` with matching pricing (prompt $0.00000125/token, completion $0.0000025/token, input_cache_read $0.0000002/token) — the API is live, no longer 'coming soon'. confidence: medium: OpenRouter's slug lacks the `-0309` snapshot suffix and reports a 2M-token context while xAI's own pages state 1M — vendor figure (1,000,000) used as authoritative for the native API. Grok 4.20 beta multi-agent variant: multiple agents run in parallel to conduct deep research, coordinate tool use internally, and synthesize a final answer (per OpenRouter's model description). supports_tool_use omitted rather than guessed: OpenRouter's supported_parameters for x-ai/grok-4.20-multi-agent do not include `tools`/`tool_choice` (client-side function calling), unlike the sibling variants — tool use appears to be internal to the agent swarm; xAI does not publish a definitive statement. `response_format`/`structured_outputs` supported (structured_output: true); `reasoning`/`reasoning_effort` supported, thinking tokens billed at the output rate (reasoning_tokens_billed: true). Input modalities text + image (OpenRouter architecture also lists `file`, taken as supports_pdf: true); output text-only; audio not supported. max_output_tokens defaulted to the 1M context cap — no separate max-output limit published. Knowledge cutoff and release date not published; omitted rather than guessed. Batch pricing not published. xAI's API is native-only (not on Bedrock/Vertex/Azure). Priced identically to `grok-4.3`. Sibling rows: `grok-4.20-0309-reasoning`, `grok-4.20-0309-non-reasoning`. Re-verified 2026-08-11: price, tiering, and context window unchanged on docs.x.ai/docs/models and docs.x.ai/docs/pricing; confidence left at medium pending re-check of the OpenRouter family-level slug/context discrepancy." }, { "provider": "Alibaba", "provider_url": "https://www.alibabacloud.com/product/modelstudio", "model_id": "qwen3-max", "display_name": "Qwen3-Max", "model_family": "Qwen 3 Max", "knowledge_cutoff": "2025-06-30", "context_window": 262144, "max_output_tokens": 32768, "input_per_mtok_usd": "1.20", "output_per_mtok_usd": "6.00", "pricing_tiers": [ { "threshold_tokens": 32768, "input_per_mtok_usd": "2.40", "output_per_mtok_usd": "12.00" }, { "threshold_tokens": 131072, "input_per_mtok_usd": "3.00", "output_per_mtok_usd": "15.00" } ], "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["dashscope"], "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.alibabacloud.com/help/en/model-studio/model-pricing", "notes": "Alibaba DashScope (International) tiered pricing by input-token bucket: 0-32K = $1.20 input / $6.00 output per MTok (base row rate); 32K-128K = $2.40 / $12.00; 128K-252K = $3.00 / $15.00 (captured in pricing_tiers) — unchanged at last_verified. Context window 262,144 tokens (DashScope publishes 252K as the top-tier pricing ceiling; Qwen team and OpenRouter publish the full 262,144 model context). Max output 32,768 tokens. `qwen3-max` is a rolling alias: as of 2026-07-02 it resolves to `qwen3-max-2026-01-23` (previously `qwen3-max-2025-09-23`, still separately listed); pricing identical across both dated snapshots, so this is a metadata-only pointer change. Knowledge cutoff 2025-06-30 (unverified against the new snapshot; vendor page does not publish per-snapshot cutoffs, carried over from prior verification). Hybrid thinking model: thinking mode disabled by default but available via `/think` (and disabled via `/no_think`); when enabled, thinking tokens are billed at the output rate, hence reasoning_tokens_billed: true. Text-only inputs and outputs (the Qwen3-VL family is a separate set of model_ids). Tool calling and structured outputs supported via the DashScope and OpenAI-compatible endpoints. Explicit context cache discounts cached input tokens to 10% of the standard rate, but DashScope does not publish a single cache_read figure across the tiered input rates, so cache_read_per_mtok_usd is omitted rather than guessed. Deployment via DashScope (Model Studio) only at last_verified; not on Bedrock, Vertex, Together, Fireworks, or Groq as a first-party offering. CHANGE 2026-07-02: batch_input_per_mtok_usd/batch_output_per_mtok_usd removed — the vendor pricing page's International/Singapore \"Qwen-Max\" table for qwen3-max no longer carries a batch-inference-discount badge (verified against raw page HTML: only the \"context caching discount\" annotation remains). The 50% batch discount badge is now only shown on the Chinese-mainland deployment row for this model family, which this dataset does not price (DashScope International is canonical). Prior batch figures ($0.60/$3.00) were the International 50%-off rate and are no longer offered there as of this verification. Re-verified 2026-07-13: prices, tiers, and batch/cache posture unchanged against alibabacloud.com/help/en/model-studio/billing; still no batch badge on the International row (China-mainland row still carries it), model_id still active (not deprecated), no newer dated snapshot beyond qwen3-max-2026-01-23. Re-verified 2026-07-19: International (Singapore) list prices and tiers unchanged ($1.20/$6.00, $2.40/$12.00, $3.00/$15.00); alias still resolves to qwen3-max-2026-01-23; no deprecation notice. New since last pass: a 'Limited-time 50% off' promo label now appears on the International qwen3-max row — promo only, list prices (recorded here) unchanged, so last_changed_at not bumped. Re-verified 2026-08-11: International tiers unchanged ($1.20/$6.00, $2.40/$12.00, $3.00/$15.00 across 0-32K/32K-128K/128K-256K); alias still qwen3-max-2026-01-23, no newer dated snapshot; model_id still listed and active on the billing page (not in the 'recommended models' table alongside qwen3.7-max/qwen3.7-plus/qwen3.6-flash per the models page, but still priced with no deprecation banner); no deprecation notice. Re-verified 2026-09-02: International tiers unchanged ($1.20/$6.00, $2.40/$12.00, $3.00/$15.00 across 0-32K/32K-128K/128K-256K); alias still resolves to qwen3-max-2026-01-23 with no newer dated snapshot; row still priced with no deprecation banner on the billing page. A new `qwen3.8-max` flagship generation launched this pass (added as a separate row below) alongside `qwen3.7-max`; qwen3-max remains a distinct, separately-priced, still-active SKU (not superseded/replaced) so no `deprecated_at`/`replaced_by_model_id` set here." }, { "provider": "Alibaba", "provider_url": "https://www.alibabacloud.com/product/modelstudio", "model_id": "qwen3-coder-plus", "display_name": "Qwen3-Coder-Plus", "model_family": "Qwen 3 Coder", "context_window": 1048576, "max_output_tokens": 65536, "input_per_mtok_usd": "1.00", "output_per_mtok_usd": "5.00", "pricing_tiers": [ { "threshold_tokens": 32768, "input_per_mtok_usd": "1.80", "output_per_mtok_usd": "9.00" }, { "threshold_tokens": 131072, "input_per_mtok_usd": "3.00", "output_per_mtok_usd": "15.00" }, { "threshold_tokens": 262144, "input_per_mtok_usd": "6.00", "output_per_mtok_usd": "60.00" } ], "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": false, "deployment_options": ["dashscope"], "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.alibabacloud.com/help/en/model-studio/model-pricing", "notes": "Alibaba DashScope (International) tiered pricing by input-token bucket: 0-32K = $1.00 / $5.00 per MTok (base row rate); 32K-128K = $1.80 / $9.00; 128K-256K = $3.00 / $15.00; 256K-1M = $6.00 / $60.00 (captured in pricing_tiers) — unchanged at last_verified. 1,000,000-token context window with 65,536 max output tokens. `qwen3-coder-plus` still resolves to `qwen3-coder-plus-2025-09-23` (no newer dated snapshot published; an older `qwen3-coder-plus-2025-07-22` snapshot remains listed separately at identical pricing). Built on the Qwen3-Coder 480B-A35B MoE base; positioned for agentic coding (robust tool calling and environment interaction). Not a thinking/reasoning SKU (no chain-of-thought billing semantics), so reasoning_tokens_billed is false. Text-only modalities. Explicit context cache discounts cached input to 10% of the standard rate; implicit cache to 20%; DashScope does not publish a single cache_read figure across tiered input rates, so cache_read_per_mtok_usd is omitted rather than guessed. Knowledge cutoff not published. Deployment via DashScope (Model Studio) only at last_verified. CHANGE 2026-07-02: batch_input_per_mtok_usd/batch_output_per_mtok_usd removed — the vendor's \"Qwen-Coder\" pricing section intro no longer mentions batch-call pricing at all (unlike the \"Qwen-Max\" section, which still documents a batch discount), and neither the International nor Chinese-mainland qwen3-coder-plus table rows carry a batch badge (verified against raw page HTML across all 4 occurrences of this model_id on the page). Batch inference discount is no longer offered for this model as of this verification; prior figures ($0.50/$2.50) are stale. Re-verified 2026-07-13: prices, tiers, and no-batch posture unchanged against alibabacloud.com/help/en/model-studio/billing; still resolves to qwen3-coder-plus-2025-09-23 with no newer dated snapshot; model_id still active. Re-verified 2026-07-19: International (Singapore) tiered prices unchanged ($1.00/$5.00, $1.80/$9.00, $3.00/$15.00, $6.00/$60.00); alias still qwen3-coder-plus-2025-09-23; context caching discount still present, no batch badge; no deprecation notice. Re-verified 2026-08-11: International tiered prices unchanged ($1.00/$5.00, $1.80/$9.00, $3.00/$15.00, $6.00/$60.00); alias still qwen3-coder-plus-2025-09-23, no newer dated snapshot; model_id still active on the billing page; no deprecation notice. Re-verified 2026-09-02 (raw-HTML scrape of the Singapore table): International tiered prices unchanged ($1.00/$5.00, $1.80/$9.00, $3.00/$15.00, $6.00/$60.00); alias still qwen3-coder-plus-2025-09-23; still no batch-discount badge on the International row (context-caching-discount blockquote only); no deprecation notice. Not in scope of this pass's requested 4-row list but re-verified alongside the other Alibaba rows for consistency since it was already pulled from the same source fetch." }, { "provider": "Alibaba", "provider_url": "https://www.alibabacloud.com/product/modelstudio", "model_id": "qwen3.7-max", "display_name": "Qwen3.7-Max", "featured": true, "model_family": "Qwen 3.7 Max", "context_window": 1048576, "max_output_tokens": 65536, "input_per_mtok_usd": "2.50", "output_per_mtok_usd": "7.50", "modalities": { "input": ["text"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": false, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["dashscope"], "aggregators": ["openrouter"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.alibabacloud.com/help/en/model-studio/model-pricing", "notes": "New flagship generation added this pass. `qwen3.7-max` is a rolling alias, currently equivalent to `qwen3.7-max-2026-05-20` per the vendor pricing page (updated Jun 26, 2026); a newer `qwen3.7-max-2026-06-08` dated snapshot is also listed at identical pricing but is not yet the default alias. Flat (non-tiered) DashScope International rate of $2.50 input / $7.50 output per MTok across the full 0-1M token range (no pricing_tiers needed, unlike qwen3-max/qwen3-coder-plus). Cross-verified against OpenRouter (openrouter.ai/qwen/qwen3.7-max), which lists the same $2.50/$7.50 standard rate (its displayed $1.25/$3.75 is an explicitly-labeled 50%-off launch promotion, not the standard price). No batch-inference-discount badge on the International row (Chinese-mainland deployment shows one; not priced here, DashScope International is canonical). Context caching discount available but DashScope does not publish a flat cache_read rate, so cache_read_per_mtok_usd is omitted. 1M-token context window confirmed on the vendor pricing page; max_output_tokens (65,536) is not published on the vendor pricing/model pages directly and is sourced from third-party trackers (OpenRouter, llm-stats.com) that agree on the figure — confidence: medium reflects this one field, not the price. Text-only modalities (MarkTechPost coverage of the 2026-05-20 launch explicitly notes no image input); native tool/function calling supported (reported over 1,000 sequential tool calls in agentic testing). Hybrid thinking model (Non-Thinking and Thinking modes both billed at the same rate), hence reasoning_tokens_billed: true. Knowledge cutoff not published by the vendor; omitted rather than guessed. Deployment via DashScope (Model Studio) only; not yet confirmed on Bedrock, Vertex, Together, Fireworks, or Groq as a first-party offering. Re-verified 2026-07-13: flat $2.50/$7.50 rate, 1M context, and no-batch-on-International posture unchanged against alibabacloud.com/help/en/model-studio/billing; cross-checked openrouter.ai/api/v1/models (qwen/qwen3.7-max: context_length 1,000,000, top_provider.max_completion_tokens 65,536) — corroborates max_output_tokens beyond the prior third-party-tracker citation. Still no qwen4-generation flagship listed by the vendor as of this pass; qwen3.7-max remains current. Re-verified 2026-07-19: flat $2.50/$7.50 list rate over 0-1M unchanged on the International row; the 'Limited-time 50% off' promo label now appears on the vendor billing page itself (previously observed only on OpenRouter) — promo only, list price recorded here, so last_changed_at not bumped. Alias still resolves to qwen3.7-max-2026-05-20; the qwen3.7-max-2026-06-08 dated snapshot remains listed at identical pricing but is still not the default alias. No deprecation notice; still no qwen4/qwen3.8 flagship listed. Re-verified 2026-08-11: flat $2.50/$7.50 list rate over 0-1M unchanged; alias still resolves to qwen3.7-max-2026-05-20 with qwen3.7-max-2026-06-08 still listed at identical pricing and still not the default; no deprecation notice; still no qwen4/qwen3.8 flagship. Re-verified 2026-09-02 (raw-HTML scrape): flat $2.50/$7.50 list rate over 0-1M unchanged on the International row; alias still resolves to qwen3.7-max-2026-05-20; no deprecation notice. A new `qwen3.8-max` flagship generation launched this pass at $2.00/$6.00 flat (added as a separate row below); qwen3.7-max remains a distinct, separately-priced, still-active SKU on the billing page — not superseded, so `deprecated_at`/`replaced_by_model_id` are not set here." }, { "provider": "Alibaba", "provider_url": "https://www.alibabacloud.com/product/modelstudio", "model_id": "qwen3.7-plus", "display_name": "Qwen3.7-Plus", "model_family": "Qwen 3.7 Plus", "context_window": 1048576, "max_output_tokens": 65536, "input_per_mtok_usd": "0.40", "output_per_mtok_usd": "1.60", "pricing_tiers": [ { "threshold_tokens": 262144, "input_per_mtok_usd": "1.20", "output_per_mtok_usd": "4.80" } ], "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["dashscope"], "aggregators": ["openrouter"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.alibabacloud.com/help/en/model-studio/billing", "notes": "New row, added on the 2026-07-13 refresh. Qwen3.7-Plus is Alibaba's cost-effective multimodal agent model in the Qwen3.7 series (released 2026-05-31/06-02), a step down from qwen3.7-max. `qwen3.7-plus` is a rolling alias, currently equivalent to `qwen3.7-plus-2026-05-26`. DashScope International tiered pricing by input-token bucket: 0-256K = $0.40 input / $1.60 output per MTok (base row rate); 256K-1M = $1.20 input / $4.80 output (captured in pricing_tiers). CHANGE 2026-07-19: 256K-1M tier output rate corrected/moved from $1.60 to $4.80 per MTok — the vendor billing page now unambiguously lists $4.8 output for the 256K-1M International tier (confirmed in two separate fetches of the page), and OpenRouter's models API corroborates via its min_prompt_tokens=256000 pricing override (completion $3.84/MTok = exactly 20% off the $4.80 list, matching the vendor's limited-time 20%-off promo; base-tier override figures likewise match $0.40/$1.60/$1.20 list at 20% off). Structured prices here remain list prices, not promo prices. The 20%-off limited-time promotion now appears on the International row itself (both tiers), not only China-mainland. Context window 1,048,576 tokens (1M) per vendor billing page. max_output_tokens (65,536) not published on the vendor billing page directly; sourced from OpenRouter's public models API (openrouter.ai/api/v1/models, qwen/qwen3.7-plus: top_provider.max_completion_tokens 65536, context_length 1,000,000) and cross-checked against llm-stats.com/api/models/qwen3.7-plus (\"up to 65,536 output tokens\", 1M context, multimodal: true) — satisfies the two-source rule; confidence: medium reflects this field, not the price. Input modalities text + image (\"multimodal interactive hybrid agent\": perceives scenes, reads screens/GUIs, writes code from visual references) per OpenRouter architecture (text+image->text) and llm-stats.com; output text-only. Vendor billing page itself does not call out modality explicitly (no separate image-token pricing tier), so supports_vision is sourced from the two secondary trackers, not the primary. Tool calling and structured outputs supported per OpenRouter's supported_parameters (tools/tool_choice, response_format/structured_outputs). Always-on/hybrid thinking (llm-stats.com tags thinking: true); thinking tokens billed at the standard output rate, hence reasoning_tokens_billed: true. Context caching discount available but DashScope does not publish a flat cache_read rate, so cache_read_per_mtok_usd is omitted rather than guessed. No batch-inference-discount badge found for the International row. Knowledge cutoff not published by the vendor; omitted rather than guessed. Deployment via DashScope (Model Studio) only; not yet confirmed on Bedrock, Vertex, Together, Fireworks, or Groq as a first-party offering. Re-verified 2026-08-11: International tiered prices unchanged ($0.40/$1.60 for 0-256K; $1.20/$4.80 for 256K-1M); alias still resolves to qwen3.7-plus-2026-05-26; no deprecation notice. Re-verified 2026-09-02 (raw-HTML scrape): International tiered list prices unchanged ($0.40/$1.60 for 0-256K; $1.20/$4.80 for 256K-1M, still shown with the limited-time 20%-off promo label — list prices recorded here); alias still resolves to qwen3.7-plus-2026-05-26; no deprecation notice; row still in the vendor's models-page 'recommended' set alongside qwen3.8-max and qwen3.8-flash." }, { "provider": "Alibaba", "provider_url": "https://www.alibabacloud.com/product/modelstudio", "model_id": "qwen3.6-flash", "display_name": "Qwen3.6-Flash", "model_family": "Qwen 3.6 Flash", "context_window": 1048576, "max_output_tokens": 65536, "input_per_mtok_usd": "0.25", "output_per_mtok_usd": "1.50", "pricing_tiers": [ { "threshold_tokens": 262144, "input_per_mtok_usd": "1.00", "output_per_mtok_usd": "4.00" } ], "modalities": { "input": ["text", "image", "video"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["dashscope"], "aggregators": ["openrouter"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-13", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.alibabacloud.com/help/en/model-studio/billing", "notes": "New row, added on the 2026-07-13 refresh. Qwen3.6-Flash is Alibaba's fast/efficient tier in the Qwen3.6 series (released 2026-04-16), positioned below the Qwen3.7 generation on price but supporting multimodal input. `qwen3.6-flash` is a rolling alias, currently equivalent to `qwen3.6-flash-2026-04-16`. DashScope International tiered pricing by input-token bucket: 0-256K = $0.25 input / $1.50 output per MTok (base row rate); 256K-1M = $1.00 / $4.00 (captured in pricing_tiers). Context window 1,048,576 tokens (1M) per vendor billing page. max_output_tokens (65,536) not published on the vendor billing page directly; sourced from OpenRouter's public models API (openrouter.ai/api/v1/models, qwen/qwen3.6-flash: top_provider.max_completion_tokens 65536, context_length 1,000,000, architecture text+image+video->text) — satisfies the two-source rule alongside the vendor's own pricing table; confidence: medium reflects this field and modality, not the price. Input modalities text + image + video per OpenRouter; output text-only. Unlike qwen3-max/qwen3.7-max, this SKU carries a 50% batch-inference-discount badge on the International/Singapore row itself (confirmed on the vendor billing page) rather than only the China-mainland row — batch_input_per_mtok_usd/batch_output_per_mtok_usd are nonetheless omitted here because the discount applies per input-token tier and the schema's batch fields are flat (non-tiered), so a single number would misrepresent the 256K/1M split; noted here instead of guessed. Context caching discount available but DashScope does not publish a flat cache_read rate, so cache_read_per_mtok_usd is omitted. Tool calling and structured outputs supported per OpenRouter's supported_parameters. Hybrid thinking (OpenRouter reasoning.mandatory: false); thinking tokens billed at the standard output rate, hence reasoning_tokens_billed: true. Knowledge cutoff not published by the vendor; omitted rather than guessed. Deployment via DashScope (Model Studio) only; not yet confirmed on Bedrock, Vertex, Together, Fireworks, or Groq as a first-party offering. Re-verified 2026-07-19: International tiered prices unchanged ($0.25/$1.50 for 0-256K; $1.00/$4.00 for 256K-1M); 50% batch-discount badge still on the International row (batch fields still omitted for the tiered-vs-flat reason above); alias still resolves to qwen3.6-flash-2026-04-16; no deprecation notice. Re-verified 2026-08-11: International tiered prices unchanged ($0.25/$1.50 for 0-256K; $1.00/$4.00 for 256K-1M); 50% batch-discount badge still on the International row; alias still qwen3.6-flash-2026-04-16; no deprecation notice. New-launch check this pass surfaced qwen3.5-flash, qwen-flash (rolling family alias), and qwen3-omni-flash on the billing page, none currently in this dataset — deferred: the models catalog page does not expose their context window / max output tokens / knowledge cutoff for a two-source-compliant add in this pass; revisit next refresh. Re-verified 2026-09-02 (raw-HTML scrape of the Singapore table): International tiered prices unchanged ($0.25/$1.50 for 0-256K; $1.00/$4.00 for 256K-1M); 50% batch-discount badge still on the International row; alias still resolves to qwen3.6-flash-2026-04-16; no deprecation notice; qwen3.6-flash is absent from the vendor's models-page 'recommended' set (now qwen3.8-max/qwen3.7-plus/qwen3.8-flash) but remains separately priced with no deprecation banner on the billing page, so not marked deprecated. New-launch check this pass found four verifiable new rows added below: `qwen3.8-max` (flagship successor), `qwen3.8-flash` (fast-tier successor), and the two Groq-cross-check models `qwen3.6-27b`/`qwen3.8-27b` (both genuinely on DashScope International, priced there as primary per the two-source rule; Groq is a second direct host, not the canonical price). Also surfaced but deferred this pass (out of the requested scope and thinly documented): `qwen3.7-flash` (billing page shows tiered $0.03/$0.13 + $0.10/$0.40, but not in the models-page recommended set and possibly already superseded by qwen3.8-flash), `qwen3.6-35b-a3b` (open-weight sibling of qwen3.6-27b, $0.375/$2.25 flat on the Singapore table), and `qwen3.8-2.4t-a95b` (open-weight sibling of qwen3.8-max, Singapore $2.00/$6.00 flat with context caching; China-Beijing $1.65/$4.951) — revisit next refresh." }, { "provider": "Alibaba", "provider_url": "https://www.alibabacloud.com/product/modelstudio", "model_id": "qwen3.8-max", "display_name": "Qwen3.8-Max", "model_family": "Qwen 3.8 Max", "context_window": 1048576, "max_output_tokens": 131072, "input_per_mtok_usd": "2.00", "output_per_mtok_usd": "6.00", "modalities": { "input": ["text", "image", "video"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["dashscope"], "aggregators": ["openrouter"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-09-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.alibabacloud.com/help/en/model-studio/billing", "notes": "New row, added on the 2026-09-02 refresh. Qwen3.8-Max is the new flagship generation, general-availability successor line to qwen3-max/qwen3.7-max (both remain separately priced and active, so not marked deprecated/replaced_by). `qwen3.8-max` is a rolling alias currently equivalent to the dated snapshot `qwen3.8-max-0902`, also listed separately at identical pricing (both confirmed via raw-HTML scrape of the vendor billing page). Flat (non-tiered) DashScope International rate of $2.00 input / $6.00 output per MTok across the full 0-1M token range; no batch-inference-discount badge on the International row (context-caching-discount blockquote only; the China-Beijing row carries both a discounted rate of $1.65/$4.951 and a batch badge — not priced here, DashScope International is canonical). Cross-verified against OpenRouter (openrouter.ai/api/v1/models, qwen/qwen3.8-max: prompt $0.000002/completion $0.000006 per token = exactly $2.00/$6.00, matching the DashScope list price exactly) and against llm-stats.com/api/models/qwen3.8-max, whose per-host provider rows (fireworks, novita) also list $2/$6 at max_output_tokens 131,072 (deepinfra and together show slightly different context/output caps and $1.65-$2.50 rates — those are third-party host repricing, not the DashScope figure used here). Context window and max_output_tokens are not published on the vendor billing/models pages directly; sourced from OpenRouter (context_length 1,000,000, top_provider.max_completion_tokens 131,072) and corroborated by llm-stats.com's fireworks/novita provider rows (131,072 max output on both) — satisfies the two-source rule; confidence: medium reflects these two fields and the modality below, not the price. Input modalities text + image + video per OpenRouter's architecture tag (text+image+video->text) and llm-stats.com (multimodal: true, vision/video tags true) — a departure from the text-only qwen3-max/qwen3.7-max lineage; OpenRouter's per-model architecture tags were spot-checked against qwen3.7-max (correctly tagged text->text) to confirm the tags are model-specific, not a generic template. `reasoning`/`reasoning_effort` present in OpenRouter's supported_parameters confirms hybrid thinking support; the vendor billing table lists a single \"Non-Thinking and Thinking modes\" price column (same rate for both), hence reasoning_tokens_billed: true. Tool calling and structured outputs supported per OpenRouter's supported_parameters (tools/tool_choice, response_format/structured_outputs). Context caching discount available but DashScope does not publish a flat cache_read rate, so cache_read_per_mtok_usd is omitted rather than guessed (OpenRouter's own cache figures, e.g. $0.25 read, do not match a clean fraction of the DashScope list price and are not used). Knowledge cutoff not published by the vendor; omitted rather than guessed. Deployment via DashScope (Model Studio) only at last_verified; also listed on Fireworks, Novita, Together, and DeepInfra as third-party hosts (per llm-stats.com) but not added to deployment_options here — only vendor-direct plus hosts explicitly cross-checked this pass (Groq, on the 27b rows below) are recorded. Not marked `featured` in this pass (flagship candidate, but flipping the flag requires regenerating web/lib/featured.lock.json, which is out of scope for a costs/llm.json-only edit — flagged for the controller as a featured candidate)." }, { "provider": "Alibaba", "provider_url": "https://www.alibabacloud.com/product/modelstudio", "model_id": "qwen3.8-flash", "display_name": "Qwen3.8-Flash", "model_family": "Qwen 3.8 Flash", "context_window": 1048576, "max_output_tokens": 131072, "input_per_mtok_usd": "0.15", "output_per_mtok_usd": "0.47", "modalities": { "input": ["text", "image", "video"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["dashscope"], "aggregators": ["openrouter"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-09-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.alibabacloud.com/help/en/model-studio/billing", "notes": "New row, added on the 2026-09-02 refresh. Qwen3.8-Flash is the fast/efficient tier of the Qwen3.8 series, on the vendor's models-page 'recommended models' list alongside qwen3.8-max and qwen3.7-plus. Flat (non-tiered) DashScope International rate of $0.15 input / $0.47 output per MTok across the full 0-1M token range (confirmed via raw-HTML scrape of the vendor billing page); context-caching-discount blockquote present, no batch-inference-discount badge on the International row. Cross-verified against OpenRouter (openrouter.ai/api/v1/models, qwen/qwen3.8-flash: prompt $0.00000015/completion $0.00000047 per token = exactly $0.15/$0.47, matching the DashScope list price exactly) and against llm-stats.com/api/models/qwen3.8-flash's novita provider row (also $0.15/$0.47 exactly, max_output_tokens 131,072). Context window and max_output_tokens not published on the vendor billing/models pages directly; sourced from OpenRouter (context_length 1,000,000, top_provider.max_completion_tokens 131,072) and llm-stats.com's novita row (131,072) and native_context_length tag (1,000,000) — satisfies the two-source rule; llm-stats.com's own summary tag lists a slightly different 128,000 max-output figure, so confidence: medium reflects this field (and the modality below), not the price. Input modalities text + image + video per OpenRouter's architecture tag and llm-stats.com (multimodal: true, video/vision tags true); output text-only. The vendor billing table's price row for this model lacks the explicit \"Non-Thinking mode / Thinking mode\" two-column split seen on the qwen3.6-flash/qwen3.7-plus/qwen3.8-max sibling tables (single output-price column instead), but OpenRouter's supported_parameters for this model include `reasoning`/`include_reasoning` (llm-stats.com's description also calls it 'a multimodal reasoning model' with a 256K thinking budget tag), so this is treated as a hybrid-thinking model with reasoning_tokens_billed: true, billed at the standard output rate; the single-column table format is read as a formatting choice rather than evidence of no-thinking support. Tool calling and structured outputs supported per OpenRouter's supported_parameters. Context caching discount available but DashScope does not publish a flat cache_read rate, so cache_read_per_mtok_usd is omitted rather than guessed. Knowledge cutoff not published by the vendor; omitted rather than guessed. Deployment via DashScope (Model Studio) only at last_verified; also listed on Novita as a third-party host per llm-stats.com but not added to deployment_options here, consistent with how this dataset prices Alibaba rows at the DashScope/vendor-direct rate." }, { "provider": "Alibaba", "provider_url": "https://www.alibabacloud.com/product/modelstudio", "model_id": "qwen3.6-27b", "display_name": "Qwen3.6-27B", "model_family": "Qwen 3.6", "context_window": 262144, "max_output_tokens": 65536, "input_per_mtok_usd": "0.60", "output_per_mtok_usd": "3.60", "modalities": { "input": ["text", "image", "video"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": false, "reasoning_tokens_billed": true, "deployment_options": ["dashscope", "groq"], "aggregators": ["openrouter"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-09-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.alibabacloud.com/help/en/model-studio/billing", "notes": "New row, added on the 2026-09-02 refresh, triggered by a cross-check against a sibling provider pass that found `qwen/qwen3.6-27b` hosted on Groq (console.groq.com/docs/models) at $0.60/$3.00 per MTok. Confirmed this is a genuine Alibaba release with its own DashScope International listing, not a Groq-exclusive open-weight drop: `qwen3.6-27b` (a dense 27B open-weight model, Apache 2.0, released 2026-04-21 per llm-stats.com) is priced on the Singapore table at flat $0.60 input / $3.60 output per MTok, 0` reasoning block followed by the answer, and those reasoning tokens are billed at the output rate, hence reasoning_tokens_billed: true. Knowledge cutoff intentionally omitted: Sonar Reasoning Pro fetches the live web at query time. supports_tool_use set conservatively to false because Perplexity exposes web search as the built-in capability and recommends the separate Agent API for production tool-using agents. Structured outputs supported via response_format. Perplexity API is native-only (no Bedrock / Vertex / Azure / Together / Fireworks / Groq first-party deployment). Cache and batch APIs are not published. Note: the older `sonar-reasoning` SKU is no longer listed in Perplexity's current model lineup at last_verified. Re-verified 2026-08-11: $2/$8 per MTok and $6/$10/$14 per-1K-request search-context tiers unchanged on the vendor pricing page; model still listed as active on the Sonar models page. `sonar-deep-research` remains the only other SKU in the current lineup; not added as a row because its schema-required max_output_tokens is not published by Perplexity (model card and OpenRouter listing both omit it) — deferred rather than guessed. Re-verified 2026-09-02: $2/$8 per MTok and $6/$10/$14 per-1K-request search-context tiers unchanged; model still active on both the pricing page and the Sonar models page. `sonar-deep-research` still lacks a published max_output_tokens (pricing page now shows input/output/citation/reasoning/search-query rates for it but no token-limit field), so it remains deferred. Heads-up for the next pass: both pages now carry the banner \"Sonar Chat Completions is now Agent API. Sonar will be supported until September 27, 2026\" — not treated as deprecated here (still priced, still listed as active, no successor model_id given), but the sunset date is close enough to the next refresh cycle that it should be re-checked then; do not set deprecated_at speculatively ahead of an actual retirement." }, { "provider": "Anthropic", "provider_url": "https://www.anthropic.com", "model_id": "claude-opus-4-6", "display_name": "Claude Opus 4.6", "model_family": "Claude 4", "knowledge_cutoff": "2025-05-31", "aliases": ["anthropic.claude-opus-4-6-v1"], "context_window": 1000000, "max_output_tokens": 128000, "input_per_mtok_usd": "5.0", "output_per_mtok_usd": "25.0", "cache_read_per_mtok_usd": "0.5", "cache_write_per_mtok_usd": "6.25", "batch_input_per_mtok_usd": "2.5", "batch_output_per_mtok_usd": "12.5", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native", "bedrock", "vertex"], "last_verified": "2026-09-02", "last_changed_at": "2026-02-05", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://platform.claude.com/docs/en/about-claude/pricing", "notes": "New row this refresh: fully active legacy listing that was present on Anthropic's pricing page and models overview but previously missing from this dataset (sits between Opus 4.5 and 4.7). Pricing identical to Opus 4.5/4.7/4.8. Cache hit is 0.1x base input ($0.50/MTok); 5-minute cache write is 1.25x ($6.25/MTok); 1-hour cache write is 2x ($10/MTok). Batch API discounts both input and output by 50%. 1M context window at standard pricing, 128k max output; uses the pre-4.7 tokenizer (~750k words per 1M tokens). Supports both extended thinking and adaptive thinking; thinking output tokens billed at the output rate. Fast mode is no longer available on this model as of 2026-06-29 (speed:'fast' requests run and bill at standard rates). Reliable knowledge cutoff May 2025; training data cutoff Aug 2025. last_changed_at is the launch date 2026-02-05, INFERRED from Anthropic's tentative retirement date 'not sooner than February 5, 2027' (launch + 1 year pattern confirmed against Fable 5, Opus 4.8, and Sonnet 5); prices verified as current, not historical. Bedrock ID recorded as shown in the models overview ('anthropic.claude-opus-4-6-v1')." }, { "provider": "Anthropic", "provider_url": "https://www.anthropic.com", "model_id": "claude-opus-4-1-20250805", "display_name": "Claude Opus 4.1", "model_family": "Claude 4", "knowledge_cutoff": "2025-01-31", "aliases": [ "claude-opus-4-1", "anthropic.claude-opus-4-1-20250805-v1:0", "claude-opus-4-1@20250805" ], "context_window": 200000, "max_output_tokens": 32000, "input_per_mtok_usd": "15.0", "output_per_mtok_usd": "75.0", "cache_read_per_mtok_usd": "1.5", "cache_write_per_mtok_usd": "18.75", "batch_input_per_mtok_usd": "7.5", "batch_output_per_mtok_usd": "37.5", "modalities": { "input": ["text", "image"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": false, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["bedrock", "vertex"], "deprecated_at": "2026-06-05", "replaced_by_model_id": "claude-opus-4-8", "last_verified": "2026-09-02", "last_changed_at": "2026-06-05", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://platform.claude.com/docs/en/about-claude/pricing", "notes": "RETIRED on the Claude API (and Claude Platform on AWS) on 2026-08-05 as scheduled; requests to claude-opus-4-1-20250805 on those surfaces now fail. Still available on Amazon Bedrock (status: deprecated, not retired, per the Bedrock legacy model table) and Google Cloud Vertex AI under their own partner retirement schedules — deployment_options narrowed to bedrock/vertex only, 'native' removed. Recommended replacement claude-opus-4-8. Pricing unchanged and still listed for reference on Anthropic's pricing page: $15/$75 base; cache hit $1.50/MTok (0.1x); 5-minute cache write $18.75/MTok (1.25x); 1-hour cache write $30/MTok (2x); batch $7.50/$37.50 (50% off). 200k context window, 32k max output. Supports extended thinking (not adaptive); thinking output tokens billed at the output rate. Reliable knowledge cutoff Jan 2025; training data cutoff Mar 2025. Cross-verified against platform.claude.com/docs/en/about-claude/pricing and platform.claude.com/docs/en/about-claude/model-deprecations, and Amazon Bedrock's legacy model table, on 2026-09-02." }, { "provider": "Google", "provider_url": "https://deepmind.google", "model_id": "gemini-3.6-flash", "display_name": "Gemini 3.6 Flash", "featured": true, "model_family": "Gemini 3.6", "context_window": 1048576, "max_output_tokens": 65536, "input_per_mtok_usd": "0.75", "output_per_mtok_usd": "3.75", "cache_read_per_mtok_usd": "0.075", "cache_storage_per_mtok_per_hour_usd": "0.50", "batch_input_per_mtok_usd": "0.375", "batch_output_per_mtok_usd": "1.875", "modalities": { "input": ["text", "image", "audio", "video"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": true, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-09-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://ai.google.dev/pricing", "notes": "New model added 2026-07-27. Context window and max_output_tokens (65,536) confirmed via ai.google.dev/gemini-api/docs/models/gemini-3.6-flash. Thinking is supported; thinking tokens billed at the output rate, hence reasoning_tokens_billed: true. Multimodal input (text, image, audio, video per the model documentation); text-only output. Tool use and structured outputs supported. Caching and batch APIs available. Deployment via native API only at this verification. Knowledge cutoff and release date not published by Google; omitted rather than guessed. No long-context tiering published (unlike gemini-2.5-pro). Free tier published as per-minute RPM/TPM only, not per-day; free_tier omitted. Cross-verified against secondary source: ai.google.dev/gemini-api/docs/models/gemini-3.6-flash for context, output token limit, modalities, and capabilities. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: ai.google.dev/pricing now lists an introductory-discount price of $0.75 input / $3.75 output / $0.075 cache-read / $0.50 cache-storage-per-hour / $0.375 batch-input / $1.875 batch-output through 2026-12-31, reverting to the previous rate ($1.50/$7.50/$0.15/$1.00/$0.75/$3.75) on 2027-01-01; input/output/cache/batch fields updated to the currently-billed discounted rate and last_changed_at bumped accordingly. Cross-checked ai.google.dev/gemini-api/docs/models." }, { "provider": "Google", "provider_url": "https://deepmind.google", "model_id": "gemini-3.5-flash-lite", "display_name": "Gemini 3.5 Flash-Lite", "model_family": "Gemini 3.5", "context_window": 1048576, "max_output_tokens": 65536, "input_per_mtok_usd": "0.30", "output_per_mtok_usd": "2.50", "batch_input_per_mtok_usd": "0.15", "batch_output_per_mtok_usd": "1.25", "modalities": { "input": ["text", "image", "audio", "video"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": true, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-07-27", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://ai.google.dev/pricing", "notes": "New model added 2026-07-27. Prices confirmed from ai.google.dev/pricing: $0.30 input / $2.50 output per MTok; batch API at 50% off ($0.15/$1.25). Context window and max_output_tokens (65,536) confirmed via ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite. Thinking is supported; thinking tokens billed at the output rate, hence reasoning_tokens_billed: true. Multimodal input (text, image, audio, video per the model documentation); text-only output. Tool use and structured outputs supported. Batch API available. Context caching not available for this model (per ai.google.dev/pricing, which omits cache pricing for this SKU unlike gemini-3.6-flash). Deployment via native API only at this verification. Knowledge cutoff and release date not published by Google; omitted rather than guessed. No long-context tiering published. Free tier published as per-minute RPM/TPM only, not per-day; free_tier omitted. Cross-verified against secondary source: ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite for context, output token limit, modalities, caching availability, and capabilities. Pricing is the primary-source figure pending broader aggregator publication. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices unchanged against ai.google.dev/pricing (explicitly re-confirmed no context-caching price is published for this SKU, unlike gemini-3.6-flash/gemini-3.7-flash); audio input confirmed billed at the same standard rate as text/image/video (no separate audio premium); cross-checked ai.google.dev/gemini-api/docs/models." }, { "provider": "Google", "provider_url": "https://deepmind.google", "model_id": "gemini-3.7-flash", "display_name": "Gemini 3.7 Flash", "model_family": "Gemini 3.7", "context_window": 1048576, "max_output_tokens": 65536, "input_per_mtok_usd": "0.75", "output_per_mtok_usd": "3.75", "cache_read_per_mtok_usd": "0.075", "cache_storage_per_mtok_per_hour_usd": "0.50", "batch_input_per_mtok_usd": "0.375", "batch_output_per_mtok_usd": "1.875", "modalities": { "input": ["text", "image", "audio", "video"], "output": ["text"] }, "supports_tool_use": true, "structured_output": true, "supports_vision": true, "supports_audio_in": true, "supports_audio_out": false, "supports_pdf": true, "reasoning_tokens_billed": true, "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-09-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://ai.google.dev/pricing", "notes": "New GA model, added 2026-09-02; released August 2026 as \"the next iteration in the Gemini 3 series\" and described by Google as the latest and most capable Flash model, superseding gemini-3.6-flash without deprecating it. Standard pricing carries an introductory discount through 2026-12-31 ($0.75 input / $3.75 output / $0.075 cache-read / $0.50 cache-storage-per-hour / $0.375 batch-input / $1.875 batch-output per ai.google.dev/pricing), reverting to $1.50/$7.50/$0.15/$1.00/$0.75/$3.75 on 2027-01-01; fields capture the currently-billed discounted rate, identical in structure to sibling gemini-3.6-flash. Thinking supported at low/medium/high (minimal not offered); thinking tokens billed at the output rate, hence reasoning_tokens_billed=true. Multimodal input (text, image, audio, video, PDF); text-only output. Tool use, structured outputs, and context caching supported. Deployment via native API only at this verification; Vertex AI availability not confirmed. No long-context (>200k token) pricing tier published. Knowledge cutoff not published by Google; omitted rather than guessed. Free tier (AI Studio) published as per-minute RPM/TPM only, not per-day; free_tier omitted. Cross-verified against secondary source: ai.google.dev/gemini-api/docs/models/gemini-3.7-flash for context window, max output tokens, modalities, tool-use/structured-output support, and GA status." } ] }