{ "version": 2, "license": "CC-BY-4.0", "models": [ { "provider": "Deepgram", "provider_url": "https://deepgram.com", "model_id": "nova-3-monolingual", "display_name": "Nova-3 Monolingual", "featured": true, "price_per_minute_usd": "0.0048", "price_per_minute_batch_usd": "0.0043", "diarization_per_minute_usd": "0.002", "languages": ["en"], "streaming": true, "realtime": true, "diarization": "extra-cost", "last_verified": "2026-09-02", "last_changed_at": "2026-09-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://deepgram.com/pricing", "notes": "Pay-as-you-go tier (English). Growth tier (volume commitment) is $0.0042/min streaming, $0.0036/min pre-recorded. Diarization add-on is captured in diarization_per_minute_usd. Correction (2026-09-02): price_per_minute_batch_usd restored to $0.0043/min, matching the current deepgram.com/pricing structured data (schema.org Offer entries) and visible pricing table. The 2026-07-02 refresh had mistakenly recorded $0.0077/min here -- that figure is actually the struck-through, non-promotional REGULAR STREAMING rate (labeled 'Regular price' next to the promotional $0.0048/min streaming rate), not a pre-recorded price. Corroborated by third-party aggregators (e.g. convertaudiototext.com, llmreference.com) quoting $0.0043/min batch for Nova-3 mono." }, { "provider": "Deepgram", "provider_url": "https://deepgram.com", "model_id": "nova-3-multilingual", "display_name": "Nova-3 Multilingual", "price_per_minute_usd": "0.0058", "price_per_minute_batch_usd": "0.0052", "diarization_per_minute_usd": "0.002", "languages": "61+", "streaming": true, "realtime": true, "diarization": "extra-cost", "last_verified": "2026-09-02", "last_changed_at": "2026-09-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://deepgram.com/pricing", "notes": "Pay-as-you-go tier (multilingual, ~61 languages). Growth tier is $0.0050/min streaming, $0.0043/min pre-recorded. Diarization add-on is captured in diarization_per_minute_usd. Correction (2026-09-02): price_per_minute_batch_usd restored to $0.0052/min, matching the current deepgram.com/pricing structured data (schema.org Offer entries) and visible pricing table. The 2026-07-02 refresh had mistakenly recorded $0.0092/min here -- that figure is actually the struck-through, non-promotional REGULAR STREAMING rate (labeled 'Regular price' next to the promotional $0.0058/min streaming rate), not a pre-recorded price." }, { "provider": "Deepgram", "provider_url": "https://deepgram.com", "model_id": "nova-3-medical", "display_name": "Nova-3 Medical", "price_per_minute_usd": "0.0048", "price_per_minute_batch_usd": "0.0043", "diarization_per_minute_usd": "0.002", "languages": [ "en", "en-US", "en-AU", "en-CA", "en-GB", "en-IE", "en-IN", "en-NZ" ], "streaming": true, "realtime": true, "diarization": "extra-cost", "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-05-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://deepgram.com/learn/introducing-nova-3-medical-speech-to-text-api", "notes": "Medical-tuned Nova-3 for clinical transcription. English variants only (en, en-US, en-AU, en-CA, en-GB, en-IE, en-IN, en-NZ). Invoked via `model=nova-3-medical` in the Deepgram API. Pricing is not separately listed on the public pricing page; Deepgram's launch announcement quotes $0.0043/min pre-recorded, which matches the Nova-3 Monolingual batch rate (now confirmed current -- see nova-3-monolingual row). Streaming rate assumed equal to Nova-3 Monolingual ($0.0048/min PAYG); verify with Deepgram sales for production commitments." }, { "provider": "Deepgram", "provider_url": "https://deepgram.com", "model_id": "flux-general-en", "display_name": "Flux (English)", "price_per_minute_usd": "0.0065", "languages": ["en"], "streaming": true, "realtime": true, "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://deepgram.com/pricing", "notes": "New conversational speech recognition (CSR) model built for voice agents; adds turn-detection/end-of-turn modeling on top of transcription. Launched free for October 2025 promo; announcement post (https://deepgram.com/learn/introducing-flux-conversational-speech-recognition) does not quote a per-minute price, so pricing is single-sourced from deepgram.com/pricing (Pay-As-You-Go streaming $0.0065/min; Growth tier $0.0057/min), corroborated by a third-party pricing aggregator quoting the same figure. Correction (2026-09-02): the $0.0077/min figure is the struck-through, non-promotional REGULAR STREAMING rate shown next to the promotional $0.0065/min streaming price on deepgram.com/pricing -- it is not a pre-recorded/batch column. The structured pricing data (schema.org Offer entries) on the page lists no Pre-Recorded offer for Flux English at all, confirming price_per_minute_batch_usd should stay omitted here." }, { "provider": "Deepgram", "provider_url": "https://deepgram.com", "model_id": "flux-general-multi", "display_name": "Flux (Multilingual)", "price_per_minute_usd": "0.0078", "languages": ["en", "es", "fr", "de", "hi", "ru", "pt", "ja", "it", "nl"], "streaming": true, "realtime": true, "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://deepgram.com/pricing", "notes": "Multilingual conversational speech recognition (CSR) variant of Flux, launched after the English-only October 2025 debut (see press release: https://deepgram.com/learn/deepgram-launches-flux-multilingual-press-release). Covers 10 languages via `language_hint` on the `flux-general-multi` model_id (en, es, fr, de, hi, ru, pt, ja, it, nl). Pay-As-You-Go streaming $0.0078/min; Growth tier $0.0068/min. No pre-recorded/batch price is listed on the pricing page for this row (confirmed again 2026-09-02: no Pre-Recorded Offer entry for Flux Multilingual in the page's structured pricing data)." }, { "provider": "Deepgram", "provider_url": "https://deepgram.com", "model_id": "whisper-large", "display_name": "Whisper Cloud (Large)", "price_per_minute_usd": "0.0048", "languages": "99+", "streaming": false, "realtime": false, "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-09-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://deepgram.com/pricing", "notes": "Deepgram-hosted OpenAI Whisper, pre-recorded only ('Live streaming is not available with Deepgram Whisper Cloud' per docs -- use Nova-3 for streaming). Invoked via `model=whisper-large` (defaults to large-v2); other sizes available (`whisper-tiny`, `whisper-base`, `whisper-small`, `whisper-medium`). Pricing page (schema.org Offer 'Pre-Recorded - Whisper Large') lists $0.0048/min flat across both Pay-As-You-Go and Growth tiers -- no volume discount, unlike the Nova/Flux rows. Officially supports 99 languages per developers.deepgram.com/docs/deepgram-whisper-cloud. Not available on the EU endpoint; docs note Whisper models are 'less scalable' than Deepgram's proprietary models. confidence: medium because the secondary source (developer docs) confirms the model_id and capabilities but not the price itself -- only the pricing page quotes a number." }, { "provider": "AssemblyAI", "provider_url": "https://www.assemblyai.com", "model_id": "universal-2", "display_name": "Universal-2", "price_per_minute_usd": "0.0025", "languages": "99+", "streaming": false, "realtime": false, "diarization": "extra-cost", "last_verified": "2026-09-02", "last_changed_at": "2026-05-05", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.assemblyai.com/pricing", "notes": "Pre-recorded (file-based) only — broadest AssemblyAI language coverage. Published as $0.15/hr. Diarization add-on +$0.02/hr; the pricing page no longer lists the +$0.065/hr experimental diarization tier seen in a prior pass (llms/pricing.md, dated 2026-05-29, still shows it — treated as stale). For real-time use, see universal-3-5-pro (streaming) / universal-streaming-english / universal-streaming-multilingual." }, { "provider": "AssemblyAI", "provider_url": "https://www.assemblyai.com", "model_id": "universal-streaming", "display_name": "Universal-Streaming", "price_per_minute_usd": "0.0025", "languages": ["en"], "streaming": true, "realtime": true, "diarization": "extra-cost", "deprecated_at": "2026-07-02", "replaced_by_model_id": "universal-streaming-english", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-05-05", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.assemblyai.com/pricing", "notes": "Superseded by explicit `universal-streaming-english` / `universal-streaming-multilingual` model_ids — the streaming AsyncAPI spec (docs.assemblyai.com) no longer enumerates a bare `universal-streaming` value. Price unchanged at $0.15/hr; treated as a rename, not a price change. confidence:medium because no explicit vendor deprecation notice was found, only enum absence." }, { "provider": "AssemblyAI", "provider_url": "https://www.assemblyai.com", "model_id": "universal-streaming-english", "display_name": "Universal-Streaming English", "price_per_minute_usd": "0.0025", "languages": ["en"], "streaming": true, "realtime": true, "diarization": "extra-cost", "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.assemblyai.com/pricing", "notes": "English-only streaming model, explicit successor id to the former bare `universal-streaming`. Published as $0.15/hr. Higher-tier streaming (Universal-3.5 Pro Realtime / Streaming, model_id `universal-3-5-pro`) is $0.45/hr ($0.0075/min)." }, { "provider": "AssemblyAI", "provider_url": "https://www.assemblyai.com", "model_id": "universal-3-pro", "display_name": "Universal-3.5 Pro", "featured": true, "price_per_minute_usd": "0.0035", "languages": [ "en", "es", "fr", "de", "it", "pt", "ar", "da", "nl", "fi", "he", "hi", "ja", "zh", "no", "sv", "tr", "vi" ], "streaming": false, "realtime": false, "diarization": "extra-cost", "deployment_options": ["native"], "deprecated_at": "2026-07-10", "replaced_by_model_id": "universal-3-5-pro", "last_verified": "2026-09-02", "last_changed_at": "2026-05-26", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.assemblyai.com/pricing", "notes": "AssemblyAI's highest-accuracy pre-recorded model with native code-switching, documented at 18 languages per assemblyai.com/docs/getting-started/models and assemblyai.com/llms/models.md; unsupported languages fall back to Universal-2. Deprecated: the changelog entry \"Deprecated: universal-3-pro and universal for New and Inactive Accounts\" (2026-07-10) cut off new/inactive accounts; the entry \"Default speech model changing to Universal-3.5 Pro on September 2, 2026\" confirms requests pinned to `universal-3-pro` now return an error as of today (2026-09-02) instead of being auto-routed — hard cutover complete. Legacy request value `best` now routes to `universal-3-5-pro` (`nano` still routes to universal-2). Superseded by the `universal-3-5-pro` row at the same $0.21/hr price — a rename, not a price change. Featured flag kept per runbook policy (indexed compare pages stay live with a deprecation banner)." }, { "provider": "AssemblyAI", "provider_url": "https://www.assemblyai.com", "model_id": "universal-3-5-pro", "display_name": "Universal-3.5 Pro", "price_per_minute_usd": "0.0035", "languages": [ "en", "es", "fr", "de", "it", "pt", "ar", "da", "nl", "fi", "he", "hi", "ja", "zh", "no", "sv", "tr", "vi" ], "streaming": false, "realtime": false, "diarization": "extra-cost", "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-09-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.assemblyai.com/pricing", "notes": "Canonical async model_id as of 2026-09-02, renamed from `universal-3-pro` (see that deprecated row). Confirmed via three current first-party sources: the docs/getting-started/models secondary source links \"Universal-3.5 Pro\" to /docs/pre-recorded-audio/universal-3-5-pro, whose code samples across Python/JS/TS/cURL all show `\"speech_models\": [\"universal-3-5-pro\"]`; the changelog entry \"Default speech model changing to Universal-3.5 Pro on September 2, 2026\" states `universal-3-pro` now hard-errors and this id is the new default when no model is pinned; and the pricing page (primary) confirms $0.21/hr unchanged. Same 18-language coverage as the prior `universal-3-pro` row; unsupported languages fall back to Universal-2. This exact model_id string is shared with the streaming row below (`universal-3-5-pro`, streaming:true, realtime:true, $0.45/hr) — AssemblyAI unified the identifier across the async (`/v2/transcript`) and streaming (`wss://streaming.assemblyai.com/v3/ws`) endpoints, confirmed on the dedicated streaming model-selection reference. The two products are billed separately and are distinguished in this dataset by the `streaming` field, not by model_id — first duplicate model_id in this file; flagging for review. A separate Sync API product also uses this id (≤2min/request, $0.45/hr via an `X-AAI-Model: universal-3-5-pro` header at sync.assemblyai.com) but is not represented here — a third row with this model_id and streaming:false at a different price is not representable in the current schema; deferred." }, { "provider": "AssemblyAI", "provider_url": "https://www.assemblyai.com", "model_id": "universal-3-pro-streaming", "display_name": "Universal-3 Pro Streaming", "price_per_minute_usd": "0.0075", "languages": [ "en", "es", "fr", "de", "it", "pt", "ar", "da", "nl", "fi", "he", "hi", "ja", "zh", "no", "sv", "tr", "vi" ], "streaming": true, "realtime": true, "diarization": "extra-cost", "deployment_options": ["native"], "deprecated_at": "2026-07-19", "replaced_by_model_id": "u3-rt-pro", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-05-26", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.assemblyai.com/pricing", "notes": "Real-time streaming tier for voice agents, now branded \"Universal-3.5 Pro Realtime\". Deprecated in this dataset 2026-07-19: for a second consecutive refresh (2026-07-13 and 2026-07-19) the pricing page and the assemblyai.com/llms/models.md quick reference both give `u3-rt-pro` (alias `u3-pro`) as the API identifier, and no vendor source uses the literal string `universal-3-pro-streaming`. This was the repeat confirmation the 2026-07-13 refresh required before renaming; superseded by the `u3-rt-pro` row at the same $0.45/hr price (a rename, not a price change). confidence:medium because this id was never published by the vendor, so there is no explicit vendor deprecation notice for it." }, { "provider": "AssemblyAI", "provider_url": "https://www.assemblyai.com", "model_id": "u3-rt-pro", "display_name": "Universal-3.5 Pro Realtime", "aliases": ["u3-pro"], "price_per_minute_usd": "0.0075", "languages": [ "en", "es", "fr", "de", "it", "pt", "ar", "da", "nl", "fi", "he", "hi", "ja", "zh", "no", "sv", "tr", "vi" ], "streaming": true, "realtime": true, "diarization": "extra-cost", "deployment_options": ["native"], "deprecated_at": "2026-09-02", "replaced_by_model_id": "universal-3-5-pro", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.assemblyai.com/pricing", "notes": "Canonical id for the streaming tier this dataset previously tracked as `universal-3-pro-streaming` (GA 2026-03-03 per assemblyai.com/llms/models.md; `u3-pro` kept as a backward-compatible alias that routes here). Deprecated: the current dedicated streaming model-selection reference (docs/streaming/select-the-speech-model) lists only three available streaming models — `universal-3-5-pro` (recommended, also the default), `universal-streaming-english`, `universal-streaming-multilingual` — `u3-rt-pro`/`u3-pro` no longer appear. The streaming migration guide \"Universal Streaming to Universal-3.5 Pro Streaming\" and the async-migration changelog entry (\"u3-pro-rt (Universal-3 Pro Realtime) Redirected to universal-3-5-pro — no action required\") corroborate the same target id. Superseded by the `universal-3-5-pro` row at the same $0.45/hr price — a rename, not a price change. confidence:medium because, unlike the async id's explicit hard-error notice, no changelog entry states `u3-rt-pro` itself now errors — only that it redirects, so treat with the same caution as the prior `universal-streaming` rename." }, { "provider": "AssemblyAI", "provider_url": "https://www.assemblyai.com", "model_id": "universal-3-5-pro", "display_name": "Universal-3.5 Pro Realtime", "price_per_minute_usd": "0.0075", "languages": [ "en", "es", "fr", "de", "it", "pt", "ar", "da", "nl", "fi", "he", "hi", "ja", "zh", "no", "sv", "tr", "vi" ], "streaming": true, "realtime": true, "diarization": "extra-cost", "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-09-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.assemblyai.com/pricing", "notes": "Canonical streaming model_id as of 2026-09-02 — unified with the async model. The identical string `universal-3-5-pro` is used both for the async endpoint (see the streaming:false row above, $0.21/hr) and for this streaming WebSocket product ($0.45/hr); this is the first duplicate model_id in this dataset, flagging for review since site tooling assumes model_id is a stable per-row key. Confirmed via the dedicated streaming model-selection reference (docs/streaming/select-the-speech-model: `\"speech_model\": \"universal-3-5-pro\"`, marked Recommended and used as the default when the parameter is omitted) and the streaming migration guide \"Universal Streaming to Universal-3.5 Pro Streaming\". Supersedes `u3-rt-pro` (alias `u3-pro`) — see that row's deprecation. Pricing page (primary) brands this \"Universal-3.5 Pro Realtime\" at $0.45/hr; current docs call it \"Universal-3.5 Pro Streaming\" — same product, same price, unchanged from the prior u3-rt-pro row. Real-time inline diarization add-on +$0.12/hr (`speaker_labels: true`), billed on session duration not audio duration." }, { "provider": "AssemblyAI", "provider_url": "https://www.assemblyai.com", "model_id": "universal-streaming-multilingual", "display_name": "Universal-Streaming Multilingual", "price_per_minute_usd": "0.0025", "languages": ["en", "es", "pt", "de", "fr", "it"], "streaming": true, "realtime": true, "diarization": "extra-cost", "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-05-26", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.assemblyai.com/pricing", "notes": "Multilingual streaming variant covering EN/ES/PT/DE/FR/IT at the same $0.15/hr rate as the English-only universal-streaming-english. Good balance of cost and latency for voice agents." }, { "provider": "AssemblyAI", "provider_url": "https://www.assemblyai.com", "model_id": "whisper-streaming", "display_name": "Whisper-Streaming", "price_per_minute_usd": "0.0050", "languages": "99+", "streaming": true, "realtime": true, "diarization": "extra-cost", "deployment_options": ["native"], "deprecated_at": "2026-07-02", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-05-26", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.assemblyai.com/pricing", "notes": "OpenAI Whisper served via AssemblyAI's streaming infrastructure with 99+ language coverage. Published as $0.30/hr as of last verification, but no longer listed on the pricing page, docs/getting-started/models, the llms/models.md quick reference, or the changelog — still absent as of 2026-09-02, reconfirming silent discontinuation. No in-file replacement (no surviving streaming model offers 99+ language coverage); confidence:medium since no explicit vendor deprecation notice was found, only absence across all checked sources." }, { "provider": "Cartesia", "provider_url": "https://cartesia.ai", "model_id": "ink-whisper", "display_name": "Ink Whisper", "price_per_minute_usd": "0.003", "price_per_minute_batch_usd": "0.0015", "languages": "99+", "streaming": true, "realtime": true, "diarization": "unsupported", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://cartesia.ai/pricing", "notes": "model_id corrected from the invented 'ink-1' to the actual API identifier 'ink-whisper' (docs.cartesia.ai STT AsyncAPI spec enumerates only ink-2 and ink-whisper; latest dated snapshot is ink-whisper-2025-06-04). cartesia.ai/pricing no longer publishes an explicit $/min table -- STT is now billed in credits (docs.cartesia.ai/pricing.md: ink-whisper = 1 credit/sec streaming, 1 credit per 2 sec via the batch endpoint) with credit value varying by plan tier. Converting at the Pro-tier rate ($5/mo for 100,000 credits = $0.00005/credit, the same rate implied by the prior $0.003/min figure) yields $0.003/min streaming (unchanged) and $0.0015/min batch (new price_per_minute_batch_usd field, derived not published, hence confidence: medium). Languages widened from '42+' to '99+' per docs.cartesia.ai/build-with-cartesia/stt/older-models. Still the same STT product as the Sonic TTS provider; supports both batch (/stt) and manual realtime websocket. 2026-07-13 re-verify: still active, no deprecation banner; docs.cartesia.ai/build-with-cartesia/stt/older-models confirms 99-language support and ink-whisper-2025-06-04 as latest snapshot unchanged; docs.cartesia.ai/pricing.md credit rates (1 credit/sec streaming, 1 credit/2sec batch) and cartesia.ai/pricing Pro-tier allotment ($5/mo, ~9h16m ink-2 hours implying $0.00005/credit) unchanged, so both prices confirmed unchanged. 2026-07-19 re-verify: docs.cartesia.ai/pricing.md still lists ink-whisper at 1 credit/sec streaming and 1 credit/2sec batch; cartesia.ai/pricing plan allotments unchanged ($5 Pro ~9h16m); no deprecation notice (only 'Older STT Models' placement); prices unchanged. 2026-07-27 re-verify: both sources confirm ink-whisper remains stable with credit rates (1 credit/sec streaming, 1 credit/2sec batch) and Pro-tier plan allotment (~9h16m for $5, implying $0.00005/credit) unchanged; no new Cartesia STT models found; prices unchanged. 2026-08-11 re-verify: docs.cartesia.ai/build-with-cartesia/stt/older-models confirms ink-whisper-2025-06-04 still Stable, 99-language support unchanged, still positioned under 'Older Models' recommending Ink 2 for English turn-detection use cases (no formal deprecation date set); docs.cartesia.ai/pricing.md credit rates (1 credit/sec streaming, 1 credit/2sec batch) unchanged; cartesia.ai/pricing Pro-tier allotment (~9h16m for $5) unchanged; prices unchanged; no new Cartesia STT models found. 2026-09-02 re-verify: docs.cartesia.ai/build-with-cartesia/stt/older-models confirms ink-whisper-2025-06-04 still Stable, 99-language support unchanged; docs.cartesia.ai/pricing.md credit rates (1 credit/sec streaming, 1 credit/2sec batch) unchanged; cartesia.ai/pricing Pro-tier allotment (~9h16m for $5) unchanged; prices unchanged; no new Cartesia STT models found; no deprecation notices." }, { "provider": "Cartesia", "provider_url": "https://cartesia.ai", "model_id": "ink-2", "display_name": "Ink 2", "featured": true, "price_per_minute_usd": "0.009", "languages": ["en"], "streaming": true, "realtime": true, "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://cartesia.ai/pricing", "notes": "New model launched 2026-05-22 (docs.cartesia.ai/build-with-cartesia/stt/latest): Cartesia's fastest streaming STT with built-in turn detection, replacing the need for separate VAD. English only ('en') as of this refresh; batch (/stt) support listed as 'coming soon', so only the streaming/auto-websocket endpoint is priced here (no price_per_minute_batch_usd yet). Billed at 3 credits/sec (docs.cartesia.ai/pricing.md); converted at the Pro-tier rate of $0.00005/credit ($5/mo for 100,000 credits) = $0.009/min, consistent with the ~9h16m Ink-2 allotment shown for the $5 Pro plan on cartesia.ai/pricing (9.27h => ~$0.009/min). Diarization support unconfirmed in docs, field omitted pending verification. Not a replacement for ink-whisper -- both remain listed as stable/active models. 2026-07-13 re-verify: still active/stable, still English-only, batch endpoint still listed as not yet available, turn-detection lifecycle events (turn.start/update/eager_end/resume/end) confirmed live on docs.cartesia.ai/build-with-cartesia/stt/latest; credit rate (3 credits/sec) and Pro-tier allotment (~9h16m for $5) unchanged, so $0.009/min confirmed unchanged. 2026-07-19 re-verify: still latest/stable on docs.cartesia.ai/build-with-cartesia/stt/latest, still English-only, batch (/stt) still not available per docs.cartesia.ai/pricing.md (3 credits/sec streaming unchanged); cartesia.ai/pricing Pro allotment ~9h16m/$5 unchanged; price unchanged; no new Cartesia STT models found on either source. 2026-07-27 re-verify: still latest/stable on docs.cartesia.ai/build-with-cartesia/stt/latest, still English-only, batch still not available per docs.cartesia.ai/pricing.md (3 credits/sec streaming unchanged); cartesia.ai/pricing Pro allotment ~9h16m/$5 unchanged; price unchanged; no new Cartesia STT models; no deprecation notices. 2026-08-11 re-verify: docs.cartesia.ai/build-with-cartesia/stt/latest confirms Ink 2 still Stable, still English-only, turn-detection lifecycle events unchanged; docs.cartesia.ai/pricing.md batch (/stt) still not available for ink-2, streaming rate still 3 credits/sec; cartesia.ai/pricing Pro-tier allotment (~9h16m for $5) unchanged; price unchanged; no new Cartesia STT models found; no deprecation notices. 2026-09-02 re-verify: docs.cartesia.ai/build-with-cartesia/stt/latest confirms Ink 2 still Stable, still English-only, released 2026-05-22, turn-detection lifecycle events unchanged; docs.cartesia.ai/pricing.md batch (/stt) still not available for ink-2, streaming rate still 3 credits/sec; cartesia.ai/pricing Pro-tier allotment (~9h16m for $5) unchanged; price unchanged; no new Cartesia STT models found; no deprecation notices." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "whisper-1", "display_name": "Whisper", "price_per_minute_usd": "0.006", "languages": "99+", "streaming": false, "realtime": false, "diarization": "unsupported", "deprecated_at": "2026-08-26", "replaced_by_model_id": "gpt-transcribe", "last_verified": "2026-09-02", "last_changed_at": "2026-05-05", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/pricing", "notes": "Billed to the nearest second. Pre-recorded only. Price unchanged at $0.006/min. openai.com/api/pricing/ still 403s and platform.openai.com/docs/pricing 301-redirects to developers.openai.com/api/docs/pricing (developers.openai.com/api/docs/models likewise 301s from platform.openai.com/docs/models). 2026-08-11 re-verify: as of that pass whisper-1 no longer appeared on the /api/docs/models catalog listing at all — rate re-confirmed directly on the model detail page developers.openai.com/api/docs/models/whisper-1 ('Default snapshot: whisper-1', $0.006/min, no deprecation banner), cross-checked against /api/docs/deprecations (whisper-1 not present at that time). 2026-09-02 re-verify: price still $0.006/min (unchanged) on developers.openai.com/api/docs/pricing and the model detail page. BUT /api/docs/deprecations now (as of the 2026-08-26 entry) lists whisper-1 under 'Transcription models': \"On August 26, 2026, we notified developers using whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-transcribe-diarize of their deprecation and removal from the API on February 26, 2027\", table row 'Feb 26, 2027 | whisper-1 | gpt-live-transcribe or gpt-transcribe'. deprecated_at set to the 2026-08-26 announcement date (matching this dataset's convention, e.g. the llm.json gpt-5 row); shutdown is 2027-02-26. replaced_by_model_id set to gpt-transcribe (same non-realtime /v1/audio/transcriptions file-upload shape and flat per-minute pricing as this row) — vendor also names gpt-live-transcribe as an acceptable replacement for realtime use cases." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-4o-transcribe", "display_name": "GPT-4o Transcribe", "featured": true, "price_per_minute_usd": "0.006", "languages": "99+", "streaming": false, "realtime": false, "diarization": "unsupported", "deprecated_at": "2026-08-26", "replaced_by_model_id": "gpt-transcribe", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-13", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/pricing", "notes": "GPT-4o-powered transcription model served via the same /v1/audio/transcriptions endpoint as whisper-1; higher quality with prompting support for improved accuracy. Billed per-token ($2.50/1M input audio tokens, $10.00/1M output tokens); developers.openai.com/api/docs/pricing publishes this as an 'Estimated cost' of $0.006/minute (confidence: medium — the real bill tracks token usage, not a flat per-minute rate like whisper-1). 25 MB file-size limit. Supports `stream=true` for incremental text output on an already-uploaded file, which is not the same as live/realtime audio-in. 2026-08-11 re-verify: still active on both sources, price unchanged. 2026-09-02 re-verify: price still $0.006/min estimated ($2.50/$10.00 per 1M tokens), unchanged, on developers.openai.com/api/docs/pricing. Now listed on /api/docs/deprecations 'Transcription models' entry (announced 2026-08-26): removal from the API on 2027-02-26, vendor-recommended replacement gpt-live-transcribe or gpt-transcribe. deprecated_at set to the 2026-08-26 announcement date; replaced_by_model_id set to gpt-transcribe (same file-upload /v1/audio/transcriptions shape as this row)." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-4o-mini-transcribe", "display_name": "GPT-4o mini Transcribe", "price_per_minute_usd": "0.003", "languages": "99+", "streaming": false, "realtime": false, "diarization": "unsupported", "deprecated_at": "2026-08-26", "replaced_by_model_id": "gpt-transcribe", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-13", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/pricing", "notes": "Cost-efficient sibling of gpt-4o-transcribe with the same prompting and (non-realtime) streaming-of-output capabilities; same /v1/audio/transcriptions endpoint and 25 MB file-size limit. Billed per-token ($1.25/1M input audio tokens, $5.00/1M output tokens); developers.openai.com/api/docs/pricing publishes this as an 'Estimated cost' of $0.003/minute (confidence: medium, same token-vs-flat-rate caveat as gpt-4o-transcribe). 2026-08-11 re-verify: base model_id `gpt-4o-mini-transcribe` still active, price unchanged; only a snapshot-rotation deprecation was on file at that point (`gpt-4o-mini-transcribe-2025-03-20` -> `gpt-4o-mini-transcribe-2025-12-15`, shutdown 2027-01-20), which did not touch this row's base-alias model_id. 2026-09-02 re-verify: price still $0.003/min estimated ($1.25/$5.00 per 1M tokens), unchanged. /api/docs/deprecations now carries a base-model entry too (announced 2026-08-26, distinct from the earlier snapshot-rotation entry): removal from the API on 2027-02-26, vendor-recommended replacement gpt-live-transcribe or gpt-transcribe. deprecated_at set to the 2026-08-26 announcement date; replaced_by_model_id set to gpt-transcribe (same file-upload /v1/audio/transcriptions shape as this row)." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-4o-transcribe-diarize", "display_name": "GPT-4o Transcribe Diarize", "price_per_minute_usd": "0.006", "languages": "99+", "streaming": false, "realtime": false, "diarization": "included", "deprecated_at": "2026-08-26", "replaced_by_model_id": "gpt-transcribe", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-13", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/pricing", "notes": "Speaker-aware transcription variant for meeting recordings / multi-speaker audio: emits `transcript.text.segment` events with speaker labels (up to 4 known-speaker reference clips accepted), no separate diarization fee — priced identically to gpt-4o-transcribe at $2.50/$10.00 per 1M audio tokens, 'Estimated cost' $0.006/minute (confidence: medium, same token-vs-flat-rate caveat). Hidden/collapsed on the main pricing table. Requires `chunking_strategy` (auto or VAD config) for audio longer than 30 seconds; does not support the `prompt` or `logprobs` parameters. 25 MB file-size limit. 2026-08-11 re-verify: still active, price unchanged, not on /api/docs/deprecations at that time. 2026-09-02 re-verify: price still $0.006/min estimated, unchanged, on developers.openai.com/api/docs/pricing. Now listed on /api/docs/deprecations 'Transcription models' entry (announced 2026-08-26): removal from the API on 2027-02-26, vendor-recommended replacement gpt-live-transcribe or gpt-transcribe. deprecated_at set to the 2026-08-26 announcement date; replaced_by_model_id set to gpt-transcribe (no diarize successor announced — replacement lineup drops the diarization feature; noting this as a functional regression for any consumer relying on diarization: included)." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-realtime-whisper", "display_name": "GPT Realtime Whisper", "price_per_minute_usd": "0.017", "languages": "99+", "streaming": true, "realtime": true, "diarization": "unsupported", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/pricing", "notes": "New row this refresh (deferred 2026-07-13 pending STT-vs-realtime categorization; vendor now lists it plainly as 'Streaming speech-to-text model for realtime transcription', so it belongs here). Low-latency streaming transcription for live audio (microphone/call/media stream) via /v1/realtime and /v1/realtime/transcription_sessions. Flat $0.017 per minute of audio duration — not token-billed; price confirmed on both developers.openai.com/api/docs/pricing and the model detail page developers.openai.com/api/docs/models/gpt-realtime-whisper. confidence: medium because the supported-language list is not explicitly published for this model — '99+' is inherited from the Whisper family (speech-to-text guide's language section covers the file-based endpoints only). No function calling / structured outputs / fine-tuning. 2026-08-11 re-verify: still active, price unchanged ($0.017/min on both sources); note a new `gpt-live-transcribe` model launched this pass at the identical $0.017/min live-transcription rate (added as its own row — distinct model_id, not a rename of this one). 2026-09-02 re-verify: still active, price unchanged ($0.017/min). Explicitly named on /api/docs/deprecations as a recommended *replacement* target for the newly-deprecated whisper-1 / gpt-4o-transcribe family (not itself deprecated)." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-transcribe", "display_name": "GPT Transcribe", "price_per_minute_usd": "0.0045", "languages": "99+", "streaming": false, "realtime": false, "diarization": "unsupported", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-08-11", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/pricing", "notes": "New row this refresh: flat-rate file transcription model on the same /v1/audio/transcriptions endpoint as gpt-4o-transcribe, billed at a straight $0.0045/minute of audio duration rather than per-token (contrast gpt-4o-transcribe's token-billed 'Estimated cost' of $0.006/min). Price confirmed on both developers.openai.com/api/docs/pricing and the model detail page developers.openai.com/api/docs/models/gpt-transcribe ('Default snapshot: gpt-transcribe', active). Streaming/realtime modelled as false to match this dataset's convention for gpt-4o-transcribe: the model supports `stream=true` incremental text output on an already-uploaded file and can be used from `/v1/realtime/transcription_sessions`, but that is not continuous live audio-in the way `gpt-live-transcribe`/`gpt-realtime-whisper` are. confidence: medium because the supported-language list is not explicitly broken out for this model on either source — '99+' inherited from the rest of the OpenAI transcription family pending a dedicated languages page. 2026-09-02 re-verify: still active, price unchanged ($0.0045/min). Now the primary vendor-recommended replacement on /api/docs/deprecations for the newly-deprecated whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-transcribe-diarize rows (see their replaced_by_model_id)." }, { "provider": "OpenAI", "provider_url": "https://openai.com", "model_id": "gpt-live-transcribe", "display_name": "GPT Live Transcribe", "price_per_minute_usd": "0.017", "languages": "99+", "streaming": true, "realtime": true, "diarization": "unsupported", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-08-11", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://developers.openai.com/api/docs/pricing", "notes": "New row this refresh: low-latency streaming transcription model for live audio via /v1/realtime/transcription_sessions, flat $0.017/minute of realtime audio duration — not token-billed. Price confirmed on both developers.openai.com/api/docs/pricing and the model detail page developers.openai.com/api/docs/models/gpt-live-transcribe ('Default snapshot: gpt-live-transcribe', active). Priced identically to the existing gpt-realtime-whisper row at the same $0.017/min; kept as a separate row since it is a distinct model_id on the vendor's models catalog, not documented as a rename or successor of gpt-realtime-whisper (both remain live and undeprecated as of this pass). confidence: medium because the supported-language list is not explicitly published for this model — '99+' inherited from the rest of the OpenAI transcription family. 2026-09-02 re-verify: still active, price unchanged ($0.017/min). Explicitly named on /api/docs/deprecations as a recommended replacement target for the newly-deprecated whisper-1 / gpt-4o-transcribe family (not itself deprecated)." }, { "provider": "Groq", "provider_url": "https://groq.com", "model_id": "whisper-large-v3", "display_name": "Whisper V3 Large (Groq)", "price_per_minute_usd": "0.00185", "languages": "99+", "streaming": false, "realtime": false, "diarization": "unsupported", "min_billed_seconds": 10, "last_verified": "2026-09-02", "last_changed_at": "2026-05-05", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://console.groq.com/docs/models", "notes": "Whisper V3 Large hosted on Groq. Published as $0.111/hr. Speed factor 217x realtime. Re-verified 2026-09-02 against console.groq.com/docs/model/whisper-large-v3 and console.groq.com/docs/models (listed as a Production Model): model_id active, price ($0.111/hr = $0.00185/min) and 10s minimum billing unchanged. Not on console.groq.com/docs/deprecations. groq.com/pricing still 308-redirects to the marketing homepage (dead as a source_url) — use console.groq.com." }, { "provider": "Groq", "provider_url": "https://groq.com", "model_id": "whisper-large-v3-turbo", "display_name": "Whisper Large v3 Turbo (Groq)", "featured": true, "price_per_minute_usd": "0.000667", "languages": "99+", "streaming": false, "realtime": false, "diarization": "unsupported", "min_billed_seconds": 10, "last_verified": "2026-09-02", "last_changed_at": "2026-05-05", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://console.groq.com/docs/models", "notes": "Whisper Large v3 Turbo hosted on Groq — faster variant. Published as $0.04/hr. Speed factor 228x realtime. Re-verified 2026-09-02 against console.groq.com/docs/model/whisper-large-v3-turbo and console.groq.com/docs/models (listed as a Production Model): model_id active, price ($0.04/hr = $0.000667/min) and 10s minimum billing unchanged. console.groq.com/docs/deprecations confirms whisper-large-v3-turbo active and lists it as the replacement for the already-deprecated distil-whisper-large-v3-en (shutdown 2025-08-23, not present in this dataset). groq.com/pricing still 308-redirects to the marketing homepage (dead as a source_url) — use console.groq.com." }, { "provider": "Microsoft Azure", "provider_url": "https://azure.microsoft.com", "model_id": "azure-speech-realtime", "display_name": "Azure Speech (Real-time)", "price_per_minute_usd": "0.016667", "diarization_per_minute_usd": "0.005", "languages": "100+", "streaming": true, "realtime": true, "diarization": "extra-cost", "deployment_options": ["azure"], "last_verified": "2026-09-02", "last_changed_at": "2026-07-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://azure.microsoft.com/en-us/pricing/details/speech/", "notes": "Azure Speech Standard (S0) real-time speech-to-text. Published as $1.00/hr pay-as-you-go (= $0.016667/min), unchanged; custom real-time endpoint is $1.20/hr ($0.02/min). Commitment tiers reduce effective rate (2,000 hrs/mo $0.80/hr; 10,000 hrs/mo $0.65/hr; 50,000 hrs/mo $0.50/hr; 100,000 hrs/mo $0.40/hr — all confirmed unchanged via Retail Prices API). Enhanced add-on features on real-time (diarization, continuous language identification, pronunciation assessment) cost $0.30/hr per feature (= $0.005/min). Real-time diarization limited to 240 min/session. Languages: 100+ (139 locales per learn.microsoft.com language-support, up from 137 previously recorded). 2026-09-02 pass: azure.microsoft.com/en-us/pricing/details/speech/ fetched via WebFetch (per-hour figures still render client-side, page shows placeholder '$-' values), exact rates re-confirmed via Azure Retail Prices API (prices.azure.com, eastus USD: 'S1 Speech To Text' $1.00/hr, 'S1 Custom Speech To Text' $1.20/hr, 'S1 Speech to Text Enhanced Feature Audio' $0.30/hr, all commitment-tier meters unchanged including newly-checked 100K tier at $0.40/hr), cross-verified via learn.microsoft.com/en-us/azure/ai-services/speech-service/releasenotes (no STT pricing changes; 2026 changes are SDK feature releases — multichannel audio, source-language autodetection — plus retirement of ConversationTranslator/MeetingTranscriber/Intent Recognition, none of which affect this row) and language-support docs. No new base STT pricing tiers found in the Retail Prices API catalog for productName 'Azure Speech'. All figures unchanged from 2026-08-11 pass; confidence high." }, { "provider": "Microsoft Azure", "provider_url": "https://azure.microsoft.com", "model_id": "azure-speech-batch", "display_name": "Azure Speech (Batch)", "price_per_minute_usd": "0.003", "languages": "100+", "streaming": false, "realtime": false, "diarization": "included", "deployment_options": ["azure"], "last_verified": "2026-09-02", "last_changed_at": "2026-07-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://azure.microsoft.com/en-us/pricing/details/speech/", "notes": "Azure Speech Standard (S0) batch transcription. Published as $0.18/hr (= $0.003/min); custom batch endpoint is $0.225/hr ($0.00375/min). Batch enhanced add-ons (diarization, continuous language identification) are included at no extra charge; diarization up to 240 min/file. Fast transcription is modelled as its own row (azure-speech-fast). 2026-09-02 pass: azure.microsoft.com/en-us/pricing/details/speech/ fetched via WebFetch (figures still render client-side, page shows placeholder '$-' values), exact rates re-confirmed via Azure Retail Prices API (prices.azure.com, eastus USD: 'S1 Speech to Text Batch' $0.18/hr, 'S1 Custom Speech to Text Batch' $0.225/hr, unchanged), cross-verified via learn.microsoft.com/en-us/azure/ai-services/speech-service/releasenotes (no STT pricing changes or deprecations in 2026 releases; only SDK feature/retirement notes unrelated to batch pricing). No new batch STT tiers found in the Retail Prices API catalog. Figures unchanged from 2026-08-11 pass; confidence high." }, { "provider": "Microsoft Azure", "provider_url": "https://azure.microsoft.com", "model_id": "azure-speech-fast", "display_name": "Azure Speech (Fast)", "price_per_minute_usd": "0.006", "languages": "100+", "streaming": false, "realtime": false, "diarization": "included", "max_audio_minutes_per_file": 300, "deployment_options": ["azure"], "last_verified": "2026-09-02", "last_changed_at": "2026-07-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://azure.microsoft.com/en-us/pricing/details/speech/", "notes": "Azure Speech fast transcription: synchronous REST API (/speechtotext/transcriptions:transcribe) returning faster-than-real-time file transcription; not streaming. Published as $0.36/hr (= $0.006/min); custom fast transcription is $0.45/hr ($0.0075/min). Diarization included at no extra charge. Limits (S0): <500 MB and <5 hrs (300 min) per file, 600 requests/min. Supported across the 139-locale Azure STT locale set (fast transcription marked supported for the large majority; a handful of niche locales unsupported) per language-support docs. model_id is a Hail-coined tier slug — Azure exposes no stable model identifier. 2026-09-02 pass: rates re-confirmed via Azure Retail Prices API (prices.azure.com, eastus USD: 'Fast Transcription Speech To Text' $0.36/hr, 'Custom - Fast Transcription' $0.45/hr, unchanged), cross-verified via learn.microsoft.com/en-us/azure/ai-services/speech-service/releasenotes (no fast-transcription pricing or limit changes in 2026 releases) and language-support docs. No new fast-transcription tiers found in the Retail Prices API catalog. Figures unchanged from 2026-08-11 pass; confidence high." }, { "provider": "Google", "provider_url": "https://cloud.google.com", "model_id": "chirp_2", "display_name": "Chirp 2", "price_per_minute_usd": "0.016", "languages": "20+", "streaming": true, "realtime": true, "diarization": "included", "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-05-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://cloud.google.com/speech-to-text/pricing", "notes": "Google Cloud Speech-to-Text v2 multilingual model. Standard tier $0.016/min for both real-time and batch (down from v1's $0.024/min, per https://cloud.google.com/blog/products/ai-machine-learning/google-cloud-speech-to-text-v2-api). Dynamic Batch tier (up to 24h SLA) listed at $0.003/min on the pricing page as of 2026-07-19 (sku 7700-6778-EF8E; previously noted $0.004/min '75% off'; page footnote enumerates the Standard dynamic-batch models as default/command_and_search/latest_short/latest_long/phone_call/video/chirp without naming chirp_2) — not modelled as price_per_minute_batch_usd because standard batch is the same as real-time. Standard volume tiers: $0.016/min (0-500k min/mo), $0.01 (500k-1M), $0.008 (1M-2M), $0.004 (2M+). Supports StreamingRecognize (~20 languages), Recognize, and BatchRecognize (broadest language coverage) per https://docs.cloud.google.com/speech-to-text/docs/models/chirp-2. GA in us-central1, europe-west4, asia-southeast1. 2026-07-19 re-verify: WebFetch of the pricing page still truncates (known issue), but a raw-HTML fetch of the primary succeeded this pass — V2 standard recognition confirmed $0.016/min (sku 3099-B70F-0949); cross-checked via https://docs.cloud.google.com/speech-to-text/v2/docs/transcription-model (secondary; chirp_2 still listed GA, streaming, not deprecated) and web-search snippets citing the Google V2 launch blog. No structured price change. Language-count sub-field left untouched — prior passes found inconsistent counts (72/90+/150+) from the truncating supported-languages table, so it does not meet the two-source bar for a non-price-field edit. 2026-08-11 re-verify: WebFetch of the primary truncated again after one retry; a raw-HTML fetch of the primary succeeded and confirms Standard recognition tiers and the $0.003/min Dynamic Batch rate (sku 7700-6778-EF8E) unchanged, with no 'deprecated'/'retired'/'sunset' text anywhere on the page. Secondary (docs.cloud.google.com/speech-to-text/v2/docs/transcription-model) still lists chirp_2 as GA with streaming and batch support. No new Google STT model_ids found (checked for a Chirp 4 launch; none). No structured price change. 2026-09-02 re-verify: WebFetch of the primary truncated again after one retry; a raw-HTML fetch (decoded unicode-escaped JSON payload) confirms Standard Recognition tiers unchanged (sku 3099-B70F-0949: $0.016/$0.01/$0.008/$0.004 by volume) and Dynamic Batch Recognition unchanged (sku 7700-6778-EF8E: $0.003/min); zero hits for 'deprecat'/'retire'/'sunset'/'discontinu' anywhere on the page. Secondary (docs.cloud.google.com/speech-to-text/v2/docs/transcription-model) still lists chirp_2 as not deprecated, streaming+batch supported. Web search confirms no Chirp 4 launch; Chirp 3 remains Google's latest-generation model as of 2026-09. No structured price change." }, { "provider": "Google", "provider_url": "https://cloud.google.com", "model_id": "chirp_3", "display_name": "Chirp 3", "price_per_minute_usd": "0.016", "languages": "98+", "streaming": true, "realtime": true, "diarization": "included", "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-05-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://cloud.google.com/speech-to-text/pricing", "notes": "Google Cloud Speech-to-Text v2 latest-generation generative ASR model. Standard tier $0.016/min for both real-time and batch; Dynamic Batch tier (up to 24h SLA) listed at $0.003/min on the pricing page as of 2026-07-19 (sku 7700-6778-EF8E; previously noted $0.004/min '75% off'; page footnote enumerates the Standard dynamic-batch models without naming chirp_3) — not modelled as price_per_minute_batch_usd because standard batch is the same as real-time. Standard volume tiers: $0.016/min (0-500k min/mo), $0.01 (500k-1M), $0.008 (1M-2M), $0.004 (2M+). Adds automatic language detection and diarization vs Chirp 2 per https://docs.cloud.google.com/speech-to-text/docs/models/chirp-3. 98+ languages and locales (24 GA + 74 preview) per prior verification; supports StreamingRecognize and BatchRecognize. 2026-07-19 re-verify: WebFetch of the pricing page still truncates (known issue), but a raw-HTML fetch of the primary succeeded this pass — V2 standard recognition confirmed $0.016/min (sku 3099-B70F-0949); cross-checked via https://docs.cloud.google.com/speech-to-text/v2/docs/transcription-model (secondary; chirp_3 still listed GA as latest generation, streaming, not deprecated) and web-search snippets confirming $0.016/min standard for Chirp 3. No structured price change. Language-count sub-field left untouched — prior passes found inconsistent GA/preview splits and totals from truncating tables, and this pass's search snippets differ again (125+), so it does not meet the two-source bar for a non-price-field edit; existing 98+ figure retained pending a cleaner source. 2026-08-11 re-verify: WebFetch of the primary truncated again after one retry; a raw-HTML fetch of the primary succeeded and confirms Standard recognition tiers and the $0.003/min Dynamic Batch rate (sku 7700-6778-EF8E) unchanged, with no 'deprecated'/'retired'/'sunset' text anywhere on the page. Secondary (docs.cloud.google.com/speech-to-text/v2/docs/transcription-model) still lists chirp_3 as GA, latest generation, streaming, batch. No new Google STT model_ids found (checked for a Chirp 4 launch; none). No structured price change. 2026-09-02 re-verify: WebFetch of the primary truncated again after one retry; a raw-HTML fetch (decoded unicode-escaped JSON payload) confirms Standard Recognition tiers unchanged (sku 3099-B70F-0949: $0.016/$0.01/$0.008/$0.004 by volume) and Dynamic Batch Recognition unchanged (sku 7700-6778-EF8E: $0.003/min); zero hits for 'deprecat'/'retire'/'sunset'/'discontinu' anywhere on the page. Secondary (docs.cloud.google.com/speech-to-text/v2/docs/transcription-model) still lists chirp_3 as not deprecated, streaming+batch supported. Web search confirms no Chirp 4 launch; Chirp 3 remains Google's latest-generation model as of 2026-09. No structured price change." }, { "provider": "Speechmatics", "provider_url": "https://www.speechmatics.com", "model_id": "enhanced", "display_name": "Speechmatics Enhanced", "price_per_minute_usd": "0.00215", "languages": "56+", "streaming": true, "realtime": true, "diarization": "included", "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.speechmatics.com/pricing", "notes": "Speechmatics offers three operating points — `enhanced` (highest accuracy), `standard` (faster/cheaper), and the newer `melia-1` (multilingual, batch-only) — selected via the `operating_point` API parameter on both Batch and Real-time APIs (Melia 1 is batch-only). Pricing page lists Pro tier from $0.129/hr ($0.00215/min) on PAYG (price cut from $0.24/hr at last check) with an automatic 20% volume discount above 500 hrs/month; the same flat rate is used for both real-time and batch and is not broken out per operating point. Free plan now includes 3,000 minutes (50 hrs)/month, up from 480 min (not modelled as `free_tier` because schema expects per-day/per-token quotas). Confidence is medium because the pricing page exposes tier names rather than per-model SKU rates; verify against contract for production. Cross-verified product structure via https://docs.speechmatics.com/. 2026-07-19 re-verify: pricing page (primary) still shows Pro tier $0.129/hr ($0.00215/min) PAYG with the same 20% >500 hrs/month volume discount and 3,000 min (50 hrs)/month free tier; per-model rates still not broken out. Cross-checked https://docs.speechmatics.com/speech-to-text/models (secondary) — Enhanced still GA, batch+real-time, EU/US/AU; legacy `operating_point` parameter remains deprecated in favor of `model` (naming only, no effect on this row's model_id or price). Docs also mention a specialized Enhanced Medical variant with no separately published price. No price change. 2026-08-11 re-verify: pricing page (primary) still shows Pro tier $0.129/hr ($0.00215/min) PAYG with the same 20% >500 hrs/month volume discount (24,000 hrs/yr threshold for additional discounts); per-model rates still not broken out. Free plan has changed structure: it is now a one-time \"$100 in credit, no card required\" on signup rather than the previously-noted recurring 3,000 min (50 hrs)/month allotment — no recurring monthly free minutes are advertised on the page any longer (not modelled as `free_tier` field per prior note). Cross-checked https://docs.speechmatics.com/speech-to-text/models (secondary) — Enhanced still GA, batch+real-time, custom dictionaries/confidence scores/speaker ID/intelligence features intact. No price change. 2026-09-02 re-verify: pricing page (primary) still shows Pro tier $0.129/hr ($0.00215/min) PAYG, $100 signup credit, and the same 20% >500 hrs/month + 24,000 hrs/yr volume-discount structure; per-model rates still not broken out. Page now also lists an opt-in \"model training discount\" of 33% off STT rates for allowing Speechmatics to use submitted data for model improvement (off by default, reversible) — not modelled as a price field since it is conditional, not a published flat rate. Pricing page's general tier-features blurb currently rounds to \"55+ languages\" platform-wide (vs this row's \"56+\"); left unchanged since that figure is not published per-model and the secondary source gives no per-model count. Cross-checked https://docs.speechmatics.com/speech-to-text/models (secondary) — Enhanced still GA, batch+real-time, EU/US/AU, custom dictionaries/confidence scores/speaker ID intact; Medical domain variant still listed with no separate price. No price change; not deprecated." }, { "provider": "Speechmatics", "provider_url": "https://www.speechmatics.com", "model_id": "melia-1", "display_name": "Speechmatics Melia 1", "price_per_minute_usd": "0.00215", "languages": "68+", "streaming": false, "realtime": false, "diarization": "included", "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.speechmatics.com/pricing", "notes": "New `operating_point=melia-1` model added since last refresh: early-access, multilingual model that auto-detects and transcribes code-switched audio into one continuous transcript without language-pack selection, per https://docs.speechmatics.com/speech-to-text/models. Batch-only (no real-time API support) — Enhanced and Standard remain active and unaffected. Matches Standard's accuracy tier; lacks custom dictionaries, confidence scores, and speaker/intelligence features available on Enhanced/Standard, though diarization and word timings are supported. Available only in EU/US regions (not AU). `68+` languages is a derived count (the docs table lists 72 entries: ~68 individual languages plus 4 bilingual packs which Melia 1 does not use) — not an officially published per-model figure. Billed at the same flat Pro-tier rate as Enhanced/Standard per https://www.speechmatics.com/pricing — the pricing page does not publish a per-model rate, hence confidence: medium. 2026-07-19 re-verify: still early-access, batch-only, EU/US-only (accuracy on par with Standard; invoked via `model: melia-1` + `language: multi`) per https://docs.speechmatics.com/speech-to-text/models (secondary); flat Pro-tier rate ($0.129/hr / $0.00215/min) unchanged on the pricing page (primary). No price change; not deprecated. 2026-08-11 re-verify: still early-access, batch-only, EU/US-only per https://docs.speechmatics.com/speech-to-text/models (secondary) — no GA promotion, no region expansion. Flat Pro-tier rate ($0.129/hr / $0.00215/min) unchanged on the pricing page (primary); same rate applies across Enhanced/Standard/Melia 1 since Speechmatics still does not publish per-model SKU pricing. No price change; not deprecated. 2026-09-02 re-verify: still early-access, batch-only, EU/US-only per https://docs.speechmatics.com/speech-to-text/models (secondary) — no GA promotion, no region expansion, still lacks custom dictionaries/confidence scores/speaker ID. Flat Pro-tier rate ($0.129/hr / $0.00215/min) unchanged on the pricing page (primary), including the same 20% >500 hrs/month volume discount; pricing page's new opt-in 33% model-training discount applies platform-wide, not modelled as a separate field. No price change; not deprecated." }, { "provider": "Speechmatics", "provider_url": "https://www.speechmatics.com", "model_id": "standard", "display_name": "Speechmatics Standard", "price_per_minute_usd": "0.00215", "languages": "56+", "streaming": true, "realtime": true, "diarization": "included", "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-09-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.speechmatics.com/pricing", "notes": "Added 2026-09-02: `operating_point=standard` (now `model=standard`) is Speechmatics' third GA operating point — faster/cheaper than Enhanced, matched to Melia 1's accuracy tier per the Enhanced-vs-Standard-vs-Melia-1 comparison at https://docs.speechmatics.com/speech-to-text/models (secondary). This tier was described in the `enhanced` row's notes since the dataset's first Speechmatics pass but never modelled as its own row; added now for completeness since it is a distinct, independently-selectable, fully-documented operating point. GA (not early-access, unlike Melia 1), Batch + Realtime, available in EU/US/AU (same regional footprint as Enhanced) and carries the same custom dictionaries/confidence scores/speaker ID/diarization feature set as Enhanced. Billed at the same flat Pro-tier rate as Enhanced/Melia 1 per https://www.speechmatics.com/pricing (primary) — $0.129/hr ($0.00215/min) PAYG with the same 20% >500 hrs/month volume discount; Speechmatics does not publish a per-model SKU rate, hence confidence: medium (same reasoning as the other two rows). `56+` languages mirrors the Enhanced row's figure since both use the same language-pack model and no separate per-model count is published; not officially confirmed distinct from Enhanced's count. Not deprecated." }, { "provider": "Rev.ai", "provider_url": "https://www.rev.ai", "model_id": "whisper-fusion", "display_name": "Rev.ai Whisper Fusion", "price_per_minute_usd": "0.005", "languages": ["en"], "streaming": true, "realtime": true, "diarization": "included", "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-05-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.rev.ai/pricing", "notes": "Rev.ai's streaming transcription product, branded `Whisper Fusion` on the pricing page at $0.005/min; the parallel Whisper Large streaming tier is also $0.005/min. Free credits equivalent to 5 hours of Reverb ASR (cross-applicable across products). Reverb (batch) is modelled separately; see https://docs.rev.ai/ for the full API surface. English-primary; foreign language support is a distinct Reverb Foreign Language product line. 2026-07-13 re-verify: still listed at $0.005/min on the pricing page (primary); no price change. Note for context: the async job API's `transcriber` enum uses `fusion`/`machine`/`low_cost` rather than these marketing names (per docs.rev.ai, secondary) — model_id here follows this dataset's pricing-page-branding convention, unchanged from prior refresh. 2026-07-19 re-verify: still listed at $0.005/min on the pricing page (primary); `fusion` still an active `transcriber` value in the async job API reference (docs.rev.ai, secondary); no price change, not deprecated. 2026-08-11 re-verify: still listed at $0.005/min on the pricing page (primary); no price change. Secondary (docs.rev.ai/api/streaming/transcribers/) now documents only `machine` (account default) and `machine_v2` (always routes to Reverb ASR) — `fusion`/`low_cost` are no longer named there (a docs restructuring), but `Whisper Fusion` remains listed and priced on the pricing page itself; not deprecated. 2026-09-02 re-verify: still listed at $0.005/min on the pricing page (primary); no price change. Secondary (docs.rev.ai/api/streaming/transcribers/) still documents only `machine`/`machine_v2`, no `fusion` value; not deprecated — pricing page remains the authoritative listing for this product." }, { "provider": "Rev.ai", "provider_url": "https://www.rev.ai", "model_id": "reverb", "display_name": "Rev.ai Reverb", "price_per_minute_usd": "0.0033", "languages": ["en"], "streaming": false, "realtime": false, "diarization": "included", "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-05-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.rev.ai/pricing", "notes": "Rev.ai's async/batch ASR model branded `Reverb` at $0.20/hr ($0.0033/min). A `Reverb Turbo` tier exists at $0.10/hr ($0.0017/min) — not modelled as a separate row since it's a latency/quality dial on the same product; `Reverb Foreign Language` ($0.30/hr, $0.005/min, 56+ languages) is also priced separately and could be added if needed. Free credits equivalent to 5 hours of Reverb ASR. See https://docs.rev.ai/ for the async transcription API. 2026-07-13 re-verify: still listed at $0.20/hr on the pricing page (primary); no price change. Reverb Turbo and Reverb Foreign Language remain distinct pricing tiers, still not modelled as separate rows (unchanged decision). 2026-07-19 re-verify: still listed at $0.20/hr on the pricing page (primary); `machine` (routing to Reverb) still the default `transcriber` value in the async job API reference (docs.rev.ai, secondary); no price change, not deprecated. 2026-08-11 re-verify: still listed at $0.20/hr on the pricing page (primary); no price change; Reverb Foreign Language language count on the pricing page now reads 56+ (previously noted as 57+ — wording/count drift, not a pricing change). Secondary (docs.rev.ai/api/asynchronous/transcribers/) now documents `machine` as routing to the Reverb ASR model (`machine_v2` on streaming always routes to Reverb); no deprecation. 2026-09-02 re-verify: still listed at $0.20/hr on the pricing page (primary); no price change. Secondary (docs.rev.ai/api/asynchronous/transcribers/) documents `machine` as the default, routed to the Reverb ASR model; `human` transcriber also documented separately ($1.99/min, not modelled — human transcription, not ASR). Not deprecated." }, { "provider": "Gladia", "provider_url": "https://www.gladia.io", "model_id": "solaria-1", "display_name": "Gladia Solaria-1", "price_per_minute_usd": "0.0125", "price_per_minute_batch_usd": "0.01017", "languages": "100+", "streaming": true, "realtime": true, "realtime_latency_ms": 300, "diarization": "included", "deployment_options": ["native"], "last_verified": "2026-09-02", "last_changed_at": "2026-05-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.gladia.io/pricing", "notes": "Gladia's first-generation universal STT model `solaria-1`, supporting 100+ languages with automatic language detection and code-switching. Default model when `model` param is omitted (confirmed via docs.gladia.io pre-recorded transcription reference). Starter (PAYG) pricing: real-time $0.75/hr ($0.0125/min), async $0.61/hr (~$0.01017/min). Growth (committed) plan lowers real-time to $0.25/hr ($0.0042/min) and async to $0.20/hr ($0.0033/min). Sub-300ms streaming latency claimed on the pricing page. Speaker diarization and word-level timestamps included on all tiers. Model name confirmed via https://docs.gladia.io/. 2026-07-19 re-verify: still listed at the same rates on the pricing page (primary) and the pre-recorded API reference on docs.gladia.io (secondary, still the default model); no price change. 2026-08-11 re-verify: still listed at the same rates on the pricing page (primary); pre-recorded API reference on docs.gladia.io (secondary) still lists `solaria-1` as the default `model` value with 100+ languages and code-switching support; no price change. Free-tier wording corrected: the pricing page now describes a one-time 50€ signup credit (~80+ async hrs or ~60+ real-time hrs, no monthly reset), not the '10 free hours per month' previously noted here — likely a prior misread rather than a policy change, since no monthly-reset credit is documented on the current page. 2026-09-02 re-verify: still listed at the same rates on the pricing page (primary); docs.gladia.io models comparison page and pre-recorded API reference (secondary) still list `solaria-1` as the default model, 100+ languages, code-switching supported, async+live; no price change; no new Gladia STT models found." }, { "provider": "Gladia", "provider_url": "https://www.gladia.io", "model_id": "solaria-3", "display_name": "Gladia Solaria-3", "featured": true, "price_per_minute_usd": "0.01017", "languages": ["en", "fr", "de", "es", "it"], "streaming": false, "realtime": false, "diarization": "included", "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://www.gladia.io/pricing", "notes": "Gladia's higher-accuracy model for European real-world audio (English, French, German, Spanish, Italian); async/pre-recorded only, requires exactly one language in `language_config.languages` (no code-switching). Confirmed as a valid `model` param value via https://docs.gladia.io (pre-recorded transcription API reference). The pricing page does not break out per-model rates, so the Starter async rate ($0.61/hr ≈ $0.01017/min) is assumed to apply; `confidence: medium` reflects that this price is not explicitly stated per-model on the pricing page. No real-time/streaming support documented for this model. 2026-07-19 re-verify: still active and async-only per the pre-recorded API reference (secondary); pricing page (primary) still lists no per-model rate, so the Starter async assumption stands; no price change. 2026-08-11 re-verify: still active and async-only per the pre-recorded API reference (secondary, still \"English, French, German, Spanish, Italian\", still no code-switching); pricing page (primary) still has no per-model breakdown, so the Starter async assumption and `confidence: medium` stand; no price change. 2026-09-02 re-verify: still active and async-only per docs.gladia.io models comparison page and pre-recorded API reference (secondary, same 5 languages, no code-switching); pricing page (primary) still has no per-model breakdown, so the Starter async assumption and `confidence: medium` stand; no price change." }, { "provider": "Soniox", "provider_url": "https://soniox.com", "model_id": "stt-rt-v4", "display_name": "Soniox STT Real-time v4", "price_per_minute_usd": "0.002", "languages": "60+", "streaming": true, "realtime": true, "diarization": "included", "deployment_options": ["native"], "deprecated_at": "2026-06-30", "replaced_by_model_id": "stt-rt-v5", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-05-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://soniox.com/pricing", "notes": "Soniox real-time STT model `stt-rt-v4`. Primary billing metric is input audio tokens at $2.00 per 1M tokens; vendor approximates ~$0.12/hour which we use as $0.002/min for comparability. Aliased from `stt-rt-v3` (deprecated 2026-02-05; removed 2026-02-28 per https://soniox.com/docs/stt/models). Removed 2026-06-30 per https://soniox.com/docs/stt/models: requests using `stt-rt-v4` now auto-route to `stt-rt-v5` with no service interruption. 60+ languages with automatic language detection. Confidence medium because per-minute is an approximation of token-based pricing; actual cost varies with audio content density. 2026-07-19 re-verify: still deprecated/auto-routing to stt-rt-v5 per soniox.com/docs/stt/models (secondary); pricing page (primary) still quotes $2.00/1M input audio tokens for real-time; no price change. 2026-08-11 re-verify: still deprecated/auto-routing to stt-rt-v5 per soniox.com/docs/stt/models (secondary, no v6 listed); pricing page (primary) still quotes $2.00/1M input audio tokens for real-time; no price change. 2026-09-02 re-verify: still deprecated/auto-routing to stt-rt-v5 per soniox.com/docs/stt/models (secondary, no v6 listed); pricing page (primary) still quotes $2.00/1M input audio tokens for real-time; no price change." }, { "provider": "Soniox", "provider_url": "https://soniox.com", "model_id": "stt-rt-v5", "display_name": "Soniox STT Real-time v5", "price_per_minute_usd": "0.002", "languages": "60+", "streaming": true, "realtime": true, "diarization": "included", "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://soniox.com/pricing", "notes": "Soniox real-time STT model `stt-rt-v4` (deprecated 2026-06-30; auto-routes to `stt-rt-v5`) per https://soniox.com/docs/stt/models: `stt-rt-v5` released 2026-06-16 with reworked speaker separation and improved spoken-language identification. Soniox's pricing page bills real-time by input audio tokens at $2.00 per 1M tokens regardless of model version (~$0.12/hour, used here as $0.002/min for comparability); the docs page confirms pricing continuity across the v4→v5 auto-route (\"no service interruption\"). 60+ languages with automatic language detection. Confidence medium because (a) per-minute is an approximation of token-based pricing and (b) the pricing page does not itemize a v5-specific rate distinct from the general real-time rate. 2026-07-19 re-verify: still the current active real-time model per soniox.com/docs/stt/models (secondary), no newer version listed (no v6); pricing page (primary) still quotes $2.00/1M input audio tokens; no price change. 2026-08-11 re-verify: still the current active real-time model per soniox.com/docs/stt/models (secondary, still no v6, still \"up to 5 hours\" per request); pricing page (primary) still quotes $2.00/1M input audio tokens; no price change. 2026-09-02 re-verify: still the current active real-time model per soniox.com/docs/stt/models (secondary, still no v6, still \"up to 5 hours\" per request); pricing page (primary) still quotes $2.00/1M input audio tokens; no price change." }, { "provider": "Soniox", "provider_url": "https://soniox.com", "model_id": "stt-async-v4", "display_name": "Soniox STT Async v4", "price_per_minute_usd": "0.00167", "languages": "60+", "streaming": false, "realtime": false, "diarization": "included", "max_audio_minutes_per_file": 300, "deployment_options": ["native"], "deprecated_at": "2026-06-30", "replaced_by_model_id": "stt-async-v5", "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-05-19", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://soniox.com/pricing", "notes": "Soniox async (file) STT model `stt-async-v4`. Primary billing metric is input audio tokens at $1.50 per 1M tokens; vendor approximates ~$0.10/hour which we use as $0.00167/min for comparability. Aliased from `stt-async-v3` (deprecated 2026-02-05; removed 2026-02-28 per https://soniox.com/docs/stt/models). Removed 2026-06-30 per https://soniox.com/docs/stt/models: requests using `stt-async-v4` now auto-route to `stt-async-v5` with no service interruption. Supports up to 5 hours of audio per request. 60+ languages with automatic language detection. Confidence medium because per-minute is an approximation of token-based pricing. 2026-07-19 re-verify: still deprecated/auto-routing to stt-async-v5 per soniox.com/docs/stt/models (secondary); pricing page (primary) still quotes $1.50/1M input audio tokens for async; no price change. 2026-08-11 re-verify: still deprecated/auto-routing to stt-async-v5 per soniox.com/docs/stt/models (secondary, no v6 listed); pricing page (primary) still quotes $1.50/1M input audio tokens for async; no price change. 2026-09-02 re-verify: still deprecated/auto-routing to stt-async-v5 per soniox.com/docs/stt/models (secondary, no v6 listed); pricing page (primary) still quotes $1.50/1M input audio tokens for async; no price change." }, { "provider": "Soniox", "provider_url": "https://soniox.com", "model_id": "stt-async-v5", "display_name": "Soniox STT Async v5", "price_per_minute_usd": "0.00167", "languages": "60+", "streaming": false, "realtime": false, "diarization": "included", "max_audio_minutes_per_file": 300, "deployment_options": ["native"], "confidence": "medium", "last_verified": "2026-09-02", "last_changed_at": "2026-07-02", "verification_method": "manual-confirmed", "verified_by": "r13i", "source_url": "https://soniox.com/pricing", "notes": "Soniox async (file) STT model `stt-async-v4` (deprecated 2026-06-30; auto-routes to `stt-async-v5`) per https://soniox.com/docs/stt/models: `stt-async-v5` released 2026-06-11 with reengineered speaker separation and improved spoken-language identification. Soniox's pricing page bills async by input audio tokens at $1.50 per 1M tokens regardless of model version (~$0.10/hour, used here as $0.00167/min for comparability); the docs page confirms pricing continuity across the v4→v5 auto-route (\"no service interruption\"). Supports up to 5 hours of audio per request (carried over from v4; not independently re-stated for v5 in the docs excerpt). 60+ languages with automatic language detection. Confidence medium because (a) per-minute is an approximation of token-based pricing and (b) the max-audio-duration and pricing figures are inferred from vendor continuity statements rather than v5-specific line items. 2026-07-19 re-verify: still the current active async model per soniox.com/docs/stt/models (secondary), no newer version listed (no v6); pricing page (primary) still quotes $1.50/1M input audio tokens; no price change. 2026-08-11 re-verify: still the current active async model per soniox.com/docs/stt/models (secondary, still no v6, still \"up to 5 hours\" per request); pricing page (primary) still quotes $1.50/1M input audio tokens; no price change. 2026-09-02 re-verify: still the current active async model per soniox.com/docs/stt/models (secondary, still no v6, still \"up to 5 hours\" per request); pricing page (primary) still quotes $1.50/1M input audio tokens; no price change." } ] }