openapi: 3.0.1 info: title: Waves TTS API description: | Model-agnostic text-to-speech API. The same routes serve every Lightning model — the model is selected via the `model` body parameter. Supersedes the model-named routes under `/waves/v1/lightning-v3.1/`. Those routes still work but are deprecated; new integrations should use this API and pick a model with the body parameter instead. version: 1.0.0 servers: - url: https://api.smallest.ai description: Waves API server x-fern-server-name: waves paths: /waves/v1/tts: post: tags: - Text to Speech operationId: synthesizeSpeech summary: Synthesize speech description: | Synthesize speech from text in a single request. Pass `text` + `voice_id`, get back binary audio. Pick the model with the `model` body parameter: default `lightning_v3.1`, or `lightning_v3.1_pro` for the Pro pool. Other request parameters are identical across models. **Language behaviour on `lightning_v3.1_pro`:** pass `language: en` for UK + American accented English, pass `language: hi` for Indian accented English + Hindi (code-switching), or omit `language` to default to `en + hi` (mixed Indian + Western English coverage). Pro supports 31 languages total (10 Indic, 8 Asian & Middle Eastern, 13 European including Dutch and Swedish). Pass the matching ISO 639-1 code (e.g. `ta`, `de`, `ja`) with a Pro voice from that language, or use `auto` to route across all supported languages with any English or Hindi voice. See the [Lightning v3.1 Pro model card](/models/model-cards/text-to-speech/lightning-v-3-1-pro#supported-languages) for the full list. On `lightning_v3.1` the model accepts 20 language codes (10 European + 10 Indic) plus `auto`; the trained voice catalog covers 12 of those directly. ## When to use this - **Use this** for short utterances you can render before playback (notifications, prompts, batch jobs, audio file generation). - **Use `/waves/v1/tts/live`** when you want playback to start before the full audio is ready (long passages, latency-sensitive apps). - **Use `/waves/v1/tts/live`** (WebSocket) when text arrives incrementally (LLM token streams, live captioning). ## Key features - 44 kHz natural, expressive synthesis - Model selectable per request via `model` body parameter - Cloned voice IDs (`voice_*`) work on `lightning_v3.1` — same param as catalog voices - 20 accepted language codes on `lightning_v3.1` (12 with trained voices, 8 additional routed via English/Hindi voices). On `lightning_v3.1_pro`: 31 languages with dedicated voices (10 Indic, 8 Asian & Middle Eastern, 13 European); `language: en` → UK + American accented English; `language: hi` → Indian accented English + Hindi; omit `language` → defaults to `en + hi`. Both models accept `language: auto` for cross-language routing. - Output formats: `pcm`, `mp3`, `wav`, `ulaw`, `alaw` - Sample rates: 8 kHz – 44.1 kHz - Speed: 0.5× – 2× - Per-call pronunciation dictionaries via `pronunciation_dicts` ## Examples **cURL — Lightning v3.1 (default)** ```bash curl -X POST "https://api.smallest.ai/waves/v1/tts" \ -H "Authorization: Bearer $SMALLEST_API_KEY" \ -H "Content-Type: application/json" \ -H "Accept: audio/wav" \ -d '{ "text": "Hello from Waves TTS.", "voice_id": "magnus", "sample_rate": 24000, "output_format": "wav" }' --output speech.wav ``` **cURL — Lightning v3.1 Pro (omit `language` → defaults to `en + hi`)** ```bash curl -X POST "https://api.smallest.ai/waves/v1/tts" \ -H "Authorization: Bearer $SMALLEST_API_KEY" \ -H "Content-Type: application/json" \ -H "Accept: audio/wav" \ -d '{ "text": "Hello from the Lightning v3.1 Pro pool.", "voice_id": "meher", "model": "lightning_v3.1_pro", "sample_rate": 24000, "output_format": "wav" }' --output speech.wav ``` **cURL — Lightning v3.1 Pro with explicit `language: en` (UK + American accented English)** ```bash curl -X POST "https://api.smallest.ai/waves/v1/tts" \ -H "Authorization: Bearer $SMALLEST_API_KEY" \ -H "Content-Type: application/json" \ -H "Accept: audio/wav" \ -d '{ "text": "Good morning, this is a Pro voice speaking.", "voice_id": "meher", "model": "lightning_v3.1_pro", "language": "en", "sample_rate": 24000, "output_format": "wav" }' --output speech.wav ``` **cURL — Lightning v3.1 Pro with explicit `language: hi` (Indian accented English + Hindi)** ```bash curl -X POST "https://api.smallest.ai/waves/v1/tts" \ -H "Authorization: Bearer $SMALLEST_API_KEY" \ -H "Content-Type: application/json" \ -H "Accept: audio/wav" \ -d '{ "text": "Namaste, this is an Indian-accented Pro voice.", "voice_id": "meher", "model": "lightning_v3.1_pro", "language": "hi", "sample_rate": 24000, "output_format": "wav" }' --output speech.wav ``` ## Common gotchas - **Set `Accept: audio/wav`.** Omitting it can return an empty or unplayable response. - **Pair voice IDs with the right model.** Voice catalogs differ between `lightning_v3.1` and `lightning_v3.1_pro`. The API does not reject mismatched pairings, but using a Pro-only `voice_id` with `model=lightning_v3.1` (or omitting `model`) can return wrong or hallucinated audio. Pair Pro voices with `model=lightning_v3.1_pro`; standard catalog voices with `model=lightning_v3.1` (the default). - **Cloned voices** (`voice_*` from `add_voice`) work with `lightning_v3.1` only; voice cloning is not available on `lightning_v3.1_pro`. - **44.1 kHz output** is supported but most playback environments are happy with 24 kHz — drop the sample rate if bandwidth matters. parameters: - name: Accept in: header required: true schema: type: string enum: - audio/wav default: audio/wav description: Must be `audio/wav` to receive binary audio. Required for proper playback. security: - bearerAuth: [] requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/TtsRequest' examples: pro-default-en-hi: summary: 'Lightning v3.1 Pro: default (omit language, uses en + hi)' description: | On `lightning_v3.1_pro`, omitting `language` defaults to `en + hi` (mixed Indian + Western English coverage, auto-detected from input text). value: text: Hello from Waves TTS. voice_id: kaitlyn model: lightning_v3.1_pro output_format: mp3 sample_rate: 44100 speed: 1.0 pro-explicit-en: summary: 'Lightning v3.1 Pro: language=en (UK + American accented English)' value: text: Hello from Waves TTS. voice_id: kaitlyn model: lightning_v3.1_pro language: en output_format: mp3 sample_rate: 44100 speed: 1.0 pro-explicit-hi: summary: 'Lightning v3.1 Pro: language=hi (Indian accented English + Hindi)' value: text: Hello from Waves TTS. voice_id: meher model: lightning_v3.1_pro language: hi output_format: mp3 sample_rate: 44100 speed: 1.0 standard: summary: 'Lightning v3.1 (standard): explicit language from 12-language catalog' value: text: Hello from Waves TTS. voice_id: jordan model: lightning_v3.1 language: en output_format: mp3 sample_rate: 44100 speed: 1.0 responses: '200': description: Synthesized speech retrieved successfully. headers: X-Session-Id: schema: type: string description: Internal session identifier (system-generated UUID). X-Request-Id: schema: type: string description: Internal request identifier (system-generated UUID). X-External-Session-Id: schema: type: string description: Echoed client-provided session_id (empty if not provided). X-External-Request-Id: schema: type: string description: Echoed client-provided request_id (empty if not provided). content: audio/wav: schema: type: string format: binary description: A PCM int16 WAV file at the specified sample rate. '400': description: Bad request. content: application/json: schema: $ref: '#/components/schemas/TtsError' example: error: InvalidRequest message: The 'text' field is required. '401': description: Unauthorized. content: application/json: schema: $ref: '#/components/schemas/TtsError' example: error: Unauthorized message: Bearer token is missing or invalid. '500': description: Server error occurred. content: application/json: schema: $ref: '#/components/schemas/TtsError' example: error: InternalServerError message: An unexpected error occurred. /waves/v1/tts/live: post: tags: - Text to Speech operationId: synthesizeSpeechSse summary: Stream speech (SSE) description: | Synthesize speech and stream the audio back over Server-Sent Events. Same body as `/waves/v1/tts` — the only difference is the response is a stream of base64-encoded PCM chunks instead of one binary blob. Pick the model with the `model` body parameter, same as the sync route. **The same URL serves the WebSocket endpoint.** `wss://api.smallest.ai/waves/v1/tts/live` accepts a WebSocket upgrade for streaming-text scenarios (LLM token streams, live captioning). The HTTP `POST` documented on this page returns SSE; use `wss://` to use the WebSocket protocol instead. See the [WebSocket reference](/models/api-reference/text-to-speech/stream-speech-web-socket). ## When to use this - **Use this** when you want playback to start before synthesis is complete — long passages, latency-sensitive UI, live narration. - **Use sync `/waves/v1/tts`** when total latency doesn't matter and you'd rather get one buffer. - **Use `/waves/v1/tts/live`** (WebSocket) when the *text* arrives incrementally (LLM token stream). SSE assumes you have the full text up front. ## How it works 1. POST your text + voice settings — same payload as `/waves/v1/tts`, plus optional `model`. 2. The response is `Content-Type: text/event-stream`. Each chunk frame is `event: audio\n` followed by `data: {"audio": ""}\n\n`. 3. Decode each chunk's `audio` field with base64 and feed the PCM bytes to your audio pipeline (browser `MediaSource`, ffmpeg pipe, raw PCM player, etc.). 4. A final `data: {"done": true}\n\n` frame marks end of stream. ## Examples **cURL** ```bash curl -N -X POST "https://api.smallest.ai/waves/v1/tts/live" \ -H "Authorization: Bearer $SMALLEST_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "text": "Streaming this paragraph chunk by chunk so playback can start sooner.", "voice_id": "magnus", "sample_rate": 24000, "output_format": "pcm" }' ``` ## Common gotchas - **Use a streaming-friendly client.** `curl -N`, Python `iter_lines`, or a `fetch` `ReadableStream` reader. Buffering clients will hide the latency win. - **Audio is base64 inside the event payload**, not the raw event bytes. Decode the `data.audio` field per event. - **`output_format=pcm`** gives the lowest overhead for streaming playback. `wav`/`mp3` work but add per-chunk framing bytes. security: - bearerAuth: [] requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/TtsRequest' responses: '200': description: Synthesized speech retrieved successfully. headers: X-Session-Id: schema: type: string description: Internal session identifier (system-generated UUID). X-Request-Id: schema: type: string description: Internal request identifier (system-generated UUID). X-External-Session-Id: schema: type: string description: Echoed client-provided session_id (empty if not provided). X-External-Request-Id: schema: type: string description: Echoed client-provided request_id (empty if not provided). content: text/event-stream: schema: type: string description: | SSE stream of `event: audio` frames carrying base64-encoded PCM chunks, terminated by a `data: {"done": true}` frame. example: | event: audio data: {"audio": ""} data: {"done": true} '400': description: Bad request. content: application/json: schema: $ref: '#/components/schemas/TtsError' example: error: InvalidRequest message: The 'text' field is required. '401': description: Unauthorized. content: application/json: schema: $ref: '#/components/schemas/TtsError' example: error: Unauthorized message: Bearer token is missing or invalid. '500': description: Server error occurred. content: application/json: schema: $ref: '#/components/schemas/TtsError' example: error: InternalServerError message: An unexpected error occurred. components: schemas: TtsRequest: type: object required: - text - voice_id properties: text: type: string description: The text to convert to speech. default: Hello from Waves TTS. voice_id: type: string description: The voice identifier to use for speech generation. See the model card for available voices per model. default: magnus model: type: string description: | TTS model to route the request to. Controls which model pool serves this synthesis. - `lightning_v3.1` (default) — standard Lightning v3.1. - `lightning_v3.1_pro` — Lightning v3.1 Pro pool. Improved audio quality and naturalness, with a curated voice catalog. See the [Lightning v3.1 Pro model card](/models/model-cards/text-to-speech/lightning-v-3-1-pro) for supported voice IDs. Same concurrency and latency profile across both. Other request parameters behave identically. enum: - lightning_v3.1 - lightning_v3.1_pro default: lightning_v3.1 sample_rate: type: integer description: The sample rate for the generated audio. enum: - 8000 - 16000 - 24000 - 44100 default: 44100 speed: type: number description: The speed of the generated speech. minimum: 0.5 maximum: 2 default: 1.0 language: type: string description: | Language code for synthesis. Influences pronunciation, number/date normalization, and phoneme selection. **Default on `lightning_v3.1_pro`:** when `language` is omitted, the Pro pool defaults to **`en + hi`** (mixed Indian + Western English coverage, auto-detected from the input text). Each voice has its own `tags.language` set in the voice catalog — query `GET /waves/v1/lightning-v3.1/get_voices`. Pass a language the voice was trained on; passing other codes is accepted by the API but produces English-pronounced output. **`auto` (recommended for cross-language use cases):** routes internally based on the input text. Any English or Hindi voice can be used across all supported languages when `auto` is set; the platform handles language-appropriate routing without needing a code per call. **On `lightning_v3.1`** — 20 supported languages: - 10 European: English, Spanish, French, German, Italian, Dutch, Swedish, Portuguese, Polish, Russian - 10 Indic: Hindi, Marathi, Gujarati, Punjabi, Bengali, Odia, Tamil, Telugu, Kannada, Malayalam **On `lightning_v3.1_pro`** — 31 supported languages (adds 11 over base): - 13 European: base 10 plus Greek, Finnish, Norwegian - 8 Asian & Middle Eastern: Chinese, Japanese, Korean, Indonesian, Malay, Vietnamese, Turkish, Arabic - 10 Indic: same as base - Pass `en` → UK + American accented English. - Pass `hi` → Indian accented English + Hindi (code-switching). - Omit `language` → defaults to `en + hi` (mixed Indian + Western English coverage, auto-detected from input text). enum: - auto - en - hi - mr - kn - ta - bn - gu - te - ml - pa - or - es - de - fr - it - nl - sv - pt - ru - el - fi - 'no' - pl - ar - zh - id - ja - ko - ms - tr - vi number_pronunciation_language: type: string description: | Optional. Sets the language used to read out numeric content — numbers, currency amounts, times, and the numeric parts of dates and years — independently of the synthesis voice. Ordinary words are not translated. - If you **omit `language`**, this value also becomes the synthesis language: model selection and voice routing follow it. - If you **set `language` explicitly**, `language` always wins for synthesis and `number_pronunciation_language` only changes how numeric content is normalized. It works both ways — read numbers in Hindi under an English voice, or in English under a Hindi voice (tuned for Indian, often mixed-script, use cases). - Omit this field to keep the existing behaviour — normalization follows `language`. Note: only numeric tokens are re-spoken; the words around them stay in the text language. On a cross-language request names may also render in the target script (e.g. "Smith" → "स्मिथ"), which is generally the desired reading for native-language voices. Accepts the same language codes as `language` (including `auto`, `nl`, `sv`). enum: - auto - en - hi - mr - kn - ta - bn - gu - te - ml - pa - or - es - de - fr - it - nl - sv - pt - ru - el - fi - 'no' - pl - ar - zh - id - ja - ko - ms - tr - vi output_format: type: string description: | Format of the returned audio. `pcm` is the lowest-latency option but requires a decoder to play; `mp3` and `wav` are directly playable in browsers and most media players. The server default is `pcm` when the field is omitted — the API playground uses `mp3` so the generated audio is directly playable. default: pcm example: mp3 enum: - mp3 - pcm - wav - ulaw - alaw pronunciation_dicts: type: array items: type: string description: The ID of the pronunciation dictionary to use for speech generation. description: The IDs of the pronunciation dictionaries to use for speech generation. Available on both `lightning_v3.1` and `lightning_v3.1_pro`. word_timestamps: type: boolean default: false description: | **WebSocket-only feature.** Accepted on this endpoint but ignored — no per-word timing information is returned in the sync HTTP or SSE response shape. To receive `status: "word_timestamp"` frames with per-word `{ id, word, start, end }` data, use the WebSocket endpoint `wss://api.smallest.ai/waves/v1/tts/live`. See [Word-level timestamps](/models/documentation/text-to-speech-lightning/word-timestamps). session_id: type: string description: Optional client-provided session identifier for correlation. Only alphanumeric characters, hyphens, underscores, and dots are allowed. Max 128 characters. Echoed back in response headers as `X-External-Session-Id`. maxLength: 128 pattern: ^[a-zA-Z0-9_\-.]+$ request_id: type: string description: Optional client-provided request identifier for correlation. Only alphanumeric characters, hyphens, underscores, and dots are allowed. Max 128 characters. Echoed back in response headers as `X-External-Request-Id`. maxLength: 128 pattern: ^[a-zA-Z0-9_\-.]+$ TtsError: type: object properties: error: type: string description: Error type. message: type: string description: Error message. securitySchemes: bearerAuth: type: http scheme: bearer bearerFormat: JWT