--- name: voice-generation description: Generate, play, and download realistic speech or compare voices across VoxBridge providers. Use when the user asks for TTS, narration, character dialogue, a spoken audio file, voice comparison, voice selection, or provider-specific voice controls. --- # VoxBridge voice generation VoxBridge is provider-neutral. Never assume a specific fictional universe, character set, business, or content workflow. When the VoxBridge MCP gateway is connected: 1. Call `list_providers` when provider availability or capabilities are unknown. 2. Call `list_voices` before generation when the user has not provided an exact voice ID. 3. For one voice, call `generate_speech` using the user's requested provider, `voice_id`, language, model, `output_format`, speed, performance direction, and `delivery` choice. 4. For an ordered script with multiple voice IDs in one file, select any configured speech provider whose `supports_dialogue` value is true, then call `generate_dialogue` once. This is not limited to ElevenLabs. Put each turn in `segments` with its text and exact `voice_id`, plus optional `instructions`, `speed`, `language`, provider-specific `options`, and `pause_after_ms`. Preserve the requested order, use one provider and one global model, and explain that the combined developer-alpha output is PCM WAV. Do not emulate this by making unrelated tool calls and asking another tool to merge files. 5. Map delivery intent explicitly: use `playback` for “recite,” “read aloud,” “play,” or “preview”; use `file` for “download,” “save,” or “give me the audio file”; use `both` when the user requests both. When no delivery preference is stated, use the gateway default, `both`. 6. Map natural-language performance direction into `instructions` only when `list_providers` reports support. Hume currently requires `model="octave-1"` for acting direction; OpenAI excludes `tts-1` and `tts-1-hd`. Respect each provider's `instructions_note`. 7. Use `options_json` for documented provider-specific controls on `generate_speech`; use the per-segment `options` object on `generate_dialogue`. Do not invent parameters. 8. Respect the selected provider's reported `speed_range`, formats, `file_delivery_formats`, character limit, allowed options, and `control_notes`. Raw headerless PCM is unavailable in this developer alpha because it lacks portable sample metadata; use WAV instead, explaining the container change if the user explicitly asked for PCM. 9. Surface only the representation(s) the user selected. `playback` returns inline audio, `file` returns the named binary audio-file resource with its Download control, and `both` gives the MCP App model-hidden playback data while retaining the same file resource for download. On hosts without MCP Apps, explain that the readable file remains available but the in-widget player is unavailable. Identify every result as synthetic AI-generated audio and preserve the requested output format. 10. If a provider is not configured, state that clearly and offer another connected provider rather than silently switching. 11. If the user asks to compare providers, keep the source text and requested performance direction as consistent as each provider permits. 12. Do not regenerate speech merely to change presentation: when `both` is selected, the player and file reference resolve to the same generated bytes. 13. For `file` or `both`, use the gateway's MCP App Download control when the host supports `ui/download-file`. Local download does not create a conversation attachment. The separate **Add file to ChatGPT** action is offered only by a compatible host. It calls the app-visible `materialize_audio_file` tool with the exact existing filename, verifies the returned bytes against the original metadata, and passes that same embedded resource through `ui/update-model-context`. Only the context update creates the removable composer attachment; the user sends the next message to share it with the model. The flow never regenerates speech and may trigger host approval. Do not claim reusable-library persistence, a reusable file ID, or support in every editing tool. 14. When ChatGPT, the user, or a compatible downstream tool needs to inspect, edit, measure, or otherwise use the generated audio bytes, call `materialize_audio_file` with the exact `materialize_resource_uri` included in the generation result. If the host hides that custom URI, pass the exact generated `file_name` instead. Supply exactly one of those values; never infer or alter either one, and never regenerate speech merely to obtain the file. This bounded path returns the same bytes and integrity metadata in an embedded resource. A direct tool call exposes those bytes to the caller; an MCP App tool call exposes them only to the App until it performs a separate model-context update. Neither fact alone proves reusable-library persistence or that every downstream editing tool accepts audio. 15. `generate_dialogue` makes one potentially billable provider call per segment. It supports at most 10 segments, 5,000 total characters, 30 seconds of inserted pauses, 600 seconds of output, and a 20 MiB final WAV by default. A later failure can occur after earlier calls; relay the gateway's partial-billing warning and never claim a partial combined file exists. The gateway suppresses a short-window identical host replay (for example, an approval retry) so it can return the existing result instead of rebilling the same request; this is not permission to issue a second tool call deliberately. 16. Combined dialogue currently requires one provider, one global model, and compatible PCM WAV responses. Do not claim MP3 composition, cross-provider mixing, automatic resampling, per-segment models, voice cloning, or speech-to-speech is available. Initial providers: ElevenLabs, Hume AI, Cartesia, Resemble AI, OpenAI, Deepgram, Google Cloud Text-to-Speech, and Microsoft Azure Speech.