openapi: 3.2.0 info: title: ElevenLabs API Documentation Text To Voice API description: This is the documentation for the ElevenLabs API. You can use this API to use our service programmatically, this is done by using your API key. You can find your API key in the dashboard at https://elevenlabs.io/app/settings/api-keys. version: '1.0' tags: - name: text-to-voice description: Design and generate custom voices from a text prompt. paths: /v1/text-to-voice/create-previews: post: tags: - text-to-voice summary: '[Deprecated] Generate A Voice Preview From Description' description: '**Deprecated.** Use `POST /v1/text-to-voice/design` instead. Generate a custom voice based on voice description. This method returns a list of voice previews. Each preview has a generated_voice_id and a sample of the voice as base64 encoded mp3 audio. To create the voice use `POST /v1/text-to-voice` with the chosen `generated_voice_id`.' operationId: text_to_voice deprecated: true parameters: - name: output_format in: query required: false schema: $ref: '#/components/schemas/AllowedOutputFormats' title: Output format of the generated audio. description: Output format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pro tier or above. Note that the μ-law format (sometimes written mu-law, often approximated as u-law) is commonly used for Twilio audio inputs. enum: - mp3_22050_32 - mp3_24000_48 - mp3_44100_32 - mp3_44100_64 - mp3_44100_96 - mp3_44100_128 - mp3_44100_192 - pcm_8000 - pcm_16000 - pcm_22050 - pcm_24000 - pcm_32000 - pcm_44100 - pcm_48000 - ulaw_8000 - alaw_8000 - opus_48000_32 - opus_48000_64 - opus_48000_96 - opus_48000_128 - opus_48000_192 default: mp3_44100_192 description: Output format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pro tier or above. Note that the μ-law format (sometimes written mu-law, often approximated as u-law) is commonly used for Twilio audio inputs. - name: xi-api-key in: header required: false schema: anyOf: - type: string - type: 'null' description: Your API key. This is required by most endpoints to access our API programmatically. You can view your xi-api-key using the 'Profile' tab on the website. title: Xi-Api-Key description: Your API key. This is required by most endpoints to access our API programmatically. You can view your xi-api-key using the 'Profile' tab on the website. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/VoicePreviewsRequestModel' responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/VoicePreviewsResponseModel' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' x-fern-sdk-group-name: text_to_voice x-fern-sdk-method-name: create_previews /v1/text-to-voice: post: tags: - text-to-voice summary: Create A New Voice From Voice Preview description: Create a voice from previously generated voice preview. This endpoint should be called after you fetched a generated_voice_id using POST /v1/text-to-voice/design or POST /v1/text-to-voice/:voice_id/remix. operationId: create_voice parameters: - name: xi-api-key in: header required: false schema: anyOf: - type: string - type: 'null' description: Your API key. This is required by most endpoints to access our API programmatically. You can view your xi-api-key using the 'Profile' tab on the website. title: Xi-Api-Key description: Your API key. This is required by most endpoints to access our API programmatically. You can view your xi-api-key using the 'Profile' tab on the website. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/Body_Create_a_new_voice_from_voice_preview_v1_text_to_voice_post' responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/VoiceResponseModel' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' x-fern-sdk-group-name: text_to_voice x-fern-sdk-method-name: create /v1/text-to-voice/design: post: tags: - text-to-voice summary: Design A Voice description: Design a voice via a prompt. This method returns a list of voice previews. Each preview has a generated_voice_id and a sample of the voice as base64 encoded mp3 audio. To create a voice use the generated_voice_id of the preferred preview with the /v1/text-to-voice endpoint. operationId: text_to_voice_design parameters: - name: output_format in: query required: false schema: $ref: '#/components/schemas/AllowedOutputFormats' title: Output format of the generated audio. description: Output format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pro tier or above. Note that the μ-law format (sometimes written mu-law, often approximated as u-law) is commonly used for Twilio audio inputs. enum: - mp3_22050_32 - mp3_24000_48 - mp3_44100_32 - mp3_44100_64 - mp3_44100_96 - mp3_44100_128 - mp3_44100_192 - pcm_8000 - pcm_16000 - pcm_22050 - pcm_24000 - pcm_32000 - pcm_44100 - pcm_48000 - ulaw_8000 - alaw_8000 - opus_48000_32 - opus_48000_64 - opus_48000_96 - opus_48000_128 - opus_48000_192 default: mp3_44100_192 description: Output format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pro tier or above. Note that the μ-law format (sometimes written mu-law, often approximated as u-law) is commonly used for Twilio audio inputs. - name: xi-api-key in: header required: false schema: anyOf: - type: string - type: 'null' description: Your API key. This is required by most endpoints to access our API programmatically. You can view your xi-api-key using the 'Profile' tab on the website. title: Xi-Api-Key description: Your API key. This is required by most endpoints to access our API programmatically. You can view your xi-api-key using the 'Profile' tab on the website. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/VoiceDesignRequestModel' responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/VoicePreviewsResponseModel' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' x-fern-sdk-group-name: text_to_voice x-fern-sdk-method-name: design /v1/text-to-voice/{voice_id}/remix: post: tags: - text-to-voice summary: Remix A Voice description: Remix an existing voice via a prompt. This method returns a list of voice previews. Each preview has a generated_voice_id and a sample of the voice as base64 encoded mp3 audio. To create a voice use the generated_voice_id of the preferred preview with the /v1/text-to-voice endpoint. operationId: text_to_voice_remix parameters: - name: voice_id in: path required: true schema: type: string description: Voice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices. examples: - 21m00Tcm4TlvDq8ikWAM title: Voice Id description: Voice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices. - name: output_format in: query required: false schema: $ref: '#/components/schemas/AllowedOutputFormats' title: Output format of the generated audio. description: Output format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pro tier or above. Note that the μ-law format (sometimes written mu-law, often approximated as u-law) is commonly used for Twilio audio inputs. enum: - mp3_22050_32 - mp3_24000_48 - mp3_44100_32 - mp3_44100_64 - mp3_44100_96 - mp3_44100_128 - mp3_44100_192 - pcm_8000 - pcm_16000 - pcm_22050 - pcm_24000 - pcm_32000 - pcm_44100 - pcm_48000 - ulaw_8000 - alaw_8000 - opus_48000_32 - opus_48000_64 - opus_48000_96 - opus_48000_128 - opus_48000_192 default: mp3_44100_192 description: Output format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pro tier or above. Note that the μ-law format (sometimes written mu-law, often approximated as u-law) is commonly used for Twilio audio inputs. - name: xi-api-key in: header required: false schema: anyOf: - type: string - type: 'null' description: Your API key. This is required by most endpoints to access our API programmatically. You can view your xi-api-key using the 'Profile' tab on the website. title: Xi-Api-Key description: Your API key. This is required by most endpoints to access our API programmatically. You can view your xi-api-key using the 'Profile' tab on the website. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/VoiceRemixRequestModel' responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/VoicePreviewsResponseModel' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' x-fern-sdk-group-name: text_to_voice x-fern-sdk-method-name: remix /v1/text-to-voice/{generated_voice_id}/stream: get: tags: - text-to-voice summary: Text To Voice Preview Streaming description: Stream a voice preview that was created via the /v1/text-to-voice/design endpoint. operationId: text_to_voice_preview_stream parameters: - name: generated_voice_id in: path required: true schema: type: string description: The generated_voice_id to stream. examples: - 37HceQefKmEi3bGovXjL embed: true title: Generated Voice Id description: The generated_voice_id to stream. - name: xi-api-key in: header required: false schema: anyOf: - type: string - type: 'null' description: Your API key. This is required by most endpoints to access our API programmatically. You can view your xi-api-key using the 'Profile' tab on the website. title: Xi-Api-Key description: Your API key. This is required by most endpoints to access our API programmatically. You can view your xi-api-key using the 'Profile' tab on the website. responses: '200': description: Streaming audio data content: audio/mpeg: schema: type: string format: binary '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' x-fern-sdk-group-name: - text_to_voice - preview x-fern-sdk-method-name: stream x-fern-streaming: true components: schemas: Body_Create_a_new_voice_from_voice_preview_v1_text_to_voice_post: properties: voice_name: type: string title: Voice Name description: Name to use for the created voice. examples: - Sassy squeaky mouse voice_description: type: string maxLength: 1000 minLength: 20 title: Voice Description description: Description to use for the created voice. examples: - A sassy squeaky mouse generated_voice_id: type: string title: Generated Voice Id description: The generated_voice_id to create; obtain it from POST /v1/text-to-voice/design, POST /v1/text-to-voice/:voice_id/remix, or the response headers when generating previews. examples: - 37HceQefKmEi3bGovXjL labels: anyOf: - additionalProperties: type: string type: object - type: 'null' title: Labels description: Optional, metadata to add to the created voice. Defaults to None. examples: - language: en name: Voice metadata played_not_selected_voice_ids: anyOf: - items: type: string type: array - type: 'null' title: Played Not Selected Voice Ids description: List of voice ids that the user has played but not selected. Used for RLHF. type: object required: - voice_name - voice_description - generated_voice_id title: Body_Create_a_new_voice_from_voice_preview_v1_text_to_voice_post VoicePreviewResponseModel: properties: audio_base_64: type: string title: Audio Base 64 description: The base64 encoded audio of the preview. generated_voice_id: type: string title: Generated Voice Id description: The ID of the generated voice. Use it to create a voice from the preview. media_type: type: string title: Media Type description: The media type of the preview. duration_secs: type: number title: Duration Secs description: The duration of the preview in seconds. language: anyOf: - type: string - type: 'null' title: Language description: The language of the preview. type: object required: - audio_base_64 - generated_voice_id - media_type - duration_secs - language title: VoicePreviewResponseModel SampleResponseModel: properties: sample_id: type: string title: Sample Id description: The ID of the sample. file_name: type: string title: File Name description: The name of the sample file. mime_type: type: string title: Mime Type description: The MIME type of the sample file. size_bytes: type: integer title: Size Bytes description: The size of the sample file in bytes. hash: type: string title: Hash description: The hash of the sample file. duration_secs: anyOf: - type: number - type: 'null' title: Duration Secs remove_background_noise: anyOf: - type: boolean - type: 'null' title: Remove Background Noise has_isolated_audio: anyOf: - type: boolean - type: 'null' title: Has Isolated Audio has_isolated_audio_preview: anyOf: - type: boolean - type: 'null' title: Has Isolated Audio Preview speaker_separation: anyOf: - $ref: '#/components/schemas/SpeakerSeparationResponseModel' - type: 'null' trim_start: anyOf: - type: integer - type: 'null' title: Trim Start trim_end: anyOf: - type: integer - type: 'null' title: Trim End type: object required: - sample_id - file_name - mime_type - size_bytes - hash title: SampleResponseModel example: file_name: sample.mp3 hash: '1234567890' mime_type: audio/mpeg sample_id: DCwhRBWXzGAHq8TQ4Fs18 size_bytes: 1000000 ManualVerificationResponseModel: properties: extra_text: type: string title: Extra Text description: The extra text of the manual verification. request_time_unix: type: integer title: Request Time Unix description: The date of the manual verification in Unix time. files: items: $ref: '#/components/schemas/ManualVerificationFileResponseModel' type: array title: Files description: The files of the manual verification. type: object required: - extra_text - request_time_unix - files title: ManualVerificationResponseModel example: extra_text: Please verify the voice is that of a female. files: - file_id: CwhRBWXzGAHq8TQ4Fs18 file_name: file.mp3 mime_type: audio/mpeg size_bytes: 1000000 upload_date_unix: 1714204800 request_time_unix: 1714204800 SpeakerSeparationResponseModel: properties: voice_id: type: string title: Voice Id description: The ID of the voice. sample_id: type: string title: Sample Id description: The ID of the sample. status: type: string enum: - not_started - pending - completed - failed title: Status description: The status of the speaker separation. speakers: anyOf: - additionalProperties: $ref: '#/components/schemas/SpeakerResponseModel' type: object - type: 'null' title: Speakers description: The speakers of the sample. selected_speaker_ids: anyOf: - items: type: string type: array - type: 'null' title: Selected Speaker Ids description: The IDs of the selected speakers. type: object required: - voice_id - sample_id - status title: SpeakerSeparationResponseModel example: sample_id: DCwhRBWXzGAHq8TQ4Fs18 status: not_started voice_id: DCwhRBWXzGAHq8TQ4Fs18 VoiceSharingModerationCheckResponseModel: properties: date_checked_unix: anyOf: - type: integer - type: 'null' title: Date Checked Unix description: The date the moderation check was made in Unix time. name_value: anyOf: - type: string - type: 'null' title: Name Value description: The name value of the voice. name_check: anyOf: - type: boolean - type: 'null' title: Name Check description: Whether the name check was successful. description_value: anyOf: - type: string - type: 'null' title: Description Value description: The description value of the voice. description_check: anyOf: - type: boolean - type: 'null' title: Description Check description: Whether the description check was successful. sample_ids: anyOf: - items: type: string type: array - type: 'null' title: Sample Ids description: A list of sample IDs. sample_checks: anyOf: - items: type: number type: array - type: 'null' title: Sample Checks description: A list of sample checks. captcha_ids: anyOf: - items: type: string type: array - type: 'null' title: Captcha Ids description: A list of captcha IDs. captcha_checks: anyOf: - items: type: number type: array - type: 'null' title: Captcha Checks description: A list of CAPTCHA check values. type: object title: VoiceSharingModerationCheckResponseModel example: captcha_checks: - 0.95 - 0.98 captcha_ids: - captcha1 - captcha2 date_checked_unix: 1714204800 description_check: true description_value: A female voice with a soft and friendly tone. name_check: true name_value: Rachel sample_checks: - 0.95 - 0.98 sample_ids: - sample1 - sample2 ValidationError: properties: loc: items: anyOf: - type: string - type: integer type: array title: Location msg: type: string title: Message type: type: string title: Error Type type: object required: - loc - msg - type title: ValidationError VoiceSettingsResponseModel: properties: stability: anyOf: - type: number maximum: 1.0 minimum: 0.0 - type: 'null' title: Stability description: Determines how stable the voice is and the randomness between each generation. Lower values introduce broader emotional range for the voice. Higher values can result in a monotonous voice with limited emotion. default: 0.5 use_speaker_boost: anyOf: - type: boolean - type: 'null' title: Use Speaker Boost description: This setting boosts the similarity to the original speaker. Using this setting requires a slightly higher computational load, which in turn increases latency. default: true similarity_boost: anyOf: - type: number maximum: 1.0 minimum: 0.0 - type: 'null' title: Similarity Boost description: Determines how closely the AI should adhere to the original voice when attempting to replicate it. default: 0.75 style: anyOf: - type: number - type: 'null' title: Style description: Determines the style exaggeration of the voice. This setting attempts to amplify the style of the original speaker. It does consume additional computational resources and might increase latency if set to anything other than 0. default: 0.0 speed: anyOf: - type: number - type: 'null' title: Speed description: Adjusts the speed of the voice. A value of 1.0 is the default speed, while values less than 1.0 slow down the speech, and values greater than 1.0 speed it up. default: 1.0 type: object title: VoiceSettingsResponseModel example: similarity_boost: 1.0 speed: 1.0 stability: 1.0 style: 0.0 use_speaker_boost: true VoiceVerificationResponseModel: properties: requires_verification: type: boolean title: Requires Verification description: Whether the voice requires verification. is_verified: type: boolean title: Is Verified description: Whether the voice has been verified. verification_failures: items: type: string type: array title: Verification Failures description: List of verification failures. verification_attempts_count: type: integer title: Verification Attempts Count description: The number of verification attempts. language: anyOf: - type: string - type: 'null' title: Language description: The language of the voice. verification_attempts: anyOf: - items: $ref: '#/components/schemas/VerificationAttemptResponseModel' type: array - type: 'null' title: Verification Attempts description: Number of times a verification was attempted. type: object required: - requires_verification - is_verified - verification_failures - verification_attempts_count title: VoiceVerificationResponseModel example: is_verified: true language: en requires_verification: false verification_attempts: - accepted: true date_unix: 1714204800 levenshtein_distance: 2 recording: mime_type: audio/mpeg recording_id: CwhRBWXzGAHq8TQ4Fs17 size_bytes: 1000000 transcription: Hello, how are you? upload_date_unix: 1714204800 similarity: 0.95 text: Hello, how are you? verification_attempts_count: 0 verification_failures: [] SpeakerResponseModel: properties: speaker_id: type: string title: Speaker Id description: The ID of the speaker. duration_secs: type: number title: Duration Secs description: The duration of the speaker segment in seconds. utterances: anyOf: - items: $ref: '#/components/schemas/UtteranceResponseModel' type: array - type: 'null' title: Utterances description: The utterances of the speaker. type: object required: - speaker_id - duration_secs title: SpeakerResponseModel example: duration_secs: 5.0 speaker_id: DCwhRBWXzGAHq8TQ4Fs18 VoiceResponseModel: properties: voice_id: type: string title: Voice Id description: The ID of the voice. name: type: string title: Name description: The name of the voice. samples: anyOf: - items: $ref: '#/components/schemas/SampleResponseModel' type: array - type: 'null' title: Samples description: List of samples associated with the voice. category: type: string enum: - generated - cloned - premade - professional - famous - high_quality title: Category description: The category of the voice. fine_tuning: anyOf: - $ref: '#/components/schemas/FineTuningResponseModel' - type: 'null' description: Fine-tuning information for the voice. labels: additionalProperties: type: string type: object title: Labels description: Labels associated with the voice. description: anyOf: - type: string - type: 'null' title: Description description: The description of the voice. preview_url: anyOf: - type: string - type: 'null' title: Preview Url description: The preview URL of the voice. available_for_tiers: items: type: string type: array title: Available For Tiers description: The tiers the voice is available for. settings: anyOf: - $ref: '#/components/schemas/VoiceSettingsResponseModel' - type: 'null' description: The settings of the voice. sharing: anyOf: - $ref: '#/components/schemas/VoiceSharingResponseModel' - type: 'null' description: The sharing information of the voice. high_quality_base_model_ids: items: type: string type: array title: High Quality Base Model Ids description: The base model IDs for high-quality voices. verified_languages: anyOf: - items: $ref: '#/components/schemas/VerifiedVoiceLanguageResponseModel' type: array - type: 'null' title: Verified Languages description: The verified languages of the voice. collection_ids: anyOf: - items: type: string type: array - type: 'null' title: Collection Ids description: The IDs of collections this voice belongs to. safety_control: anyOf: - type: string enum: - NONE - BAN - CAPTCHA - ENTERPRISE_BAN - ENTERPRISE_CAPTCHA - type: 'null' title: Safety Control description: The safety controls of the voice. voice_verification: anyOf: - $ref: '#/components/schemas/VoiceVerificationResponseModel' - type: 'null' description: The voice verification of the voice. permission_on_resource: anyOf: - type: string - type: 'null' title: Permission On Resource description: The permission on the resource of the voice. is_owner: anyOf: - type: boolean - type: 'null' title: Is Owner description: Whether the voice is owned by the user. is_legacy: type: boolean title: Is Legacy description: Whether the voice is legacy. default: false is_mixed: type: boolean title: Is Mixed description: Whether the voice is mixed. default: false favorited_at_unix: anyOf: - type: integer - type: 'null' title: Favorited At Unix description: Timestamp when the voice was marked as favorite in Unix time. created_at_unix: anyOf: - type: integer - type: 'null' title: Created At Unix description: The creation time of the voice in Unix time. is_bookmarked: anyOf: - type: boolean - type: 'null' title: Is Bookmarked description: Whether the voice is bookmarked by the current user. Only relevant for community (library-copied) voices. recording_quality: anyOf: - type: string enum: - studio - good - ok - poor - bad - type: 'null' title: Recording Quality description: The recording quality of the voice as determined by the review pipeline. labelling_status: anyOf: - type: string enum: - in_review - review_complete - type: 'null' title: Labelling Status description: The review pipeline status of the voice. recording_quality_reason: anyOf: - type: string - type: 'null' title: Recording Quality Reason description: The reason for the recording quality assessment, as determined by the review pipeline. type: object required: - voice_id - name - category - labels - available_for_tiers - high_quality_base_model_ids title: VoiceResponseModel example: available_for_tiers: - creator - enterprise category: professional description: A warm, expressive voice with a touch of humor. fine_tuning: is_allowed_to_fine_tune: true manual_verification_requested: false state: eleven_multilingual_v2: fine_tuned verification_attempts_count: 2 verification_failures: [] high_quality_base_model_ids: - eleven_v2_flash - eleven_flash_v2 - eleven_turbo_v2_5 - eleven_multilingual_v2 - eleven_v2_5_flash - eleven_flash_v2_5 - eleven_turbo_v2 is_legacy: false is_mixed: false is_owner: false labels: accent: American age: middle-aged description: expressive gender: female use_case: social media name: Rachel preview_url: https://storage.googleapis.com/eleven-public-prod/premade/voices/9BWtsMINqrJLrRacOk9x/405766b8-1f4e-4d3c-aba1-6f25333823ec.mp3 settings: similarity_boost: 1.0 speed: 1.0 stability: 1.0 style: 0.0 use_speaker_boost: true sharing: category: professional cloned_by_count: 50 date_unix: 1714204800 description: A female voice with a soft and friendly tone. disable_at_unix: 1714204800 enabled_in_library: true featured: true financial_rewards_enabled: true free_users_allowed: true history_item_sample_id: DCwhRBWXzGAHq8TQ4Fs18 labels: accent: American gender: female liked_by_count: 100 live_moderation_enabled: true moderation_check: captcha_checks: - 0.95 - 0.98 captcha_ids: - captcha1 - captcha2 date_checked_unix: 1714204800 description_check: true description_value: A female voice with a soft and friendly tone. name_check: true name_value: Rachel sample_checks: - 0.95 - 0.98 sample_ids: - sample1 - sample2 name: Rachel notice_period: 30 original_voice_id: DCwhRBWXzGAHq8TQ4Fs18 public_owner_id: DCwhRBWXzGAHq8TQ4Fs18 rate: 0.05 reader_app_enabled: true reader_restricted_on: - resource_id: FCwhRBWXzGAHq8TQ4Fs18 resource_type: read review_status: allowed status: enabled voice_mixing_allowed: false whitelisted_emails: - example@example.com verified_languages: - accent: american language: en locale: en-US model_id: eleven_multilingual_v2 preview_url: https://storage.googleapis.com/eleven-public-prod/premade/voices/9BWtsMINqrJLrRacOk9x/405766b8-1f4e-4d3c-aba1-6f25333823ec.mp3 voice_id: 21m00Tcm4TlvDq8ikWAM voice_verification: is_verified: true language: en requires_verification: false verification_attempts: - accepted: true date_unix: 1714204800 levenshtein_distance: 2 recording: mime_type: audio/mpeg recording_id: CwhRBWXzGAHq8TQ4Fs17 size_bytes: 1000000 transcription: Hello, how are you? upload_date_unix: 1714204800 similarity: 0.95 text: Hello, how are you? verification_attempts_count: 0 verification_failures: [] FineTuningResponseModel: properties: is_allowed_to_fine_tune: type: boolean title: Is Allowed To Fine Tune description: Whether the user is allowed to fine-tune the voice. state: additionalProperties: type: string enum: - not_started - queued - fine_tuning - fine_tuned - failed - delayed type: object title: State description: The state of the fine-tuning process for each model. verification_failures: items: type: string type: array title: Verification Failures description: List of verification failures in the fine-tuning process. verification_attempts_count: type: integer title: Verification Attempts Count description: The number of verification attempts in the fine-tuning process. manual_verification_requested: type: boolean title: Manual Verification Requested description: Whether a manual verification was requested for the fine-tuning process. language: anyOf: - type: string - type: 'null' title: Language description: The language of the fine-tuning process. progress: anyOf: - additionalProperties: type: number type: object - type: 'null' title: Progress description: The progress of the fine-tuning process. message: anyOf: - additionalProperties: type: string type: object - type: 'null' title: Message description: The message of the fine-tuning process. dataset_duration_seconds: anyOf: - type: number - type: 'null' title: Dataset Duration Seconds description: The duration of the dataset in seconds. verification_attempts: anyOf: - items: $ref: '#/components/schemas/VerificationAttemptResponseModel' type: array - type: 'null' title: Verification Attempts description: The number of verification attempts. slice_ids: anyOf: - items: type: string type: array - type: 'null' title: Slice Ids description: List of slice IDs. manual_verification: anyOf: - $ref: '#/components/schemas/ManualVerificationResponseModel' - type: 'null' description: The manual verification of the fine-tuning process. max_verification_attempts: anyOf: - type: integer - type: 'null' title: Max Verification Attempts description: The maximum number of verification attempts. next_max_verification_attempts_reset_unix_ms: anyOf: - type: integer - type: 'null' title: Next Max Verification Attempts Reset Unix Ms description: The next maximum verification attempts reset time in Unix milliseconds. type: object required: - is_allowed_to_fine_tune - state - verification_failures - verification_attempts_count - manual_verification_requested title: FineTuningResponseModel example: is_allowed_to_fine_tune: true manual_verification_requested: false state: eleven_multilingual_v2: fine_tuned verification_attempts_count: 2 verification_failures: [] VerificationAttemptResponseModel: properties: text: type: string title: Text description: The text of the verification attempt. date_unix: type: integer title: Date Unix description: The date of the verification attempt in Unix time. accepted: type: boolean title: Accepted description: Whether the verification attempt was accepted. similarity: type: number title: Similarity description: The similarity of the verification attempt. levenshtein_distance: type: number title: Levenshtein Distance description: The Levenshtein distance of the verification attempt. recording: anyOf: - $ref: '#/components/schemas/RecordingResponseModel' - type: 'null' description: The recording of the verification attempt. type: object required: - text - date_unix - accepted - similarity - levenshtein_distance title: VerificationAttemptResponseModel example: accepted: true date_unix: 1714204800 levenshtein_distance: 2 recording: mime_type: audio/mpeg recording_id: CwhRBWXzGAHq8TQ4Fs17 size_bytes: 1000000 transcription: Hello, how are you? upload_date_unix: 1714204800 similarity: 0.95 text: Hello, how are you? VoicePreviewsResponseModel: properties: previews: items: $ref: '#/components/schemas/VoicePreviewResponseModel' type: array title: Previews description: The previews of the generated voices. text: type: string title: Text description: The text used to preview the voices. type: object required: - previews - text title: VoicePreviewsResponseModel VoiceDesignRequestModel: properties: voice_description: type: string maxLength: 1000 minLength: 20 title: Voice Description description: Description to use for the created voice. examples: - A sassy squeaky mouse model_id: type: string enum: - eleven_multilingual_ttv_v2 - eleven_ttv_v3 title: Model Id description: 'Model to use for the voice generation. Possible values: eleven_multilingual_ttv_v2, eleven_ttv_v3.' default: eleven_multilingual_ttv_v2 examples: - eleven_multilingual_ttv_v2 text: anyOf: - type: string maxLength: 1000 minLength: 100 - type: 'null' title: Text description: Text to generate, text length has to be between 100 and 1000. examples: - Every act of kindness, no matter how small, carries value and can make a difference, as no gesture of goodwill is ever wasted. auto_generate_text: type: boolean title: Auto Generate Text description: Whether to automatically generate a text suitable for the voice description. default: false loudness: type: number maximum: 1.0 minimum: -1.0 title: Loudness description: Controls the volume level of the generated voice. -1 is quietest, 1 is loudest, 0 corresponds to roughly -24 LUFS. default: 0.5 examples: - 0.5 seed: anyOf: - type: integer maximum: 2147483647.0 minimum: 0.0 - type: 'null' title: Seed description: Random number that controls the voice generation. Same seed with same inputs produces same voice. examples: - 11 guidance_scale: type: number maximum: 100.0 minimum: 0.0 title: Guidance Scale description: Controls how closely the AI follows the prompt. Lower numbers give the AI more freedom to be creative, while higher numbers force it to stick more to the prompt. High numbers can cause voice to sound artificial or robotic. We recommend to use longer, more detailed prompts at lower Guidance Scale. default: 5 examples: - 5 stream_previews: type: boolean title: Stream Previews description: Determines whether the Text to Voice previews should be included in the response. If true, only the generated IDs will be returned which can then be streamed via the /v1/text-to-voice/:generated_voice_id/stream endpoint. default: false examples: - true should_enhance: type: boolean title: Should Enhance description: Whether to enhance the voice description using AI to add more detail and improve voice generation quality. When enabled, the system will automatically expand simple prompts into more detailed voice descriptions. Defaults to False default: false examples: - true remixing_session_id: anyOf: - type: string - type: 'null' title: Remixing Session Id description: The remixing session id. examples: - '123' remixing_session_iteration_id: anyOf: - type: string - type: 'null' title: Remixing Session Iteration Id description: The id of the remixing session iteration where these generations should be attached to. If not provided, a new iteration will be created. examples: - '123' quality: anyOf: - type: number maximum: 1.0 minimum: -1.0 - type: 'null' title: Quality description: Higher quality results in better voice output but less variety. examples: - 0.9 reference_audio_base64: anyOf: - type: string - type: 'null' title: Reference Audio Base64 description: Reference audio to use for the voice generation. The audio should be base64 encoded. Only supported when using the eleven_ttv_v3 model. prompt_strength: anyOf: - type: number maximum: 1.0 minimum: 0.0 - type: 'null' title: Prompt Strength description: Controls the balance of prompt versus reference audio when generating voice samples. 0 means almost no prompt influence, 1 means almost no reference audio influence. Only supported when using the eleven_ttv_v3 model. examples: - 0.25 type: object required: - voice_description title: VoiceDesignRequestModel ReaderResourceResponseModel: properties: resource_type: type: string enum: - read - collection title: Resource Type description: The type of resource. resource_id: type: string title: Resource Id description: The ID of the resource. type: object required: - resource_type - resource_id title: ReaderResourceResponseModel example: resource_id: FCwhRBWXzGAHq8TQ4Fs18 resource_type: read HTTPValidationError: properties: detail: items: $ref: '#/components/schemas/ValidationError' type: array title: Detail type: object title: HTTPValidationError RecordingResponseModel: properties: recording_id: type: string title: Recording Id description: The ID of the recording. mime_type: type: string title: Mime Type description: The MIME type of the recording. size_bytes: type: integer title: Size Bytes description: The size of the recording in bytes. upload_date_unix: type: integer title: Upload Date Unix description: The date of the recording in Unix time. transcription: type: string title: Transcription description: The transcription of the recording. type: object required: - recording_id - mime_type - size_bytes - upload_date_unix - transcription title: RecordingResponseModel example: mime_type: audio/mpeg recording_id: CwhRBWXzGAHq8TQ4Fs17 size_bytes: 1000000 transcription: Hello, how are you? upload_date_unix: 1714204800 UtteranceResponseModel: properties: start: type: number title: Start description: The start time of the utterance in seconds. end: type: number title: End description: The end time of the utterance in seconds. type: object required: - start - end title: UtteranceResponseModel example: end: 1.0 start: 0.0 VerifiedVoiceLanguageResponseModel: properties: language: type: string title: Language description: The language of the voice. model_id: type: string title: Model Id description: The voice's model ID. accent: anyOf: - type: string - type: 'null' title: Accent description: The voice's accent, if applicable. locale: anyOf: - type: string - type: 'null' title: Locale description: The voice's locale, if applicable. preview_url: anyOf: - type: string - type: 'null' title: Preview Url description: The voice's preview URL, if applicable. type: object required: - language - model_id title: VerifiedVoiceLanguageResponseModel example: accent: American language: en model_id: eleven_turbo_v2_5 AllowedOutputFormats: type: string enum: - mp3_22050_32 - mp3_24000_48 - mp3_44100_32 - mp3_44100_64 - mp3_44100_96 - mp3_44100_128 - mp3_44100_192 - pcm_8000 - pcm_16000 - pcm_22050 - pcm_24000 - pcm_32000 - pcm_44100 - pcm_48000 - ulaw_8000 - alaw_8000 - opus_48000_32 - opus_48000_64 - opus_48000_96 - opus_48000_128 - opus_48000_192 VoiceRemixRequestModel: properties: voice_description: type: string maxLength: 1000 minLength: 5 title: Voice Description description: Description of the changes to make to the voice. examples: - Make the voice have a higher pitch. text: anyOf: - type: string maxLength: 1000 minLength: 100 - type: 'null' title: Text description: Text to generate, text length has to be between 100 and 1000. examples: - Every act of kindness, no matter how small, carries value and can make a difference, as no gesture of goodwill is ever wasted. auto_generate_text: type: boolean title: Auto Generate Text description: Whether to automatically generate a text suitable for the voice description. default: false loudness: type: number maximum: 1.0 minimum: -1.0 title: Loudness description: Controls the volume level of the generated voice. -1 is quietest, 1 is loudest, 0 corresponds to roughly -24 LUFS. default: 0.5 examples: - 0.5 seed: anyOf: - type: integer maximum: 2147483647.0 minimum: 0.0 - type: 'null' title: Seed description: Random number that controls the voice generation. Same seed with same inputs produces same voice. examples: - 11 guidance_scale: type: number maximum: 100.0 minimum: 0.0 title: Guidance Scale description: Controls how closely the AI follows the prompt. Lower numbers give the AI more freedom to be creative, while higher numbers force it to stick more to the prompt. High numbers can cause voice to sound artificial or robotic. We recommend to use longer, more detailed prompts at lower Guidance Scale. default: 2 examples: - 5 stream_previews: type: boolean title: Stream Previews description: Determines whether the Text to Voice previews should be included in the response. If true, only the generated IDs will be returned which can then be streamed via the /v1/text-to-voice/:generated_voice_id/stream endpoint. default: false examples: - true remixing_session_id: anyOf: - type: string - type: 'null' title: Remixing Session Id description: The remixing session id. examples: - '123' remixing_session_iteration_id: anyOf: - type: string - type: 'null' title: Remixing Session Iteration Id description: The id of the remixing session iteration where these generations should be attached to. If not provided, a new iteration will be created. examples: - '123' prompt_strength: anyOf: - type: number maximum: 1.0 minimum: 0.0 - type: 'null' title: Prompt Strength description: Controls the balance of prompt versus reference audio when generating voice samples. 0 means almost no prompt influence, 1 means almost no reference audio influence. Only supported when using the eleven_ttv_v3 model. examples: - 0.25 type: object required: - voice_description title: VoiceRemixRequestModel VoiceSharingResponseModel: properties: status: type: string enum: - enabled - disabled - copied - copied_disabled title: Status description: The status of the voice sharing. history_item_sample_id: anyOf: - type: string - type: 'null' title: History Item Sample Id description: The sample ID of the history item. date_unix: type: integer title: Date Unix description: The date of the voice sharing in Unix time. whitelisted_emails: items: type: string type: array title: Whitelisted Emails description: A list of whitelisted emails. public_owner_id: type: string title: Public Owner Id description: The ID of the public owner. original_voice_id: type: string title: Original Voice Id description: The ID of the original voice. financial_rewards_enabled: type: boolean title: Financial Rewards Enabled description: Whether financial rewards are enabled. free_users_allowed: type: boolean title: Free Users Allowed description: Whether free users are allowed. live_moderation_enabled: type: boolean title: Live Moderation Enabled description: Whether live moderation is enabled. rate: anyOf: - type: number - type: 'null' title: Rate description: The rate of the voice sharing. fiat_rate: anyOf: - type: number - type: 'null' title: Fiat Rate description: The rate of the voice sharing in USD per 1000 credits. notice_period: type: integer title: Notice Period description: The notice period of the voice sharing. disable_at_unix: anyOf: - type: integer - type: 'null' title: Disable At Unix description: The date of the voice sharing in Unix time. voice_mixing_allowed: type: boolean title: Voice Mixing Allowed description: Whether voice mixing is allowed. featured: type: boolean title: Featured description: Whether the voice is featured. category: type: string enum: - generated - cloned - premade - professional - famous - high_quality title: Category description: The category of the voice. reader_app_enabled: anyOf: - type: boolean - type: 'null' title: Reader App Enabled description: Whether the reader app is enabled. image_url: anyOf: - type: string - type: 'null' title: Image Url description: The image URL of the voice. ban_reason: anyOf: - type: string - type: 'null' title: Ban Reason description: The ban reason of the voice. liked_by_count: type: integer title: Liked By Count description: The number of likes on the voice. cloned_by_count: type: integer title: Cloned By Count description: The number of clones on the voice. name: type: string title: Name description: The name of the voice. description: anyOf: - type: string - type: 'null' title: Description description: The description of the voice. labels: additionalProperties: type: string type: object title: Labels description: The labels of the voice. review_status: type: string enum: - not_requested - pending - declined - allowed - allowed_with_changes title: Review Status description: The review status of the voice. review_message: anyOf: - type: string - type: 'null' title: Review Message description: The review message of the voice. enabled_in_library: type: boolean title: Enabled In Library description: Whether the voice is enabled in the library. instagram_username: anyOf: - type: string - type: 'null' title: Instagram Username description: The Instagram username of the voice. twitter_username: anyOf: - type: string - type: 'null' title: Twitter Username description: The Twitter/X username of the voice. youtube_username: anyOf: - type: string - type: 'null' title: Youtube Username description: The YouTube username of the voice. tiktok_username: anyOf: - type: string - type: 'null' title: Tiktok Username description: The TikTok username of the voice. moderation_check: anyOf: - $ref: '#/components/schemas/VoiceSharingModerationCheckResponseModel' - type: 'null' description: The moderation check of the voice. reader_restricted_on: anyOf: - items: $ref: '#/components/schemas/ReaderResourceResponseModel' type: array - type: 'null' title: Reader Restricted On description: The reader restricted on of the voice. type: object required: - status - date_unix - whitelisted_emails - public_owner_id - original_voice_id - financial_rewards_enabled - free_users_allowed - live_moderation_enabled - notice_period - voice_mixing_allowed - featured - category - liked_by_count - cloned_by_count - name - labels - review_status - enabled_in_library title: VoiceSharingResponseModel example: category: professional cloned_by_count: 50 date_unix: 1714204800 description: A female voice with a soft and friendly tone. disable_at_unix: 1714204800 enabled_in_library: true featured: true financial_rewards_enabled: true free_users_allowed: true history_item_sample_id: DCwhRBWXzGAHq8TQ4Fs18 labels: accent: American gender: female liked_by_count: 100 live_moderation_enabled: true moderation_check: captcha_checks: - 0.95 - 0.98 captcha_ids: - captcha1 - captcha2 date_checked_unix: 1714204800 description_check: true description_value: A female voice with a soft and friendly tone. name_check: true name_value: Rachel sample_checks: - 0.95 - 0.98 sample_ids: - sample1 - sample2 name: Rachel notice_period: 30 original_voice_id: DCwhRBWXzGAHq8TQ4Fs18 public_owner_id: DCwhRBWXzGAHq8TQ4Fs18 rate: 0.05 reader_app_enabled: true reader_restricted_on: - resource_id: FCwhRBWXzGAHq8TQ4Fs18 resource_type: read review_status: allowed status: enabled voice_mixing_allowed: false whitelisted_emails: - example@example.com ManualVerificationFileResponseModel: properties: file_id: type: string title: File Id description: The ID of the file. file_name: type: string title: File Name description: The name of the file. mime_type: type: string title: Mime Type description: The MIME type of the file. size_bytes: type: integer title: Size Bytes description: The size of the file in bytes. upload_date_unix: type: integer title: Upload Date Unix description: The date of the file in Unix time. type: object required: - file_id - file_name - mime_type - size_bytes - upload_date_unix title: ManualVerificationFileResponseModel example: file_id: CwhRBWXzGAHq8TQ4Fs18 file_name: file.mp3 mime_type: audio/mpeg size_bytes: 1000000 upload_date_unix: 1714204800 VoicePreviewsRequestModel: properties: voice_description: type: string maxLength: 1000 minLength: 20 title: Voice Description description: Description to use for the created voice. examples: - A sassy squeaky mouse text: anyOf: - type: string maxLength: 1000 minLength: 100 - type: 'null' title: Text description: Text to generate, text length has to be between 100 and 1000. examples: - Every act of kindness, no matter how small, carries value and can make a difference, as no gesture of goodwill is ever wasted. auto_generate_text: type: boolean title: Auto Generate Text description: Whether to automatically generate a text suitable for the voice description. default: false loudness: type: number maximum: 1.0 minimum: -1.0 title: Loudness description: Controls the volume level of the generated voice. -1 is quietest, 1 is loudest, 0 corresponds to roughly -24 LUFS. default: 0.5 examples: - 0.5 quality: type: number maximum: 1.0 minimum: -1.0 title: Quality description: Higher quality results in better voice output but less variety. default: 0.9 examples: - 0.9 seed: anyOf: - type: integer maximum: 2147483647.0 minimum: 0.0 - type: 'null' title: Seed description: Random number that controls the voice generation. Same seed with same inputs produces same voice. examples: - 11 guidance_scale: type: number maximum: 100.0 minimum: 0.0 title: Guidance Scale description: Controls how closely the AI follows the prompt. Lower numbers give the AI more freedom to be creative, while higher numbers force it to stick more to the prompt. High numbers can cause voice to sound artificial or robotic. We recommend to use longer, more detailed prompts at lower Guidance Scale. default: 5 examples: - 5 should_enhance: type: boolean title: Should Enhance description: Whether to enhance the voice description using AI to add more detail and improve voice generation quality. When enabled, the system will automatically expand simple prompts into more detailed voice descriptions. Defaults to False default: false examples: - true type: object required: - voice_description title: VoicePreviewsRequestModel