openapi: 3.2.0 info: title: Sakura Internet Tts API version: 1.0.1 description: 'Operations tagged Tts across 2 of this provider''s published API definitions: ai-engine-inference-api.yaml, sakura-internet-ai-engine-inference-openapi.yml. Each path carries the servers of the definition it was published in.' servers: - url: https://api.ai.sakura.ad.jp security: - BearerAuth: [] tags: - name: TTS paths: /tts/v1/audio_query: post: summary: Create audio query (TTS) operationId: createTtsAudioQuery description: '音声合成用のクエリ(JSON)を作成します。 典型的には、/tts/v1/audio_queryでクエリを作成し、/tts/v1/synthesisに渡して音声(wav)を生成します。 このAPIはVOICEVOX Engine APIの/audio_query仕様を参考にした互換インターフェースを提供します。 公式仕様: https://voicevox.github.io/voicevox_engine/api/' parameters: - name: text in: query required: true schema: type: string minLength: 1 maxLength: 1000 description: 音声合成するテキスト - name: speaker in: query required: true schema: type: integer minimum: 0 description: 話者/スタイルID(利用可能な値はコントロールパネル等をご確認ください) - name: enable_katakana_english in: query required: false schema: type: boolean default: true description: 'カタカナ英語を有効にする。 ' - name: core_version in: query required: false schema: type: string description: '音声合成のバージョン指定。 ※現在は指定できますが無視されます。 ' responses: '200': description: Audio query created content: application/json: schema: $ref: '#/components/schemas/TtsAudioQuery' examples: exampleOk: summary: Successful audio_query response value: accent_phrases: [] speedScale: 1.0 pitchScale: 0.0 intonationScale: 1.0 volumeScale: 1.0 prePhonemeLength: 0.1 postPhonemeLength: 0.1 outputSamplingRate: 24000 outputStereo: false kana: '' '400': description: Bad request '401': description: Unauthorized '429': description: Rate limited '500': description: Server error '504': description: Server error tags: - TTS servers: - url: https://api.ai.sakura.ad.jp /tts/v1/synthesis: post: summary: Synthesize speech from audio query (TTS) operationId: synthesizeTtsSpeech description: '音声合成を行います。 /tts/v1/audio_query で作成したクエリ(JSON)をリクエストボディに渡して、音声(wav)を生成します。 このAPIはVOICEVOX Engine APIの/synthesis仕様を参考にした互換インターフェースを提供します。 公式仕様: https://voicevox.github.io/voicevox_engine/api/' parameters: - name: speaker in: query required: true schema: type: integer minimum: 0 description: 話者/スタイルID(利用可能な値はコントロールパネル等をご確認ください) - name: enable_interrogative_upspeak in: query required: false schema: type: boolean default: true description: '疑問系のテキストが与えられたら語尾を自動調整する ' - name: core_version in: query required: false schema: type: string description: 'Core Version。 ※現在は指定できますが無視されます。 ' requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/TtsSynthesisRequest' examples: exampleOk: summary: Example synthesis request value: accent_phrases: [] speedScale: 1 pitchScale: 0 intonationScale: 1 volumeScale: 1 prePhonemeLength: 0.1 postPhonemeLength: 0.1 pauseLength: null pauseLengthScale: 1 outputSamplingRate: 24000 outputStereo: false kana: string responses: '200': description: Speech audio generated (wav) headers: Content-Type: description: audio/wav schema: type: string example: audio/wav content: audio/wav: schema: type: string format: binary '400': description: Bad request '401': description: Unauthorized '429': description: Rate limited '500': description: Server error '504': description: Server error tags: - TTS servers: - url: https://api.ai.sakura.ad.jp components: schemas: TtsAccentPhrase: type: object required: - moras - accent - is_interrogative properties: moras: type: array description: モーラ列 items: type: object additionalProperties: true accent: type: integer description: アクセント位置 pause_mora: type: - object - 'null' description: 無音モーラ additionalProperties: true is_interrogative: type: boolean description: 疑問形かどうか TtsSynthesisRequest: type: object required: - accent_phrases - speedScale - pitchScale - intonationScale - volumeScale - prePhonemeLength - postPhonemeLength - outputSamplingRate - outputStereo - kana additionalProperties: true properties: accent_phrases: type: array description: アクセント句のリスト items: type: object description: Accent Phrase additionalProperties: true speedScale: type: number description: 全体の話速 pitchScale: type: number description: 全体の音高 intonationScale: type: number description: 全体の抑揚 volumeScale: type: number description: 全体の音量 prePhonemeLength: type: number description: 音声の前の無音時間 postPhonemeLength: type: number description: 音声の後の無音時間 pauseLength: description: '句読点などの無音時間。nullのときは無視される。デフォルト値はnull ' anyOf: - type: number - type: 'null' pauseLengthScale: type: number description: 句読点などの無音時間(倍率)。デフォルト値は1 default: 1 outputSamplingRate: type: integer description: 音声データの出力サンプリングレート outputStereo: type: boolean description: 音声データをステレオ出力するか否か kana: type: string description: '読み(かな)。 [読み取り専用] AquesTalk風記法によるテキスト。音声合成用のクエリとしては無視される ' TtsAudioQuery: type: object required: - accent_phrases - speedScale - pitchScale - intonationScale - volumeScale - prePhonemeLength - postPhonemeLength - outputSamplingRate - outputStereo properties: accent_phrases: type: array description: アクセント句のリスト items: $ref: '#/components/schemas/TtsAccentPhrase' speedScale: type: number description: 全体の話速 pitchScale: type: number description: 全体の音高 intonationScale: type: number description: 全体の抑揚 volumeScale: type: number description: 全体の音量 prePhonemeLength: type: number description: 音声の前の無音時間 postPhonemeLength: type: number description: 音声の後の無音時間 pauseLength: description: '句読点などの無音時間。 null のときは無視される。デフォルト値は null ' anyOf: - type: number - type: 'null' pauseLengthScale: type: number description: 句読点などの無音時間(倍率) default: 1 outputSamplingRate: type: integer description: 音声データの出力サンプリングレート outputStereo: type: boolean description: 音声データをステレオ出力するか否か kana: type: string description: '読み(かな)。 [読み取り専用] AquesTalk風記法によるテキスト。 音声合成用のクエリとしては無視される ' securitySchemes: BearerAuth: type: http scheme: bearer x-refined-from: - ai-engine-inference-api.yaml - sakura-internet-ai-engine-inference-openapi.yml