openapi: 3.0.3 paths: /v2/upload: post: operationId: FileController_upload_v2 parameters: [] requestBody: required: true content: multipart/form-data: schema: type: object properties: audio: type: string format: binary description: The file to be uploaded. This should be an audio or video file. application/json: schema: type: object properties: audio_url: type: string description: The URL of the audio or video file to be uploaded. responses: '200': description: '' content: application/json: schema: $ref: '#/components/schemas/AudioUploadResponse' security: - x_gladia_key: [] summary: Upload an audio file or provide an audio URL for processing tags: - File Management /v2/pre-recorded: post: operationId: PreRecordedController_initPreRecordedJob_v2 parameters: [] requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/InitTranscriptionRequest' responses: '201': description: The pre recorded job has been initiated content: application/json: schema: $ref: '#/components/schemas/InitPreRecordedTranscriptionResponse' '400': description: Something is wrong with the request content: application/json: schema: $ref: '#/components/schemas/BadRequestErrorResponse' '401': description: You don't have the permissions to initiate a new pre recorded job content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' '422': description: The parameters you gave are incorrect content: application/json: schema: $ref: '#/components/schemas/UnprocessableEntityErrorResponse' security: - x_gladia_key: [] summary: Initiate a new pre recorded job tags: - Pre-recorded V2 get: operationId: PreRecordedController_getPreRecordedJobs_v2 parameters: - name: offset required: false in: query description: The starting point for pagination. A value of 0 starts from the first item. schema: minimum: 0 default: 0 type: integer - name: limit required: false in: query description: The maximum number of items to return. Useful for pagination and controlling data payload size. schema: minimum: 1 default: 20 type: integer - name: date required: false in: query description: Filter items relevant to a specific date in ISO format (YYYY-MM-DD). schema: format: date-time example: '2026-06-12' type: string - name: before_date required: false in: query description: Include items that occurred before the specified date in ISO format. schema: format: date-time example: '2026-06-12T21:00:09.947Z' type: string - name: after_date required: false in: query description: Filter for items after the specified date. Use with `before_date` for a range. Date in ISO format. schema: format: date-time example: '2026-06-12T21:00:09.947Z' type: string - name: status required: false in: query description: Filter the list based on item status. Accepts multiple values from the predefined list. schema: example: - done type: array items: type: string enum: - queued - processing - done - error - name: custom_metadata required: false in: query schema: additionalProperties: true example: user: John Doe type: object responses: '200': description: A list of pre recorded jobs matching the parameters. content: application/json: schema: $ref: '#/components/schemas/ListPreRecordedResponse' '401': description: You don't have the permissions to access pre recorded jobs content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' security: - x_gladia_key: [] summary: Get pre recorded jobs based on query parameters tags: - Pre-recorded V2 /v2/pre-recorded/{id}: get: operationId: PreRecordedController_getPreRecordedJob_v2 parameters: - name: id required: true in: path description: Id of the pre recorded job schema: example: 45463597-20b7-4af7-b3b3-f5fb778203ab type: string responses: '200': description: The pre recorded job's metadata content: application/json: schema: $ref: '#/components/schemas/PreRecordedResponse' '401': description: You don't have the permissions to access the pre recorded job content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' '404': description: The pre recorded job doesn't exist or has been deleted content: application/json: schema: $ref: '#/components/schemas/NotFoundErrorResponse' security: - x_gladia_key: [] summary: Get the pre recorded job's metadata tags: - Pre-recorded V2 delete: operationId: PreRecordedController_deletePreRecordedJob_v2 parameters: - name: id required: true in: path description: Id of the pre recorded job schema: example: 45463597-20b7-4af7-b3b3-f5fb778203ab type: string responses: '202': description: The pre recorded job has been successfully deleted '401': description: You don't have the permissions to delete this pre recorded job content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' '403': description: The pre recorded job is not in a deletable state content: application/json: schema: $ref: '#/components/schemas/ForbiddenErrorResponse' '404': description: The pre recorded job doesn't exist or has been deleted content: application/json: schema: $ref: '#/components/schemas/NotFoundErrorResponse' security: - x_gladia_key: [] summary: Delete the pre recorded job tags: - Pre-recorded V2 /v2/pre-recorded/{id}/file: get: operationId: PreRecordedController_getAudio_v2 parameters: - name: id required: true in: path description: Id of the pre recorded job schema: example: 45463597-20b7-4af7-b3b3-f5fb778203ab type: string responses: '200': description: The audio file used for this pre recorded job content: application/octet-stream: schema: type: string format: binary example: '401': description: You don't have the permissions to access this pre recorded job or its audio file content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' '404': description: The pre recorded job or its audio file doesn't exist or has been deleted content: application/json: schema: $ref: '#/components/schemas/NotFoundErrorResponse' security: - x_gladia_key: [] summary: Download the audio file used for this pre recorded job tags: - Pre-recorded V2 /v2/transcription: post: operationId: TranscriptionController_initPreRecordedJob_v2 parameters: [] requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/InitTranscriptionRequest' responses: '201': description: The transcription job has been initiated content: application/json: schema: $ref: '#/components/schemas/InitPreRecordedTranscriptionResponse' '400': description: Something is wrong with the request content: application/json: schema: $ref: '#/components/schemas/BadRequestErrorResponse' '401': description: You don't have the permissions to initiate a new transcription job content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' '422': description: The parameters you gave are incorrect content: application/json: schema: $ref: '#/components/schemas/UnprocessableEntityErrorResponse' security: - x_gladia_key: [] summary: Initiate a new transcription job tags: - Transcription V2 get: operationId: TranscriptionController_list_v2 parameters: - name: offset required: false in: query description: The starting point for pagination. A value of 0 starts from the first item. schema: minimum: 0 default: 0 type: integer - name: limit required: false in: query description: The maximum number of items to return. Useful for pagination and controlling data payload size. schema: minimum: 1 default: 20 type: integer - name: date required: false in: query description: Filter items relevant to a specific date in ISO format (YYYY-MM-DD). schema: format: date-time example: '2026-06-12' type: string - name: before_date required: false in: query description: Include items that occurred before the specified date in ISO format. schema: format: date-time example: '2026-06-12T21:00:09.947Z' type: string - name: after_date required: false in: query description: Filter for items after the specified date. Use with `before_date` for a range. Date in ISO format. schema: format: date-time example: '2026-06-12T21:00:09.947Z' type: string - name: status required: false in: query description: Filter the list based on item status. Accepts multiple values from the predefined list. schema: example: - done type: array items: type: string enum: - queued - processing - done - error - name: custom_metadata required: false in: query schema: additionalProperties: true example: user: John Doe type: object - name: kind required: false in: query description: Filter the list based on the item type. Supports multiple values from the predefined list. schema: example: - pre-recorded type: array items: type: string enum: - pre-recorded - live responses: '200': description: A list of transcription jobs matching the parameters. content: application/json: schema: $ref: '#/components/schemas/ListTranscriptionResponse' '401': description: You don't have the permissions to access transcription jobs content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' security: - x_gladia_key: [] summary: Get transcription jobs based on query parameters tags: - Transcription V2 /v2/transcription/{id}: get: operationId: TranscriptionController_getTranscript_v2 parameters: - name: id required: true in: path description: Id of the transcription job schema: example: 45463597-20b7-4af7-b3b3-f5fb778203ab type: string responses: '200': description: The transcription job's metadata content: application/json: schema: example: id: 45463597-20b7-4af7-b3b3-f5fb778203ab request_id: G-45463597 version: 2 kind: pre-recorded created_at: '2023-12-28T09:04:17.210Z' status: queued file: id: f0dcZE10-23d8-47f0-a25d-74a6eed88721 filename: split_infinity.wav source: http://files.gladia.io/example/audio-transcription/split_infinity.wav audio_duration: 20 number_of_channels: 1 request_params: audio_url: http://files.gladia.io/example/audio-transcription/split_infinity.wav subtitles: false diarization: false translation: false summarization: false sentences: false moderation: false named_entity_recognition: false name_consistency: false custom_spelling: false structured_data_extraction: false chapterization: false sentiment_analysis: false display_mode: false audio_enhancer: false language_config: code_switching: false languages: - fr - en accurate_words_timestamps: false diarization_enhanced: false punctuation_enhanced: false completed_at: null custom_metadata: null error_code: null result: null oneOf: - $ref: '#/components/schemas/PreRecordedResponse' - $ref: '#/components/schemas/StreamingResponse' discriminator: propertyName: kind mapping: pre-recorded: '#/components/schemas/PreRecordedResponse' live: '#/components/schemas/StreamingResponse' '401': description: You don't have the permissions to access the transcription job content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' '404': description: The transcription job doesn't exist or has been deleted content: application/json: schema: $ref: '#/components/schemas/NotFoundErrorResponse' security: - x_gladia_key: [] summary: Get the transcription job's metadata tags: - Transcription V2 delete: operationId: TranscriptionController_deleteTranscript_v2 parameters: - name: id required: true in: path description: Id of the transcription job schema: example: 45463597-20b7-4af7-b3b3-f5fb778203ab type: string responses: '202': description: The transcription job has been successfully deleted '401': description: You don't have the permissions to delete this transcription job content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' '403': description: The transcription job is not in a deletable state content: application/json: schema: $ref: '#/components/schemas/ForbiddenErrorResponse' '404': description: The transcription job doesn't exist or has been deleted content: application/json: schema: $ref: '#/components/schemas/NotFoundErrorResponse' security: - x_gladia_key: [] summary: Delete the transcription job tags: - Transcription V2 /v2/transcription/{id}/file: get: operationId: TranscriptionController_getAudio_v2 parameters: - name: id required: true in: path description: Id of the transcription job schema: example: 45463597-20b7-4af7-b3b3-f5fb778203ab type: string responses: '200': description: The audio file used for this transcription job content: application/octet-stream: schema: type: string format: binary example: '401': description: You don't have the permissions to access this transcription job or its audio file content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' '404': description: The transcription job or its audio file doesn't exist or has been deleted content: application/json: schema: $ref: '#/components/schemas/NotFoundErrorResponse' security: - x_gladia_key: [] summary: Download the audio file used for this transcription job tags: - Transcription V2 /audio/text/audio-transcription: post: operationId: AudioToTextController_audioTranscription parameters: [] requestBody: required: true content: multipart/form-data: schema: type: object properties: audio: type: string format: binary audio_url: type: string default: http://files.gladia.io/example/audio-transcription/split_infinity.wav language_behaviour: type: string enum: - automatic single language - automatic multiple languages - manual default: automatic single language language: type: string enum: - afrikaans - albanian - amharic - arabic - armenian - assamese - azerbaijani - bashkir - basque - belarusian - bengali - bosnian - breton - bulgarian - catalan - chinese - croatian - czech - danish - dutch - english - estonian - faroese - finnish - french - galician - georgian - german - greek - gujarati - haitian creole - hausa - hawaiian - hebrew - hindi - hungarian - icelandic - indonesian - italian - japanese - javanese - kannada - kazakh - khmer - korean - lao - latin - latvian - lingala - lithuanian - luxembourgish - macedonian - malagasy - malay - malayalam - maltese - maori - marathi - mongolian - myanmar - nepali - norwegian - nynorsk - occitan - pashto - persian - polish - portuguese - punjabi - romanian - russian - sanskrit - serbian - shona - sindhi - sinhala - slovak - slovenian - somali - spanish - sundanese - swahili - swedish - tagalog - tajik - tamil - tatar - telugu - thai - tibetan - turkish - turkmen - ukrainian - urdu - uzbek - vietnamese - welsh - yiddish - yoruba transcription_hint: type: string toggle_diarization: type: boolean default: false diarization_num_speakers: type: integer diarization_min_speakers: type: integer diarization_max_speakers: type: integer toggle_direct_translate: type: boolean default: false target_translation_language: type: string enum: - afrikaans - albanian - amharic - arabic - armenian - assamese - azerbaijani - bashkir - basque - belarusian - bengali - bosnian - breton - bulgarian - catalan - chinese - croatian - czech - danish - dutch - english - estonian - faroese - finnish - french - galician - georgian - german - greek - gujarati - haitian creole - hausa - hawaiian - hebrew - hindi - hungarian - icelandic - indonesian - italian - japanese - javanese - kannada - kazakh - khmer - korean - lao - latin - latvian - lingala - lithuanian - luxembourgish - macedonian - malagasy - malay - malayalam - maltese - maori - marathi - mongolian - myanmar - nepali - norwegian - nynorsk - occitan - pashto - persian - polish - portuguese - punjabi - romanian - russian - sanskrit - serbian - shona - sindhi - sinhala - slovak - slovenian - somali - spanish - sundanese - swahili - swedish - tagalog - tajik - tamil - tatar - telugu - thai - tibetan - turkish - turkmen - ukrainian - urdu - uzbek - vietnamese - welsh - wolof - yiddish - yoruba output_format: type: string enum: - json - srt - vtt - plain - txt default: json toggle_noise_reduction: type: boolean default: false toggle_accurate_words_timestamps: type: boolean default: false webhook_url: type: string responses: '200': description: '' security: - x_gladia_key: [] tags: - AudioToText - Transcription V1 /video/text/video-transcription: post: operationId: VideoToTextController_videoTranscription parameters: [] requestBody: required: true content: multipart/form-data: schema: type: object properties: video: type: string format: binary video_url: type: string default: http://files.gladia.io/example/audio-transcription/split_infinity.wav language_behaviour: type: string enum: - automatic single language - automatic multiple languages - manual default: automatic single language language: type: string enum: - afrikaans - albanian - amharic - arabic - armenian - assamese - azerbaijani - bashkir - basque - belarusian - bengali - bosnian - breton - bulgarian - catalan - chinese - croatian - czech - danish - dutch - english - estonian - faroese - finnish - french - galician - georgian - german - greek - gujarati - haitian creole - hausa - hawaiian - hebrew - hindi - hungarian - icelandic - indonesian - italian - japanese - javanese - kannada - kazakh - khmer - korean - lao - latin - latvian - lingala - lithuanian - luxembourgish - macedonian - malagasy - malay - malayalam - maltese - maori - marathi - mongolian - myanmar - nepali - norwegian - nynorsk - occitan - pashto - persian - polish - portuguese - punjabi - romanian - russian - sanskrit - serbian - shona - sindhi - sinhala - slovak - slovenian - somali - spanish - sundanese - swahili - swedish - tagalog - tajik - tamil - tatar - telugu - thai - tibetan - turkish - turkmen - ukrainian - urdu - uzbek - vietnamese - welsh - yiddish - yoruba transcription_hint: type: string toggle_diarization: type: boolean default: false diarization_num_speakers: type: integer diarization_min_speakers: type: integer diarization_max_speakers: type: integer toggle_direct_translate: type: boolean default: false target_translation_language: type: string enum: - afrikaans - albanian - amharic - arabic - armenian - assamese - azerbaijani - bashkir - basque - belarusian - bengali - bosnian - breton - bulgarian - catalan - chinese - croatian - czech - danish - dutch - english - estonian - faroese - finnish - french - galician - georgian - german - greek - gujarati - haitian creole - hausa - hawaiian - hebrew - hindi - hungarian - icelandic - indonesian - italian - japanese - javanese - kannada - kazakh - khmer - korean - lao - latin - latvian - lingala - lithuanian - luxembourgish - macedonian - malagasy - malay - malayalam - maltese - maori - marathi - mongolian - myanmar - nepali - norwegian - nynorsk - occitan - pashto - persian - polish - portuguese - punjabi - romanian - russian - sanskrit - serbian - shona - sindhi - sinhala - slovak - slovenian - somali - spanish - sundanese - swahili - swedish - tagalog - tajik - tamil - tatar - telugu - thai - tibetan - turkish - turkmen - ukrainian - urdu - uzbek - vietnamese - welsh - wolof - yiddish - yoruba output_format: type: string enum: - json - srt - vtt - plain - txt default: json toggle_noise_reduction: type: boolean default: false toggle_accurate_words_timestamps: type: boolean default: false webhook_url: type: string responses: '200': description: '' security: - x_gladia_key: [] tags: - Transcription V1 /v1/history: get: operationId: HistoryController_getList_v1 parameters: - name: offset required: false in: query description: The starting point for pagination. A value of 0 starts from the first item. schema: minimum: 0 default: 0 type: integer - name: limit required: false in: query description: The maximum number of items to return. Useful for pagination and controlling data payload size. schema: minimum: 1 default: 20 type: integer - name: date required: false in: query description: Filter items relevant to a specific date in ISO format (YYYY-MM-DD). schema: format: date-time example: '2026-06-12' type: string - name: before_date required: false in: query description: Include items that occurred before the specified date in ISO format. schema: format: date-time example: '2026-06-12T21:00:09.947Z' type: string - name: after_date required: false in: query description: Filter for items after the specified date. Use with `before_date` for a range. Date in ISO format. schema: format: date-time example: '2026-06-12T21:00:09.947Z' type: string - name: status required: false in: query description: Filter the list based on item status. Accepts multiple values from the predefined list. schema: example: - done type: array items: type: string enum: - queued - processing - done - error - name: custom_metadata required: false in: query schema: additionalProperties: true example: user: John Doe type: object - name: kind required: false in: query description: Filter the list based on the item type. Supports multiple values from the predefined list. schema: example: - pre-recorded type: array items: type: string enum: - pre-recorded - live responses: '200': description: A list of jobs content: application/json: schema: $ref: '#/components/schemas/ListHistoryResponse' security: - x_gladia_key: [] summary: Get the history of all your jobs tags: - Job History /v2/live: post: operationId: StreamingController_initStreamingSession_v2 parameters: - name: region required: false in: query description: The region used to process the audio. schema: $ref: '#/components/schemas/StreamingSupportedRegions' requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/StreamingRequest' responses: '201': description: The live job has been initiated content: application/json: schema: $ref: '#/components/schemas/InitStreamingResponse' '400': description: Something is wrong with the request content: application/json: schema: $ref: '#/components/schemas/BadRequestErrorResponse' '401': description: You don't have the permissions to initiate a new live job content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' '422': description: The parameters you gave are incorrect content: application/json: schema: $ref: '#/components/schemas/UnprocessableEntityErrorResponse' security: - x_gladia_key: [] summary: Initiate a new live job tags: - Live V2 get: operationId: StreamingController_getStreamingJobs_v2 parameters: - name: offset required: false in: query description: The starting point for pagination. A value of 0 starts from the first item. schema: minimum: 0 default: 0 type: integer - name: limit required: false in: query description: The maximum number of items to return. Useful for pagination and controlling data payload size. schema: minimum: 1 default: 20 type: integer - name: date required: false in: query description: Filter items relevant to a specific date in ISO format (YYYY-MM-DD). schema: format: date-time example: '2026-06-12' type: string - name: before_date required: false in: query description: Include items that occurred before the specified date in ISO format. schema: format: date-time example: '2026-06-12T21:00:09.947Z' type: string - name: after_date required: false in: query description: Filter for items after the specified date. Use with `before_date` for a range. Date in ISO format. schema: format: date-time example: '2026-06-12T21:00:09.947Z' type: string - name: status required: false in: query description: Filter the list based on item status. Accepts multiple values from the predefined list. schema: example: - done type: array items: type: string enum: - queued - processing - done - error - name: custom_metadata required: false in: query schema: additionalProperties: true example: user: John Doe type: object responses: '200': description: A list of live jobs matching the parameters. content: application/json: schema: $ref: '#/components/schemas/ListStreamingResponse' '401': description: You don't have the permissions to access live jobs content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' security: - x_gladia_key: [] summary: Get live jobs based on query parameters tags: - Live V2 /v2/live/{id}: get: operationId: StreamingController_getStreamingJob_v2 parameters: - name: id required: true in: path description: Id of the live job schema: example: 45463597-20b7-4af7-b3b3-f5fb778203ab type: string responses: '200': description: The live job's metadata content: application/json: schema: $ref: '#/components/schemas/StreamingResponse' '401': description: You don't have the permissions to access the live job content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' '404': description: The live job doesn't exist or has been deleted content: application/json: schema: $ref: '#/components/schemas/NotFoundErrorResponse' security: - x_gladia_key: [] summary: Get the live job's metadata tags: - Live V2 delete: operationId: StreamingController_deleteStreamingJob_v2 parameters: - name: id required: true in: path description: Id of the live job schema: example: 45463597-20b7-4af7-b3b3-f5fb778203ab type: string responses: '202': description: The live job has been successfully deleted '401': description: You don't have the permissions to delete this live job content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' '403': description: The live job is not in a deletable state content: application/json: schema: $ref: '#/components/schemas/ForbiddenErrorResponse' '404': description: The live job doesn't exist or has been deleted content: application/json: schema: $ref: '#/components/schemas/NotFoundErrorResponse' security: - x_gladia_key: [] summary: Delete the live job tags: - Live V2 patch: operationId: StreamingController_patchRequestParams_v2 parameters: - name: id required: true in: path description: Id of the live job schema: example: 45463597-20b7-4af7-b3b3-f5fb778203ab type: string requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/PatchRequestParamsDTO' responses: '204': description: Successfully patched the job '400': description: Something is wrong with the request content: application/json: schema: $ref: '#/components/schemas/BadRequestErrorResponse' '401': description: You don't have the permissions to update the job content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' '404': description: The live job doesn't exist or has been deleted content: application/json: schema: $ref: '#/components/schemas/NotFoundErrorResponse' '413': description: The post_request_metadata parameter must be a json object no longer that 100kb content: application/json: schema: $ref: '#/components/schemas/PayloadTooLargeErrorResponse' security: - x_gladia_key: [] summary: For debugging purposes, send post session metadata in the request params of the job tags: - Live V2 /v2/live/{id}/file: get: operationId: StreamingController_getAudio_v2 parameters: - name: id required: true in: path description: Id of the live job schema: example: 45463597-20b7-4af7-b3b3-f5fb778203ab type: string responses: '200': description: The audio file used for this live job content: application/octet-stream: schema: type: string format: binary example: '401': description: You don't have the permissions to access this live job or its audio file content: application/json: schema: $ref: '#/components/schemas/UnauthorizedErrorResponse' '404': description: The live job or its audio file doesn't exist or has been deleted content: application/json: schema: $ref: '#/components/schemas/NotFoundErrorResponse' security: - x_gladia_key: [] summary: Download the audio file used for this live job tags: - Live V2 /v1/models: get: operationId: ModelsController_list_v1 parameters: [] responses: '200': description: List of available models summary: List Gladia's available transcription models as per OpenRouter integration spec tags: - OpenRouter info: title: Gladia Control API description: Gladia AI audio infrastructure API for speech-to-text transcription via REST and WebSocket. Supports asynchronous pre-recorded audio processing and real-time live transcription with speaker diarization, automatic language detection across 100+ languages, and audio intelligence features. version: '1.0' contact: {} tags: [] servers: - url: https://api.gladia.io/ description: Gladia API production URL components: securitySchemes: x_gladia_key: type: apiKey in: header name: x-gladia-key description: Your personal Gladia API key schemas: PreRecordedEventPayload: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab required: - id LiveEventPayload: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab required: - id UploadBody: type: object properties: {} AudioUploadMetadataDTO: type: object properties: id: type: string description: Uploaded audio file ID format: uuid example: 6c09400e-23d2-4bd2-be55-96a5ececfa3b filename: type: string description: Uploaded audio filename example: short-audio-en-16000.wav source: type: string description: Uploaded audio source format: uri example: http://files.gladia.io/example/audio-transcription/split_infinity.wav extension: type: string description: Uploaded audio detected extension format: uuid example: wav size: type: integer description: Uploaded audio size example: 365702 audio_duration: type: number description: Uploaded audio duration example: 4.145782 number_of_channels: type: integer description: Uploaded audio channel numbers example: 1 required: - id - filename - extension - size - audio_duration - number_of_channels AudioUploadResponse: type: object properties: audio_url: type: string description: Uploaded audio file Gladia URL format: uri example: https://api.gladia.io/file/6c09400e-23d2-4bd2-be55-96a5ececfa3b audio_metadata: description: Uploaded audio file detected metadata allOf: - $ref: '#/components/schemas/AudioUploadMetadataDTO' required: - audio_url - audio_metadata TranscriptionLanguageCodeEnum: type: string enum: - af - am - ar - as - az - ba - be - bg - bn - bo - br - bs - ca - cs - cy - da - de - el - en - es - et - eu - fa - fi - fo - fr - gl - gu - ha - haw - he - hi - hr - ht - hu - hy - id - is - it - ja - jw - ka - kk - km - kn - ko - la - lb - ln - lo - lt - lv - mg - mi - mk - ml - mn - mr - ms - mt - my - ne - nl - nn - 'no' - oc - pa - pl - ps - pt - ro - ru - sa - sd - si - sk - sl - sn - so - sq - sr - su - sv - sw - ta - te - tg - th - tk - tl - tr - tt - uk - ur - uz - vi - yi - yo - zh description: Specify the language in which it will be pronounced when sound comparison occurs. Default to transcription language. CustomVocabularyEntryDTO: type: object properties: value: type: string description: The text used to replace in the transcription. example: Gladia intensity: type: number description: The global intensity of the feature. example: 0.5 minimum: 0 maximum: 1 pronunciations: description: The pronunciations used in the transcription. type: array items: type: string language: description: Specify the language in which it will be pronounced when sound comparison occurs. Default to transcription language. example: en allOf: - $ref: '#/components/schemas/TranscriptionLanguageCodeEnum' required: - value CustomVocabularyConfigDTO: type: object properties: vocabulary: type: array description: 'Specific vocabulary list to feed the transcription model with. Each item can be a string or an object with the following properties: value, intensity, pronunciations, language.' example: - Westeros - value: Stark - value: Night's Watch pronunciations: - Nightz Watch intensity: 0.4 language: en items: oneOf: - $ref: '#/components/schemas/CustomVocabularyEntryDTO' - type: string default_intensity: type: number description: Default intensity for the custom vocabulary example: 0.5 minimum: 0 maximum: 1 required: - vocabulary CallbackMethodEnum: type: string enum: - POST - PUT description: 'The HTTP method to be used. Allowed values are `POST` or `PUT` (default: `POST`)' CallbackConfigDto: type: object properties: url: type: string description: The URL to be called with the result of the transcription example: http://callback.example format: uri method: description: 'The HTTP method to be used. Allowed values are `POST` or `PUT` (default: `POST`)' example: POST default: POST allOf: - $ref: '#/components/schemas/CallbackMethodEnum' required: - url SubtitlesFormatEnum: type: string enum: - srt - vtt description: Subtitles formats you want your transcription to be formatted to SubtitlesStyleEnum: type: string enum: - default - compliance description: 'Style of the subtitles. Compliance mode refers to : https://loc.gov/preservation/digital/formats//fdd/fdd000569.shtml#:~:text=SRT%20files%20are%20basic%20text,alongside%2C%20example%3A%20%22MyVideo123' SubtitlesConfigDTO: type: object properties: formats: type: array description: Subtitles formats you want your transcription to be formatted to default: - srt minItems: 1 example: - srt items: $ref: '#/components/schemas/SubtitlesFormatEnum' minimum_duration: type: number description: Minimum duration of a subtitle in seconds minimum: 0 maximum_duration: type: number description: Maximum duration of a subtitle in seconds minimum: 1 maximum: 30 maximum_characters_per_row: type: integer description: Maximum number of characters per row in a subtitle minimum: 1 maximum_rows_per_caption: type: integer description: Maximum number of rows per caption minimum: 1 maximum: 5 style: description: 'Style of the subtitles. Compliance mode refers to : https://loc.gov/preservation/digital/formats//fdd/fdd000569.shtml#:~:text=SRT%20files%20are%20basic%20text,alongside%2C%20example%3A%20%22MyVideo123' default: default allOf: - $ref: '#/components/schemas/SubtitlesStyleEnum' DiarizationConfigDTO: type: object properties: number_of_speakers: type: integer description: Exact number of speakers in the audio example: 3 minimum: 1 min_speakers: type: integer description: Minimum number of speakers in the audio example: 1 minimum: 0 max_speakers: type: integer description: Maximum number of speakers in the audio example: 2 minimum: 0 TranslationLanguageCodeEnum: type: string enum: - af - am - ar - as - az - ba - be - bg - bn - bo - br - bs - ca - cs - cy - da - de - el - en - es - et - eu - fa - fi - fo - fr - gl - gu - ha - haw - he - hi - hr - ht - hu - hy - id - is - it - ja - jw - ka - kk - km - kn - ko - la - lb - ln - lo - lt - lv - mg - mi - mk - ml - mn - mr - ms - mt - my - ne - nl - nn - 'no' - oc - pa - pl - ps - pt - ro - ru - sa - sd - si - sk - sl - sn - so - sq - sr - su - sv - sw - ta - te - tg - th - tk - tl - tr - tt - uk - ur - uz - vi - wo - yi - yo - zh description: Target language in `iso639-1` format you want the transcription translated to TranslationModelEnum: type: string enum: - base - enhanced description: Model you want the translation model to use to translate TranslationConfigDTO: type: object properties: target_languages: type: array description: Target language in `iso639-1` format you want the transcription translated to example: - en minItems: 1 items: $ref: '#/components/schemas/TranslationLanguageCodeEnum' model: description: Model you want the translation model to use to translate default: base allOf: - $ref: '#/components/schemas/TranslationModelEnum' match_original_utterances: type: boolean description: Align translated utterances with the original ones default: true lipsync: type: boolean description: 'Whether to apply lipsync to the translated transcription. ' default: true context_adaptation: type: boolean description: Enables or disables context-aware translation features that allow the model to adapt translations based on provided context. default: true context: type: string description: Context information to improve translation accuracy informal: type: boolean description: Forces the translation to use informal language forms when available in the target language. default: false required: - target_languages SummaryTypesEnum: type: string enum: - general - bullet_points - concise description: The type of summarization to apply SummarizationConfigDTO: type: object properties: type: description: The type of summarization to apply default: general allOf: - $ref: '#/components/schemas/SummaryTypesEnum' CustomSpellingConfigDTO: type: object properties: spelling_dictionary: type: object description: The list of spelling applied on the audio transcription example: Gettleman: - gettleman SQL: - Sequel additionalProperties: type: array items: type: string required: - spelling_dictionary AudioToLlmListConfigDTO: type: object properties: prompts: description: The list of prompts applied on the audio transcription example: - Extract the key points from the transcription minItems: 1 type: array items: type: array model: type: string description: The model to use for the prompt execution. You can find the list of supported models [here](https://openrouter.ai/models). default: openai/gpt-5.4-nano required: - prompts PiiRedactionEntityTypeEnum: type: string enum: - APPI - APPI_SENSITIVE - CCI - CORE_ENTITIES - CPRA - GDPR - GDPR_SENSITIVE - HEALTH_INFORMATION - HIPAA_SAFE_HARBOR - LIDI - NUMERICAL_EXCL_PCI - PCI - QUEBEC_PRIVACY_ACT - ACCOUNT_NUMBER - AGE - DATE - DATE_INTERVAL - DOB - DRIVER_LICENSE - DURATION - EMAIL_ADDRESS - EVENT - FILENAME - GENDER - HEALTHCARE_NUMBER - IP_ADDRESS - LANGUAGE - LOCATION - LOCATION_ADDRESS - LOCATION_ADDRESS_STREET - LOCATION_CITY - LOCATION_COORDINATE - LOCATION_COUNTRY - LOCATION_STATE - LOCATION_ZIP - MARITAL_STATUS - MONEY - NAME - NAME_FAMILY - NAME_GIVEN - NAME_MEDICAL_PROFESSIONAL - NUMERICAL_PII - OCCUPATION - ORGANIZATION - ORGANIZATION_MEDICAL_FACILITY - ORIGIN - PASSPORT_NUMBER - PASSWORD - PHONE_NUMBER - PHYSICAL_ATTRIBUTE - POLITICAL_AFFILIATION - RELIGION - SEXUALITY - SSN - TIME - URL - USERNAME - VEHICLE_ID - ZODIAC_SIGN - BLOOD_TYPE - CONDITION - DOSE - DRUG - INJURY - MEDICAL_PROCESS - STATISTICS - BANK_ACCOUNT - CREDIT_CARD - CREDIT_CARD_EXPIRATION - CVV - ROUTING_NUMBER - CORPORATE_ACTION - DAY - EFFECT - FINANCIAL_METRIC - MEDICAL_CODE - MONTH - ORGANIZATION_ID - PRODUCT - PROJECT - TREND - YEAR description: The entity types to redact PiiRedactionConfigDTO: type: object properties: entity_types: description: The entity types to redact example: - GDPR - HEALTH_INFORMATION - HIPAA_SAFE_HARBOR - QUEBEC_PRIVACY_ACT - EMAIL_ADDRESS - NAME - PHONE_NUMBER allOf: - $ref: '#/components/schemas/PiiRedactionEntityTypeEnum' processed_text_type: type: string description: The type of processed text to return (marker or mask) enum: - MARKER - MASK example: MARKER LanguageConfig: type: object properties: languages: type: array description: If one language is set, it will be used for the transcription. Otherwise, language will be auto-detected by the model. default: [] items: $ref: '#/components/schemas/TranscriptionLanguageCodeEnum' code_switching: type: boolean description: If true, language will be auto-detected on each utterance. Otherwise, language will be auto-detected on first utterance and then used for the rest of the transcription. If one language is set, this option will be ignored. default: false InitTranscriptionRequest: type: object properties: custom_vocabulary: type: boolean description: '**[Beta]** Can be either boolean to enable custom_vocabulary for this audio or an array with specific vocabulary list to feed the transcription model with' default: false custom_vocabulary_config: description: '**[Beta]** Custom vocabulary configuration, if `custom_vocabulary` is enabled' allOf: - $ref: '#/components/schemas/CustomVocabularyConfigDTO' callback_url: type: string description: '**[Deprecated]** Use `callback`/`callback_config` instead. Callback URL we will do a `POST` request to with the result of the transcription' example: http://callback.example format: uri deprecated: true callback: type: boolean description: Enable callback for this transcription. If true, the `callback_config` property will be used to customize the callback behaviour default: false callback_config: description: Customize the callback behaviour (url and http method) allOf: - $ref: '#/components/schemas/CallbackConfigDto' subtitles: type: boolean description: Enable subtitles generation for this transcription default: false subtitles_config: description: Configuration for subtitles generation if `subtitles` is enabled allOf: - $ref: '#/components/schemas/SubtitlesConfigDTO' diarization: type: boolean description: Enable speaker recognition (diarization) for this audio default: false diarization_config: description: Speaker recognition configuration, if `diarization` is enabled allOf: - $ref: '#/components/schemas/DiarizationConfigDTO' translation: type: boolean description: '**[Beta]** Enable translation for this audio' default: false translation_config: description: '**[Beta]** Translation configuration, if `translation` is enabled' allOf: - $ref: '#/components/schemas/TranslationConfigDTO' summarization: type: boolean description: Enable summarization for this audio default: false summarization_config: description: Summarization configuration, if `summarization` is enabled allOf: - $ref: '#/components/schemas/SummarizationConfigDTO' named_entity_recognition: type: boolean description: '**[Alpha]** Enable named entity recognition for this audio' default: false custom_spelling: type: boolean description: '**[Alpha]** Enable custom spelling for this audio' default: false custom_spelling_config: description: '**[Alpha]** Custom spelling configuration, if `custom_spelling` is enabled' allOf: - $ref: '#/components/schemas/CustomSpellingConfigDTO' sentiment_analysis: type: boolean description: Enable sentiment analysis for this audio default: false audio_to_llm: type: boolean description: Enable audio to LLM processing for this audio default: false audio_to_llm_config: description: Audio to LLM configuration, if `audio_to_llm` is enabled allOf: - $ref: '#/components/schemas/AudioToLlmListConfigDTO' pii_redaction: type: boolean description: Enable PII redaction for this audio default: false pii_redaction_config: description: PII redaction configuration, if `pii_redaction` is enabled allOf: - $ref: '#/components/schemas/PiiRedactionConfigDTO' custom_metadata: type: object description: Custom metadata you can attach to this transcription example: user: John Doe additionalProperties: true sentences: type: boolean description: Enable sentences for this audio default: false punctuation_enhanced: type: boolean description: '**[Alpha]** Use enhanced punctuation for this audio' default: false language_config: description: Specify the language configuration allOf: - $ref: '#/components/schemas/LanguageConfig' audio_url: type: string description: URL to a Gladia file or to an external audio or video file example: http://files.gladia.io/example/audio-transcription/split_infinity.wav format: uri required: - audio_url InitPreRecordedTranscriptionResponse: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab result_url: type: string description: Prebuilt URL with your transcription `id` to fetch the result example: https://api.gladia.io/v2/transcription/45463597-20b7-4af7-b3b3-f5fb778203ab format: uri required: - id - result_url BadRequestErrorResponse: type: object properties: timestamp: type: string description: Date of when the error occurred example: '2023-12-28T09:04:17.210Z' path: type: string description: Path to the API endpoint example: /v2/transcription/45463597-20b7-4af7-b3b3-f5fb778203ab request_id: type: string description: Debug id example: G-821fe9df statusCode: type: number description: HTTP status code of the error example: 400 message: type: string description: Error message example: Content-Type is missing Multipart Boundary. validation_errors: description: List of validation errors, if any example: - Field "language" must be a string - Field "min_speakers" must be a number type: array items: type: string required: - timestamp - path - request_id - statusCode - message UnauthorizedErrorResponse: type: object properties: timestamp: type: string description: Date of when the error occurred example: '2023-12-28T09:04:17.210Z' path: type: string description: Path to the API endpoint example: /v2/transcription/45463597-20b7-4af7-b3b3-f5fb778203ab request_id: type: string description: Debug id example: G-821fe9df statusCode: type: number description: HTTP status code of the error example: 401 message: type: string description: Error message example: gladia key not found required: - timestamp - path - request_id - statusCode - message UnprocessableEntityErrorResponse: type: object properties: timestamp: type: string description: Date of when the error occurred example: '2023-12-28T09:04:17.210Z' path: type: string description: Path to the API endpoint example: /v2/transcription/45463597-20b7-4af7-b3b3-f5fb778203ab request_id: type: string description: Debug id example: G-821fe9df statusCode: type: number description: HTTP status code of the error example: 422 message: type: string description: Error message example: Invalid parameter required: - timestamp - path - request_id - statusCode - message FileResponse: type: object properties: id: type: string description: The file id filename: type: string nullable: true description: The name of the uploaded file source: type: string nullable: true description: The link used to download the file if audio_url was used audio_duration: type: number nullable: true description: Duration of the audio file example: 3600 number_of_channels: type: integer nullable: true description: Number of channels in the audio file minimum: 1 example: 1 required: - id - filename - source - audio_duration - number_of_channels PreRecordedRequestParamsResponse: type: object properties: custom_vocabulary: type: boolean description: '**[Beta]** Can be either boolean to enable custom_vocabulary for this audio or an array with specific vocabulary list to feed the transcription model with' default: false custom_vocabulary_config: description: '**[Beta]** Custom vocabulary configuration, if `custom_vocabulary` is enabled' allOf: - $ref: '#/components/schemas/CustomVocabularyConfigDTO' callback_url: type: string description: '**[Deprecated]** Use `callback`/`callback_config` instead. Callback URL we will do a `POST` request to with the result of the transcription' example: http://callback.example format: uri deprecated: true callback: type: boolean description: Enable callback for this transcription. If true, the `callback_config` property will be used to customize the callback behaviour default: false callback_config: description: Customize the callback behaviour (url and http method) allOf: - $ref: '#/components/schemas/CallbackConfigDto' subtitles: type: boolean description: Enable subtitles generation for this transcription default: false subtitles_config: description: Configuration for subtitles generation if `subtitles` is enabled allOf: - $ref: '#/components/schemas/SubtitlesConfigDTO' diarization: type: boolean description: Enable speaker recognition (diarization) for this audio default: false diarization_config: description: Speaker recognition configuration, if `diarization` is enabled allOf: - $ref: '#/components/schemas/DiarizationConfigDTO' translation: type: boolean description: '**[Beta]** Enable translation for this audio' default: false translation_config: description: '**[Beta]** Translation configuration, if `translation` is enabled' allOf: - $ref: '#/components/schemas/TranslationConfigDTO' summarization: type: boolean description: Enable summarization for this audio default: false summarization_config: description: Summarization configuration, if `summarization` is enabled allOf: - $ref: '#/components/schemas/SummarizationConfigDTO' named_entity_recognition: type: boolean description: '**[Alpha]** Enable named entity recognition for this audio' default: false custom_spelling: type: boolean description: '**[Alpha]** Enable custom spelling for this audio' default: false custom_spelling_config: description: '**[Alpha]** Custom spelling configuration, if `custom_spelling` is enabled' allOf: - $ref: '#/components/schemas/CustomSpellingConfigDTO' sentiment_analysis: type: boolean description: Enable sentiment analysis for this audio default: false audio_to_llm: type: boolean description: Enable audio to LLM processing for this audio default: false audio_to_llm_config: description: Audio to LLM configuration, if `audio_to_llm` is enabled allOf: - $ref: '#/components/schemas/AudioToLlmListConfigDTO' pii_redaction: type: boolean description: Enable PII redaction for this audio default: false pii_redaction_config: description: PII redaction configuration, if `pii_redaction` is enabled allOf: - $ref: '#/components/schemas/PiiRedactionConfigDTO' sentences: type: boolean description: Enable sentences for this audio default: false punctuation_enhanced: type: boolean description: '**[Alpha]** Use enhanced punctuation for this audio' default: false language_config: description: Specify the language configuration allOf: - $ref: '#/components/schemas/LanguageConfig' audio_url: type: string format: uri nullable: true required: - audio_url TranscriptionMetadataDTO: type: object properties: audio_duration: type: number description: Duration of the transcribed audio file example: 3600 number_of_distinct_channels: type: integer description: Number of distinct channels in the transcribed audio file minimum: 1 example: 1 billing_time: type: number description: Billed duration in seconds (audio_duration * number_of_distinct_channels) example: 3600 transcription_time: type: number description: Duration of the transcription in seconds example: 20 required: - audio_duration - number_of_distinct_channels - billing_time - transcription_time AddonErrorDTO: type: object properties: status_code: type: integer description: Status code of the addon error example: 500 exception: type: string description: Reason of the addon error message: type: string description: Detailed message of the addon error required: - status_code - exception - message SentencesDTO: type: object properties: success: type: boolean description: The audio intelligence model succeeded to get a valid output is_empty: type: boolean description: The audio intelligence model returned an empty value exec_time: type: number description: Time audio intelligence model took to complete the task error: description: '`null` if `success` is `true`. Contains the error details of the failed model' nullable: true allOf: - $ref: '#/components/schemas/AddonErrorDTO' results: description: If `sentences` has been enabled, transcription as sentences. nullable: true type: array items: type: string required: - success - is_empty - exec_time - error - results SubtitleDTO: type: object properties: format: description: Format of the current subtitle example: srt allOf: - $ref: '#/components/schemas/SubtitlesFormatEnum' subtitles: type: string description: Transcription on the asked subtitle format required: - format - subtitles WordDTO: type: object properties: word: type: string description: Spoken word start: type: number description: Start timestamps in seconds of the spoken word end: type: number description: End timestamps in seconds of the spoken word confidence: type: number description: Confidence on the transcribed word (1 = 100% confident) required: - word - start - end - confidence UtteranceDTO: type: object properties: start: type: number description: Start timestamp in seconds of this utterance end: type: number description: End timestamp in seconds of this utterance confidence: type: number description: Confidence on the transcribed utterance (1 = 100% confident) channel: type: integer description: Audio channel of where this utterance has been transcribed from minimum: 0 speaker: type: integer description: If `diarization` enabled, speaker identification number minimum: 0 words: description: List of words of the utterance, split by timestamp type: array items: $ref: '#/components/schemas/WordDTO' text: type: string description: Transcription for this utterance language: description: Spoken language in this utterance example: en allOf: - $ref: '#/components/schemas/TranscriptionLanguageCodeEnum' required: - start - end - confidence - channel - words - text - language TranscriptionDTO: type: object properties: full_transcript: type: string description: All transcription on text format without any other information languages: type: array description: All the detected languages in the audio sorted from the most detected to the less detected example: - en items: $ref: '#/components/schemas/TranscriptionLanguageCodeEnum' sentences: description: If `sentences` has been enabled, sentences results type: array items: $ref: '#/components/schemas/SentencesDTO' subtitles: description: If `subtitles` has been enabled, subtitles results type: array items: $ref: '#/components/schemas/SubtitleDTO' utterances: description: Transcribed speech utterances present in the audio type: array items: $ref: '#/components/schemas/UtteranceDTO' required: - full_transcript - languages - utterances TranslationResultDTO: type: object properties: error: description: Contains the error details of the failed addon nullable: true allOf: - $ref: '#/components/schemas/AddonErrorDTO' full_transcript: type: string description: All transcription on text format without any other information languages: type: array description: All the detected languages in the audio sorted from the most detected to the less detected example: - en items: $ref: '#/components/schemas/TranslationLanguageCodeEnum' sentences: description: If `sentences` has been enabled, sentences results for this translation type: array items: $ref: '#/components/schemas/SentencesDTO' subtitles: description: If `subtitles` has been enabled, subtitles results for this translation type: array items: $ref: '#/components/schemas/SubtitleDTO' utterances: description: Transcribed speech utterances present in the audio type: array items: $ref: '#/components/schemas/UtteranceDTO' required: - error - full_transcript - languages - utterances TranslationDTO: type: object properties: success: type: boolean description: The audio intelligence model succeeded to get a valid output is_empty: type: boolean description: The audio intelligence model returned an empty value exec_time: type: number description: Time audio intelligence model took to complete the task error: description: '`null` if `success` is `true`. Contains the error details of the failed model' nullable: true allOf: - $ref: '#/components/schemas/AddonErrorDTO' results: description: List of translated transcriptions, one for each `target_languages` nullable: true type: array items: $ref: '#/components/schemas/TranslationResultDTO' required: - success - is_empty - exec_time - error - results SummarizationDTO: type: object properties: success: type: boolean description: The audio intelligence model succeeded to get a valid output is_empty: type: boolean description: The audio intelligence model returned an empty value exec_time: type: number description: Time audio intelligence model took to complete the task error: description: '`null` if `success` is `true`. Contains the error details of the failed model' nullable: true allOf: - $ref: '#/components/schemas/AddonErrorDTO' results: type: string description: If `summarization` has been enabled, summary of the transcription nullable: true required: - success - is_empty - exec_time - error - results ModerationDTO: type: object properties: success: type: boolean description: The audio intelligence model succeeded to get a valid output is_empty: type: boolean description: The audio intelligence model returned an empty value exec_time: type: number description: Time audio intelligence model took to complete the task error: description: '`null` if `success` is `true`. Contains the error details of the failed model' nullable: true allOf: - $ref: '#/components/schemas/AddonErrorDTO' results: type: string description: If `moderation` has been enabled, moderated transcription nullable: true required: - success - is_empty - exec_time - error - results NamedEntityRecognitionResult: type: object properties: entity_type: type: string text: type: string start: type: number end: type: number required: - entity_type - text - start - end NamedEntityRecognitionDTO: type: object properties: success: type: boolean description: The audio intelligence model succeeded to get a valid output is_empty: type: boolean description: The audio intelligence model returned an empty value exec_time: type: number description: Time audio intelligence model took to complete the task error: description: '`null` if `success` is `true`. Contains the error details of the failed model' nullable: true allOf: - $ref: '#/components/schemas/AddonErrorDTO' results: description: If `named_entity_recognition` has been enabled, the detected entities. nullable: true type: array items: $ref: '#/components/schemas/NamedEntityRecognitionResult' required: - success - is_empty - exec_time - error - results NamesConsistencyDTO: type: object properties: success: type: boolean description: The audio intelligence model succeeded to get a valid output is_empty: type: boolean description: The audio intelligence model returned an empty value exec_time: type: number description: Time audio intelligence model took to complete the task error: description: '`null` if `success` is `true`. Contains the error details of the failed model' nullable: true allOf: - $ref: '#/components/schemas/AddonErrorDTO' results: type: string description: Deprecated, If `name_consistency` has been enabled, Gladia will improve the consistency of the names across the transcription required: - success - is_empty - exec_time - error - results StructuredDataExtractionDTO: type: object properties: success: type: boolean description: The audio intelligence model succeeded to get a valid output is_empty: type: boolean description: The audio intelligence model returned an empty value exec_time: type: number description: Time audio intelligence model took to complete the task error: description: '`null` if `success` is `true`. Contains the error details of the failed model' nullable: true allOf: - $ref: '#/components/schemas/AddonErrorDTO' results: type: string description: If `structured_data_extraction` has been enabled, results of the AI structured data extraction for the defined classes. required: - success - is_empty - exec_time - error - results SentimentAnalysisDTO: type: object properties: success: type: boolean description: The audio intelligence model succeeded to get a valid output is_empty: type: boolean description: The audio intelligence model returned an empty value exec_time: type: number description: Time audio intelligence model took to complete the task error: description: '`null` if `success` is `true`. Contains the error details of the failed model' nullable: true allOf: - $ref: '#/components/schemas/AddonErrorDTO' results: type: string description: If `sentiment_analysis` has been enabled, Gladia will analyze the sentiments and emotions of the audio required: - success - is_empty - exec_time - error - results AudioToLlmResultDTO: type: object properties: prompt: type: string description: The prompt used nullable: true response: type: string description: The result of the AI analysis nullable: true required: - prompt - response AudioToLlmDTO: type: object properties: success: type: boolean description: The audio intelligence model succeeded to get a valid output is_empty: type: boolean description: The audio intelligence model returned an empty value exec_time: type: number description: Time audio intelligence model took to complete the task error: description: '`null` if `success` is `true`. Contains the error details of the failed model' nullable: true allOf: - $ref: '#/components/schemas/AddonErrorDTO' results: description: The result from a specific prompt nullable: true allOf: - $ref: '#/components/schemas/AudioToLlmResultDTO' required: - success - is_empty - exec_time - error - results AudioToLlmListDTO: type: object properties: success: type: boolean description: The audio intelligence model succeeded to get a valid output is_empty: type: boolean description: The audio intelligence model returned an empty value exec_time: type: number description: Time audio intelligence model took to complete the task error: description: '`null` if `success` is `true`. Contains the error details of the failed model' nullable: true allOf: - $ref: '#/components/schemas/AddonErrorDTO' results: description: If `audio_to_llm` has been enabled, results of the AI custom analysis nullable: true type: array items: $ref: '#/components/schemas/AudioToLlmDTO' required: - success - is_empty - exec_time - error - results DisplayModeDTO: type: object properties: success: type: boolean description: The audio intelligence model succeeded to get a valid output is_empty: type: boolean description: The audio intelligence model returned an empty value exec_time: type: number description: Time audio intelligence model took to complete the task error: description: '`null` if `success` is `true`. Contains the error details of the failed model' nullable: true allOf: - $ref: '#/components/schemas/AddonErrorDTO' results: description: If `display_mode` has been enabled, proposes an alternative display output. nullable: true type: array items: type: string required: - success - is_empty - exec_time - error - results ChapterizationDTO: type: object properties: success: type: boolean description: The audio intelligence model succeeded to get a valid output is_empty: type: boolean description: The audio intelligence model returned an empty value exec_time: type: number description: Time audio intelligence model took to complete the task error: description: '`null` if `success` is `true`. Contains the error details of the failed model' nullable: true allOf: - $ref: '#/components/schemas/AddonErrorDTO' results: type: object description: If `chapterization` has been enabled, will generate chapters name for different parts of the given audio. additionalProperties: true required: - success - is_empty - exec_time - error - results DiarizationDTO: type: object properties: success: type: boolean description: The audio intelligence model succeeded to get a valid output is_empty: type: boolean description: The audio intelligence model returned an empty value exec_time: type: number description: Time audio intelligence model took to complete the task error: description: '`null` if `success` is `true`. Contains the error details of the failed model' nullable: true allOf: - $ref: '#/components/schemas/AddonErrorDTO' results: description: '[Deprecated] If `diarization` has been enabled, the diarization result will appear here' type: array items: $ref: '#/components/schemas/UtteranceDTO' required: - success - is_empty - exec_time - error - results TranscriptionResultDTO: type: object properties: metadata: description: Metadata for the given transcription & audio file allOf: - $ref: '#/components/schemas/TranscriptionMetadataDTO' transcription: description: Transcription of the audio speech allOf: - $ref: '#/components/schemas/TranscriptionDTO' translation: description: If `translation` has been enabled, translation of the audio speech transcription allOf: - $ref: '#/components/schemas/TranslationDTO' summarization: description: If `summarization` has been enabled, summarization of the audio speech transcription allOf: - $ref: '#/components/schemas/SummarizationDTO' moderation: description: If `moderation` has been enabled, moderation of the audio speech transcription allOf: - $ref: '#/components/schemas/ModerationDTO' named_entity_recognition: description: If `named_entity_recognition` has been enabled, the detected entities allOf: - $ref: '#/components/schemas/NamedEntityRecognitionDTO' name_consistency: description: If `name_consistency` has been enabled, Gladia will improve consistency of the names accross the transcription allOf: - $ref: '#/components/schemas/NamesConsistencyDTO' structured_data_extraction: description: If `structured_data_extraction` has been enabled, structured data extraction results allOf: - $ref: '#/components/schemas/StructuredDataExtractionDTO' sentiment_analysis: description: If `sentiment_analysis` has been enabled, sentiment analysis of the audio speech transcription allOf: - $ref: '#/components/schemas/SentimentAnalysisDTO' audio_to_llm: description: If `audio_to_llm` has been enabled, audio to llm results of the audio speech transcription allOf: - $ref: '#/components/schemas/AudioToLlmListDTO' sentences: description: 'If `sentences` has been enabled, sentences of the audio speech transcription. Deprecated: content will move to the `transcription` object.' deprecated: true allOf: - $ref: '#/components/schemas/SentencesDTO' display_mode: description: If `display_mode` has been enabled, the output will be reordered, creating new utterances when speakers overlapped allOf: - $ref: '#/components/schemas/DisplayModeDTO' chapterization: description: If `chapterization` has been enabled, will generate chapters name for different parts of the given audio. allOf: - $ref: '#/components/schemas/ChapterizationDTO' diarization: description: If `diarization` has been requested and an error has occurred, the result will appear here allOf: - $ref: '#/components/schemas/DiarizationDTO' required: - metadata PreRecordedResponse: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab request_id: type: string description: Debug id example: G-45463597 version: type: integer description: API version example: 2 status: type: string description: '"queued": the job has been queued. "processing": the job is being processed. "done": the job has been processed and the result is available. "error": an error occurred during the job''s processing.' enum: - queued - processing - done - error created_at: type: string description: Creation date format: date-time example: '2023-12-28T09:04:17.210Z' completed_at: type: string description: Completion date when status is "done" or "error" format: date-time example: '2023-12-28T09:04:37.210Z' nullable: true custom_metadata: type: object description: Custom metadata given in the initial request example: user: John Doe additionalProperties: true error_code: type: integer description: HTTP status code of the error if status is "error" minimum: 400 maximum: 599 example: 500 nullable: true post_session_metadata: type: object description: For debugging purposes, send data that could help to identify issues kind: type: string enum: - pre-recorded example: pre-recorded default: pre-recorded file: description: The file data you uploaded. Can be null if status is "error" nullable: true allOf: - $ref: '#/components/schemas/FileResponse' request_params: description: Parameters used for this pre-recorded transcription. Can be null if status is "error" nullable: true allOf: - $ref: '#/components/schemas/PreRecordedRequestParamsResponse' result: description: Pre-recorded transcription's result when status is "done" nullable: true allOf: - $ref: '#/components/schemas/TranscriptionResultDTO' required: - id - request_id - version - status - created_at - post_session_metadata - kind NotFoundErrorResponse: type: object properties: timestamp: type: string description: Date of when the error occurred example: '2023-12-28T09:04:17.210Z' path: type: string description: Path to the API endpoint example: /v2/transcription/45463597-20b7-4af7-b3b3-f5fb778203ab request_id: type: string description: Debug id example: G-821fe9df statusCode: type: number description: HTTP status code of the error example: 404 message: type: string description: Error message example: Not found required: - timestamp - path - request_id - statusCode - message ListPreRecordedResponse: type: object properties: first: type: string description: URL to fetch the first page format: uri example: https://api.gladia.io/v2/transcription?status=done&offset=0&limit=20 current: type: string description: URL to fetch the current page format: uri example: https://api.gladia.io/v2/transcription?status=done&offset=0&limit=20 next: type: string description: URL to fetch the next page format: uri example: https://api.gladia.io/v2/transcription?status=done&offset=20&limit=20 nullable: true items: description: List of pre-recorded transcriptions type: array items: $ref: '#/components/schemas/PreRecordedResponse' required: - first - current - next - items ForbiddenErrorResponse: type: object properties: timestamp: type: string description: Date of when the error occurred example: '2023-12-28T09:04:17.210Z' path: type: string description: Path to the API endpoint example: /v2/transcription/45463597-20b7-4af7-b3b3-f5fb778203ab request_id: type: string description: Debug id example: G-821fe9df statusCode: type: number description: HTTP status code of the error example: 403 message: type: string description: Forbidden request example: Invalid parameter required: - timestamp - path - request_id - statusCode - message CallbackTranscriptionSuccessPayload: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string description: Type of event enum: - transcription.success example: transcription.success default: transcription.success payload: description: Result of the transcription allOf: - $ref: '#/components/schemas/TranscriptionResultDTO' custom_metadata: type: object description: Custom metadata given in the initial request nullable: true example: user: John Doe additionalProperties: true required: - id - event - payload ErrorDTO: type: object properties: code: type: integer description: Error code example: 400 message: type: string description: Error message example: Bad Request required: - code - message CallbackTranscriptionErrorPayload: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string description: Type of event enum: - transcription.error example: transcription.error default: transcription.error error: description: The error that occurred during the transcription allOf: - $ref: '#/components/schemas/ErrorDTO' custom_metadata: type: object description: Custom metadata given in the initial request nullable: true example: user: John Doe additionalProperties: true required: - id - event - error StreamingSupportedEncodingEnum: type: string enum: - wav/pcm - wav/alaw - wav/ulaw description: "The encoding format of the audio stream. Supported formats: \n- PCM: 8, 16, 24, and 32 bits \n- A-law:\ \ 8 bits \n- μ-law: 8 bits \n\nNote: No need to add WAV headers to raw audio as the API supports both formats." StreamingSupportedBitDepthEnum: type: number enum: - 8 - 16 - 24 - 32 description: The bit depth of the audio stream StreamingSupportedSampleRateEnum: type: number enum: - 8000 - 16000 - 32000 - 44100 - 48000 description: The sample rate of the audio stream StreamingSupportedModels: type: string enum: - solaria-1 description: The model used to process the audio. "solaria-1" is used by default. PreProcessingConfig: type: object properties: audio_enhancer: type: boolean description: If true, apply pre-processing to the audio stream to enhance the quality. default: false speech_threshold: type: number description: Sensitivity configuration for Speech Threshold. A value close to 1 will apply stricter thresholds, making it less likely to detect background sounds as speech. default: 0.6 minimum: 0 maximum: 1 RealtimeProcessingConfig: type: object properties: custom_vocabulary: type: boolean description: If true, enable custom vocabulary for the transcription. default: false custom_vocabulary_config: description: Custom vocabulary configuration, if `custom_vocabulary` is enabled allOf: - $ref: '#/components/schemas/CustomVocabularyConfigDTO' custom_spelling: type: boolean description: If true, enable custom spelling for the transcription. default: false custom_spelling_config: description: Custom spelling configuration, if `custom_spelling` is enabled allOf: - $ref: '#/components/schemas/CustomSpellingConfigDTO' translation: type: boolean description: If true, enable translation for the transcription default: false translation_config: description: Translation configuration, if `translation` is enabled allOf: - $ref: '#/components/schemas/TranslationConfigDTO' named_entity_recognition: type: boolean description: If true, enable named entity recognition for the transcription. default: false sentiment_analysis: type: boolean description: If true, enable sentiment analysis for the transcription. default: false PostProcessingConfig: type: object properties: summarization: type: boolean description: If true, generates summarization for the whole transcription. default: false summarization_config: description: Summarization configuration, if `summarization` is enabled allOf: - $ref: '#/components/schemas/SummarizationConfigDTO' chapterization: type: boolean description: If true, generates chapters for the whole transcription. default: false MessagesConfig: type: object properties: receive_partial_transcripts: type: boolean description: If true, partial transcript will be sent to websocket. default: false receive_final_transcripts: type: boolean description: If true, final transcript will be sent to websocket. default: true receive_speech_events: type: boolean description: If true, begin and end speech events will be sent to websocket. default: true receive_pre_processing_events: type: boolean description: If true, pre-processing events will be sent to websocket. default: true receive_realtime_processing_events: type: boolean description: If true, realtime processing events will be sent to websocket. default: true receive_post_processing_events: type: boolean description: If true, post-processing events will be sent to websocket. default: true receive_acknowledgments: type: boolean description: If true, acknowledgments will be sent to websocket. default: true receive_errors: type: boolean description: If true, errors will be sent to websocket. default: true receive_lifecycle_events: type: boolean description: If true, lifecycle events will be sent to websocket. default: false CallbackConfig: type: object properties: url: type: string description: URL on which we will do a `POST` request with configured messages example: https://callback.example format: uri receive_partial_transcripts: type: boolean description: If true, partial transcript will be sent to the defined callback. default: false receive_final_transcripts: type: boolean description: If true, final transcript will be sent to the defined callback. default: true receive_speech_events: type: boolean description: If true, begin and end speech events will be sent to the defined callback. default: false receive_pre_processing_events: type: boolean description: If true, pre-processing events will be sent to the defined callback. default: true receive_realtime_processing_events: type: boolean description: If true, realtime processing events will be sent to the defined callback. default: true receive_post_processing_events: type: boolean description: If true, post-processing events will be sent to the defined callback. default: true receive_acknowledgments: type: boolean description: If true, acknowledgments will be sent to the defined callback. default: false receive_errors: type: boolean description: If true, errors will be sent to the defined callback. default: false receive_lifecycle_events: type: boolean description: If true, lifecycle events will be sent to the defined callback. default: true StreamingRequestParamsResponse: type: object properties: encoding: description: "The encoding format of the audio stream. Supported formats: \n- PCM: 8, 16, 24, and 32 bits \n- A-law:\ \ 8 bits \n- μ-law: 8 bits \n\nNote: No need to add WAV headers to raw audio as the API supports both formats." default: wav/pcm allOf: - $ref: '#/components/schemas/StreamingSupportedEncodingEnum' bit_depth: description: The bit depth of the audio stream default: 16 allOf: - $ref: '#/components/schemas/StreamingSupportedBitDepthEnum' sample_rate: description: The sample rate of the audio stream default: 16000 allOf: - $ref: '#/components/schemas/StreamingSupportedSampleRateEnum' channels: type: integer description: The number of channels of the audio stream default: 1 minimum: 1 maximum: 8 model: description: The model used to process the audio. "solaria-1" is used by default. default: solaria-1 allOf: - $ref: '#/components/schemas/StreamingSupportedModels' endpointing: type: number description: The endpointing duration in seconds. Endpointing is the duration of silence which will cause an utterance to be considered as finished default: 0.05 minimum: 0.01 maximum: 10 maximum_duration_without_endpointing: type: number description: The maximum duration in seconds without endpointing. If endpointing is not detected after this duration, current utterance will be considered as finished default: 5 minimum: 5 maximum: 60 language_config: description: Specify the language configuration allOf: - $ref: '#/components/schemas/LanguageConfig' pre_processing: description: Specify the pre-processing configuration allOf: - $ref: '#/components/schemas/PreProcessingConfig' realtime_processing: description: Specify the realtime processing configuration allOf: - $ref: '#/components/schemas/RealtimeProcessingConfig' post_processing: description: Specify the post-processing configuration allOf: - $ref: '#/components/schemas/PostProcessingConfig' messages_config: description: Specify the websocket messages configuration allOf: - $ref: '#/components/schemas/MessagesConfig' callback: type: boolean description: If true, messages will be sent to configured url. default: false callback_config: description: Specify the callback configuration allOf: - $ref: '#/components/schemas/CallbackConfig' StreamingTranscriptionResultWithMessagesDTO: type: object properties: metadata: description: Metadata for the given transcription & audio file allOf: - $ref: '#/components/schemas/TranscriptionMetadataDTO' transcription: description: Transcription of the audio speech allOf: - $ref: '#/components/schemas/TranscriptionDTO' translation: description: If `translation` has been enabled, translation of the audio speech transcription allOf: - $ref: '#/components/schemas/TranslationDTO' summarization: description: If `summarization` has been enabled, summarization of the audio speech transcription allOf: - $ref: '#/components/schemas/SummarizationDTO' named_entity_recognition: description: If `named_entity_recognition` has been enabled, the detected entities allOf: - $ref: '#/components/schemas/NamedEntityRecognitionDTO' sentiment_analysis: description: If `sentiment_analysis` has been enabled, sentiment analysis of the audio speech transcription allOf: - $ref: '#/components/schemas/SentimentAnalysisDTO' chapterization: description: If `chapterization` has been enabled, will generate chapters name for different parts of the given audio. allOf: - $ref: '#/components/schemas/ChapterizationDTO' messages: description: Real-Time messages sent by the server during the live transcription type: array items: type: string required: - metadata StreamingResponse: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab request_id: type: string description: Debug id example: G-45463597 version: type: integer description: API version example: 2 status: type: string description: '"queued": the job has been queued. "processing": the job is being processed. "done": the job has been processed and the result is available. "error": an error occurred during the job''s processing.' enum: - queued - processing - done - error created_at: type: string description: Creation date format: date-time example: '2023-12-28T09:04:17.210Z' completed_at: type: string description: Completion date when status is "done" or "error" format: date-time example: '2023-12-28T09:04:37.210Z' nullable: true custom_metadata: type: object description: Custom metadata given in the initial request example: user: John Doe additionalProperties: true error_code: type: integer description: HTTP status code of the error if status is "error" minimum: 400 maximum: 599 example: 500 nullable: true post_session_metadata: type: object description: For debugging purposes, send data that could help to identify issues kind: type: string enum: - live example: live default: live file: description: The file data you uploaded. Can be null if status is "error" nullable: true allOf: - $ref: '#/components/schemas/FileResponse' request_params: description: Parameters used for this live transcription. Can be null if status is "error" nullable: true allOf: - $ref: '#/components/schemas/StreamingRequestParamsResponse' result: description: Live transcription's result when status is "done" nullable: true allOf: - $ref: '#/components/schemas/StreamingTranscriptionResultWithMessagesDTO' required: - id - request_id - version - status - created_at - post_session_metadata - kind ListTranscriptionResponse: type: object properties: first: type: string description: URL to fetch the first page format: uri example: https://api.gladia.io/v2/transcription?status=done&offset=0&limit=20 current: type: string description: URL to fetch the current page format: uri example: https://api.gladia.io/v2/transcription?status=done&offset=0&limit=20 next: type: string description: URL to fetch the next page format: uri example: https://api.gladia.io/v2/transcription?status=done&offset=20&limit=20 nullable: true items: description: List of transcriptions discriminator: propertyName: kind mapping: pre-recorded: '#/components/schemas/PreRecordedResponse' live: '#/components/schemas/StreamingResponse' type: array items: oneOf: - $ref: '#/components/schemas/PreRecordedResponse' - $ref: '#/components/schemas/StreamingResponse' required: - first - current - next - items ListHistoryResponse: type: object properties: first: type: string description: URL to fetch the first page format: uri example: https://api.gladia.io/v2/transcription?status=done&offset=0&limit=20 current: type: string description: URL to fetch the current page format: uri example: https://api.gladia.io/v2/transcription?status=done&offset=0&limit=20 next: type: string description: URL to fetch the next page format: uri example: https://api.gladia.io/v2/transcription?status=done&offset=20&limit=20 nullable: true items: description: List of jobs discriminator: propertyName: kind mapping: pre-recorded: '#/components/schemas/PreRecordedResponse' live: '#/components/schemas/StreamingResponse' type: array items: oneOf: - $ref: '#/components/schemas/PreRecordedResponse' - $ref: '#/components/schemas/StreamingResponse' required: - first - current - next - items AudioChunkActionData: type: object properties: chunk: type: string description: Chunk encoded in base64. The chunk must contains complete frames example: aGVsbG8= required: - chunk AudioChunkAction: type: object properties: type: type: string enum: - audio_chunk example: audio_chunk default: audio_chunk data: description: Payload of the audio chunk action allOf: - $ref: '#/components/schemas/AudioChunkActionData' required: - type - data StopRecordingAction: type: object properties: type: type: string enum: - stop_recording example: stop_recording default: stop_recording required: - type Error: type: object properties: message: type: string description: The error message required: - message AudioChunkAckData: type: object properties: byte_range: description: Range in bytes length of the audio chunk (relative to the whole session) minItems: 2 maxItems: 2 example: - 1024 - 2048 type: array items: type: integer time_range: description: Range in seconds of the audio chunk (relative to the whole session) minItems: 2 maxItems: 2 example: - 0.8 - 0.9 type: array items: type: number required: - byte_range - time_range AudioChunkAckMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' acknowledged: type: boolean description: Flag to indicate if the action was successfully acknowledged example: true error: description: Error message if the action was not successfully acknowledged nullable: true example: null allOf: - $ref: '#/components/schemas/Error' type: type: string enum: - audio_chunk example: audio_chunk default: audio_chunk data: description: The message data. "null" if the action was not successfully acknowledged nullable: true allOf: - $ref: '#/components/schemas/AudioChunkAckData' required: - session_id - created_at - acknowledged - error - type - data CallbackLiveAudioChunkAckMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.audio_chunk example: live.audio_chunk default: live.audio_chunk payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/AudioChunkAckMessage' required: - id - event - payload EndRecordingMessageData: type: object properties: recording_duration: type: number description: Total audio duration in seconds example: 344.45 required: - recording_duration EndRecordingMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' type: type: string enum: - end_recording example: end_recording default: end_recording data: description: The message data allOf: - $ref: '#/components/schemas/EndRecordingMessageData' required: - session_id - created_at - type - data CallbackLiveEndRecordingMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.end_recording example: live.end_recording default: live.end_recording payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/EndRecordingMessage' required: - id - event - payload EndSessionMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' type: type: string enum: - end_session example: end_session default: end_session required: - session_id - created_at - type CallbackLiveEndSessionMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.end_session example: live.end_session default: live.end_session payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/EndSessionMessage' required: - id - event - payload TranslationData: type: object properties: utterance_id: type: string description: Id of the utterance used for this result example: 00-00000011 utterance: description: The transcribed utterance allOf: - $ref: '#/components/schemas/UtteranceDTO' original_language: description: The original language in `iso639-1` or `iso639-2` format depending on the language allOf: - $ref: '#/components/schemas/TranscriptionLanguageCodeEnum' target_language: description: The target language in `iso639-1` or `iso639-2` format depending on the language allOf: - $ref: '#/components/schemas/TranslationLanguageCodeEnum' translated_utterance: description: The translated utterance allOf: - $ref: '#/components/schemas/UtteranceDTO' required: - utterance_id - utterance - original_language - target_language - translated_utterance TranslationMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' error: description: Error message if the addon failed nullable: true example: null allOf: - $ref: '#/components/schemas/Error' type: type: string enum: - translation example: translation default: translation data: description: The message data. "null" if the addon failed nullable: true allOf: - $ref: '#/components/schemas/TranslationData' required: - session_id - created_at - error - type - data CallbackLiveTranslationMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.translation example: live.translation default: live.translation payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/TranslationMessage' required: - id - event - payload NamedEntityRecognitionData: type: object properties: utterance_id: type: string description: Id of the utterance used for this result example: 00-00000011 utterance: description: The transcribed utterance allOf: - $ref: '#/components/schemas/UtteranceDTO' results: description: The NER results type: array items: $ref: '#/components/schemas/NamedEntityRecognitionResult' required: - utterance_id - utterance - results NamedEntityRecognitionMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' error: description: Error message if the addon failed nullable: true example: null allOf: - $ref: '#/components/schemas/Error' type: type: string enum: - named_entity_recognition example: named_entity_recognition default: named_entity_recognition data: description: The message data. "null" if the addon failed nullable: true allOf: - $ref: '#/components/schemas/NamedEntityRecognitionData' required: - session_id - created_at - error - type - data CallbackLiveNamedEntityRecognitionMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.named_entity_recognition example: live.named_entity_recognition default: live.named_entity_recognition payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/NamedEntityRecognitionMessage' required: - id - event - payload ChapterizationSentence: type: object properties: sentence: type: string start: type: number end: type: number words: type: array items: $ref: '#/components/schemas/WordDTO' required: - sentence - start - end - words PostChapterizationResult: type: object properties: abstractive_summary: type: string extractive_summary: type: string summary: type: string headline: type: string gist: type: string keywords: type: array items: type: string start: type: number end: type: number sentences: type: array items: $ref: '#/components/schemas/ChapterizationSentence' text: type: string required: - headline - gist - keywords - start - end - sentences - text PostChapterizationMessageData: type: object properties: results: description: The chapters type: array items: $ref: '#/components/schemas/PostChapterizationResult' required: - results PostChapterizationMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' error: description: Error message if the addon failed nullable: true example: null allOf: - $ref: '#/components/schemas/Error' type: type: string enum: - post_chapterization example: post_chapterization default: post_chapterization data: description: The message data. "null" if the addon failed nullable: true allOf: - $ref: '#/components/schemas/PostChapterizationMessageData' required: - session_id - created_at - error - type - data CallbackLivePostChapterizationMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.post_chapterization example: live.post_chapterization default: live.post_chapterization payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/PostChapterizationMessage' required: - id - event - payload StreamingTranscriptionResultDTO: type: object properties: metadata: description: Metadata for the given transcription & audio file allOf: - $ref: '#/components/schemas/TranscriptionMetadataDTO' transcription: description: Transcription of the audio speech allOf: - $ref: '#/components/schemas/TranscriptionDTO' translation: description: If `translation` has been enabled, translation of the audio speech transcription allOf: - $ref: '#/components/schemas/TranslationDTO' summarization: description: If `summarization` has been enabled, summarization of the audio speech transcription allOf: - $ref: '#/components/schemas/SummarizationDTO' named_entity_recognition: description: If `named_entity_recognition` has been enabled, the detected entities allOf: - $ref: '#/components/schemas/NamedEntityRecognitionDTO' sentiment_analysis: description: If `sentiment_analysis` has been enabled, sentiment analysis of the audio speech transcription allOf: - $ref: '#/components/schemas/SentimentAnalysisDTO' chapterization: description: If `chapterization` has been enabled, will generate chapters name for different parts of the given audio. allOf: - $ref: '#/components/schemas/ChapterizationDTO' required: - metadata PostFinalTranscriptMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' type: type: string enum: - post_final_transcript example: post_final_transcript default: post_final_transcript data: description: The message data allOf: - $ref: '#/components/schemas/StreamingTranscriptionResultDTO' required: - session_id - created_at - type - data CallbackLivePostFinalTranscriptMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.post_final_transcript example: live.post_final_transcript default: live.post_final_transcript payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/PostFinalTranscriptMessage' required: - id - event - payload PostSummarizationMessageData: type: object properties: results: type: string description: The summarization required: - results PostSummarizationMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' error: description: Error message if the addon failed nullable: true example: null allOf: - $ref: '#/components/schemas/Error' type: type: string enum: - post_summarization example: post_summarization default: post_summarization data: description: The message data. "null" if the addon failed nullable: true allOf: - $ref: '#/components/schemas/PostSummarizationMessageData' required: - session_id - created_at - error - type - data CallbackLivePostSummarizationMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.post_summarization example: live.post_summarization default: live.post_summarization payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/PostSummarizationMessage' required: - id - event - payload PostTranscriptMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' type: type: string enum: - post_transcript example: post_transcript default: post_transcript data: description: The message data allOf: - $ref: '#/components/schemas/TranscriptionDTO' required: - session_id - created_at - type - data CallbackLivePostTranscriptMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.post_transcript example: live.post_transcript default: live.post_transcript payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/PostTranscriptMessage' required: - id - event - payload SentimentAnalysisResult: type: object properties: sentiment: type: string emotion: type: string text: type: string start: type: number end: type: number channel: type: number required: - sentiment - emotion - text - start - end - channel SentimentAnalysisData: type: object properties: utterance_id: type: string description: Id of the utterance used for this result example: 00-00000011 utterance: description: The transcribed utterance allOf: - $ref: '#/components/schemas/UtteranceDTO' results: description: The sentiment analysis results type: array items: $ref: '#/components/schemas/SentimentAnalysisResult' required: - utterance_id - utterance - results SentimentAnalysisMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' error: description: Error message if the addon failed nullable: true example: null allOf: - $ref: '#/components/schemas/Error' type: type: string enum: - sentiment_analysis example: sentiment_analysis default: sentiment_analysis data: description: The message data. "null" if the addon failed nullable: true allOf: - $ref: '#/components/schemas/SentimentAnalysisData' required: - session_id - created_at - error - type - data CallbackLiveSentimentAnalysisMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.sentiment_analysis example: live.sentiment_analysis default: live.sentiment_analysis payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/SentimentAnalysisMessage' required: - id - event - payload StartRecordingMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' type: type: string enum: - start_recording example: start_recording default: start_recording required: - session_id - created_at - type CallbackLiveStartRecordingMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.start_recording example: live.start_recording default: live.start_recording payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/StartRecordingMessage' required: - id - event - payload StartSessionMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' type: type: string enum: - start_session example: start_session default: start_session required: - session_id - created_at - type CallbackLiveStartSessionMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.start_session example: live.start_session default: live.start_session payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/StartSessionMessage' required: - id - event - payload StopRecordingAckData: type: object properties: recording_duration: type: number description: Total audio duration in seconds example: 344.45 recording_left_to_process: type: number description: Audio duration left to process in seconds example: 11.23 required: - recording_duration - recording_left_to_process StopRecordingAckMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' acknowledged: type: boolean description: Flag to indicate if the action was successfully acknowledged example: true error: description: Error message if the action was not successfully acknowledged nullable: true example: null allOf: - $ref: '#/components/schemas/Error' type: type: string enum: - stop_recording example: stop_recording default: stop_recording data: description: The message data. "null" if the action was not successfully acknowledged nullable: true allOf: - $ref: '#/components/schemas/StopRecordingAckData' required: - session_id - created_at - acknowledged - error - type - data CallbackLiveStopRecordingAckMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.stop_recording example: live.stop_recording default: live.stop_recording payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/StopRecordingAckMessage' required: - id - event - payload TranscriptMessageData: type: object properties: id: type: string description: Id of the utterance example: 00-00000011 is_final: type: boolean description: Flag to indicate if the transcript is final or not example: true utterance: description: The transcribed utterance allOf: - $ref: '#/components/schemas/UtteranceDTO' required: - id - is_final - utterance TranscriptMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' type: type: string enum: - transcript example: transcript default: transcript data: description: The message data allOf: - $ref: '#/components/schemas/TranscriptMessageData' required: - session_id - created_at - type - data CallbackLiveTranscriptMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.transcript example: live.transcript default: live.transcript payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/TranscriptMessage' required: - id - event - payload SpeechMessageData: type: object properties: time: type: number description: Timestamp in seconds of the speech event example: 12.56 channel: type: number description: Channel of the speech event example: 1 required: - time - channel SpeechStartMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' type: type: string enum: - speech_start example: speech_start default: speech_start data: description: The message data allOf: - $ref: '#/components/schemas/SpeechMessageData' required: - session_id - created_at - type - data CallbackLiveSpeechStartMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.speech_start example: live.speech_start default: live.speech_start payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/SpeechStartMessage' required: - id - event - payload SpeechEndMessage: type: object properties: session_id: type: string description: Id of the live session example: 4a39145c-2844-4557-8f34-34883f7be7d9 created_at: type: string description: Date of creation of the message. The date is formatted as an ISO 8601 string example: '2021-09-01T12:00:00.123Z' type: type: string enum: - speech_end example: speech_end default: speech_end data: description: The message data allOf: - $ref: '#/components/schemas/SpeechMessageData' required: - session_id - created_at - type - data CallbackLiveSpeechEndMessage: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab event: type: string enum: - live.speech_end example: live.speech_end default: live.speech_end payload: description: The live message payload as sent to the WebSocket allOf: - $ref: '#/components/schemas/SpeechEndMessage' required: - id - event - payload StreamingRequest: type: object properties: encoding: description: "The encoding format of the audio stream. Supported formats: \n- PCM: 8, 16, 24, and 32 bits \n- A-law:\ \ 8 bits \n- μ-law: 8 bits \n\nNote: No need to add WAV headers to raw audio as the API supports both formats." default: wav/pcm allOf: - $ref: '#/components/schemas/StreamingSupportedEncodingEnum' bit_depth: description: The bit depth of the audio stream default: 16 allOf: - $ref: '#/components/schemas/StreamingSupportedBitDepthEnum' sample_rate: description: The sample rate of the audio stream default: 16000 allOf: - $ref: '#/components/schemas/StreamingSupportedSampleRateEnum' channels: type: integer description: The number of channels of the audio stream default: 1 minimum: 1 maximum: 8 custom_metadata: type: object description: Custom metadata you can attach to this live transcription example: user: John Doe additionalProperties: true model: description: The model used to process the audio. "solaria-1" is used by default. default: solaria-1 allOf: - $ref: '#/components/schemas/StreamingSupportedModels' endpointing: type: number description: The endpointing duration in seconds. Endpointing is the duration of silence which will cause an utterance to be considered as finished default: 0.05 minimum: 0.01 maximum: 10 maximum_duration_without_endpointing: type: number description: The maximum duration in seconds without endpointing. If endpointing is not detected after this duration, current utterance will be considered as finished default: 5 minimum: 5 maximum: 60 language_config: description: Specify the language configuration allOf: - $ref: '#/components/schemas/LanguageConfig' pre_processing: description: Specify the pre-processing configuration allOf: - $ref: '#/components/schemas/PreProcessingConfig' realtime_processing: description: Specify the realtime processing configuration allOf: - $ref: '#/components/schemas/RealtimeProcessingConfig' post_processing: description: Specify the post-processing configuration allOf: - $ref: '#/components/schemas/PostProcessingConfig' messages_config: description: Specify the websocket messages configuration allOf: - $ref: '#/components/schemas/MessagesConfig' callback: type: boolean description: If true, messages will be sent to configured url. default: false callback_config: description: Specify the callback configuration allOf: - $ref: '#/components/schemas/CallbackConfig' StreamingSupportedRegions: type: string enum: - us-west - eu-west InitStreamingResponse: type: object properties: id: type: string description: Id of the job format: uuid example: 45463597-20b7-4af7-b3b3-f5fb778203ab created_at: type: string description: Creation date format: date-time example: '2023-12-28T09:04:17.210Z' url: type: string description: The websocket url to connect to for sending audio data. The url will contain the temporary token to authenticate the session. example: wss://api.gladia.io/v2/live?token=4a39145c-2844-4557-8f34-34883f7be7d9 format: uri required: - id - created_at - url ListStreamingResponse: type: object properties: first: type: string description: URL to fetch the first page format: uri example: https://api.gladia.io/v2/transcription?status=done&offset=0&limit=20 current: type: string description: URL to fetch the current page format: uri example: https://api.gladia.io/v2/transcription?status=done&offset=0&limit=20 next: type: string description: URL to fetch the next page format: uri example: https://api.gladia.io/v2/transcription?status=done&offset=20&limit=20 nullable: true items: description: List of live transcriptions type: array items: $ref: '#/components/schemas/StreamingResponse' required: - first - current - next - items PatchRequestParamsDTO: type: object properties: {} PayloadTooLargeErrorResponse: type: object properties: timestamp: type: string description: Date of when the error occurred example: '2023-12-28T09:04:17.210Z' path: type: string description: Path to the API endpoint example: /v2/transcription/45463597-20b7-4af7-b3b3-f5fb778203ab request_id: type: string description: Debug id example: G-821fe9df statusCode: type: number description: HTTP status code of the error example: 413 message: type: string description: Payload too large example: payload too large required: - timestamp - path - request_id - statusCode - message WebhookTranscriptionCreatedPayload: type: object properties: event: type: string enum: - transcription.created default: transcription.created example: transcription.created payload: $ref: '#/components/schemas/PreRecordedEventPayload' required: - event - payload WebhookTranscriptionSuccessPayload: type: object properties: event: type: string enum: - transcription.success default: transcription.success example: transcription.success payload: $ref: '#/components/schemas/PreRecordedEventPayload' required: - event - payload WebhookTranscriptionErrorPayload: type: object properties: event: type: string enum: - transcription.error default: transcription.error example: transcription.error payload: $ref: '#/components/schemas/PreRecordedEventPayload' required: - event - payload WebhookLiveStartSessionPayload: type: object properties: event: type: string enum: - live.start_session default: live.start_session example: live.start_session payload: $ref: '#/components/schemas/LiveEventPayload' required: - event - payload WebhookLiveStartRecordingPayload: type: object properties: event: type: string enum: - live.start_recording default: live.start_recording example: live.start_recording payload: $ref: '#/components/schemas/LiveEventPayload' required: - event - payload WebhookLiveEndRecordingPayload: type: object properties: event: type: string enum: - live.end_recording default: live.end_recording example: live.end_recording payload: $ref: '#/components/schemas/LiveEventPayload' required: - event - payload WebhookLiveEndSessionPayload: type: object properties: event: type: string enum: - live.end_session default: live.end_session example: live.end_session payload: $ref: '#/components/schemas/LiveEventPayload' required: - event - payload