openapi: 3.2.0 info: title: AIML Video API version: 1.0.0 servers: - url: https://api.aimlapi.com tags: - name: Video paths: /v2/video/generations: post: operationId: _v2_video_generations requestBody: required: true content: application/json: schema: anyOf: - type: object properties: model: type: string enum: - sora-2-t2v - openai/sora-2-t2v - sora-2 - openai/sora-2 prompt: type: string description: The text description of the scene, subject, or action to generate in the video. resolution: type: string enum: - 720p default: 720p description: The resolution of the output video, where the number refers to the short side in pixels. aspect_ratio: type: string enum: - '16:9' - '9:16' default: '16:9' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 4 - 8 - 12 default: '4' required: - model - prompt title: sora-2-t2v, openai/sora-2-t2v, sora-2, openai/sora-2 - type: object properties: model: type: string enum: - sora-2-i2v - openai/sora-2-i2v prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: 'A URL or a Base64-encoded image file used as the initial frame for video generation. The image dimensions must match the selected video resolution and aspect ratio. Supported configurations include: 720p with aspect ratios: - 16:9 — 1280x720 - 9:16 — 720x1280 1080p with aspect ratios: - 16:9 — 1792x1024 - 9:16 — 1024x1792' resolution: type: string enum: - 720p default: 720p description: The resolution of the output video, where the number refers to the short side in pixels. aspect_ratio: type: string enum: - '16:9' - '9:16' default: '16:9' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 4 - 8 - 12 default: '4' required: - model - prompt - image_url title: sora-2-i2v, openai/sora-2-i2v - type: object properties: model: type: string enum: - sora-2-pro-t2v - openai/sora-2-pro-t2v - sora-2-pro - openai/sora-2-pro prompt: type: string description: The text description of the scene, subject, or action to generate in the video. resolution: type: string enum: - 720p - 1080p default: 1080p description: The resolution of the output video, where the number refers to the short side in pixels. aspect_ratio: type: string enum: - '16:9' - '9:16' default: '16:9' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 4 - 8 - 12 default: '4' required: - model - prompt title: sora-2-pro-t2v, openai/sora-2-pro-t2v, sora-2-pro, openai/sora-2-pro - type: object properties: model: type: string enum: - sora-2-pro-i2v - openai/sora-2-pro-i2v prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: 'A URL or a Base64-encoded image file used as the initial frame for video generation. The image dimensions must match the selected video resolution and aspect ratio. Supported configurations include: 720p with aspect ratios: - 16:9 — 1280x720 - 9:16 — 720x1280 1080p with aspect ratios: - 16:9 — 1792x1024 - 9:16 — 1024x1792' resolution: type: string enum: - 720p - 1080p default: 1080p description: The resolution of the output video, where the number refers to the short side in pixels. aspect_ratio: type: string enum: - '16:9' - '9:16' default: '16:9' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 4 - 8 - 12 default: '4' required: - model - prompt - image_url title: sora-2-pro-i2v, openai/sora-2-pro-i2v - type: object properties: model: type: string enum: - bytedance/seedance-1-0-pro - bytedance/seedance-1-0-pro-fast - bytedance/seedance-1-0-pro-t2v - bytedance/seedance-1-0-pro-i2v image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. prompt: type: string description: The text description of the scene, subject, or action to generate in the video. resolution: type: string enum: - 480p - 720p - 1080p default: 1080p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 watermark: type: boolean default: false description: Whether the video contains a watermark. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. camerafixed: type: boolean default: false description: 'Whether to fix the camera position. - true: Fix the camera position. The platform will append instructions to fix the camera position in the user''s prompt, but the actual effect is not guaranteed. - false: Do not fix the camera position.' last_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. required: - model - prompt title: bytedance/seedance-1-0-pro, bytedance/seedance-1-0-pro-fast, bytedance/seedance-1-0-pro-t2v, bytedance/seedance-1-0-pro-i2v - type: object properties: model: type: string enum: - bytedance/seedance-1-5-pro prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. last_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. aspect_ratio: type: string enum: - '16:9' - '4:3' - '1:1' - '3:4' - '9:16' - '21:9' - adaptive default: '16:9' description: The aspect ratio of the generated video. resolution: type: string enum: - 480p - 720p - 1080p default: 720p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The length of the output video in seconds. enum: - 4 - 5 - 6 - 7 - 8 - 9 - 10 - 11 - 12 seed: type: integer default: -1 description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. camera_fixed: type: boolean default: false description: 'Whether to fix the camera position. - true: Fix the camera position. The platform will append instructions to fix the camera position in the user''s prompt, but the actual effect is not guaranteed. - false: Do not fix the camera position.' watermark: type: boolean default: false description: Whether the video contains a watermark. generate_audio: type: boolean default: true description: Whether to generate audio for the video. required: - model - prompt title: bytedance/seedance-1-5-pro - type: object properties: model: type: string enum: - bytedance/dreamina-seedance-2-0 - bytedance/seedance-2-0 - bytedance/seedance-2.0/image-to-video - bytedance/seedance-2.0/reference-to-video - bytedance/seedance-2.0/text-to-video provider: type: string description: Provider routing override, pinning the request to one provider and disabling fallback. `auto` (default) uses the full Tencent Cloud VOD -> native ByteDance -> fal.ai chain. `tencent_vod` is the only provider that accepts reference images containing real people (it registers them as assets first); pinning `bytedance` or `fal` sends the images straight to a provider that rejects them with "the input image may contain real person". Case-insensitive. example: auto prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. last_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. audio_url: type: string format: uri description: URL, Base64-encoded string or asset ID of the audio. The duration of the audio file specified in the parameter must not exceed 15.2 seconds. If audio is provided, at least one reference image or video is required. video_url: type: string format: uri description: The public URL of the video. Only video URLs are supported. image_urls: type: array items: type: string format: uri minItems: 1 maxItems: 9 description: Reference images to guide video generation. Refer to them in the prompt as @Image1, @Image2, etc. audio_urls: type: array items: type: string format: uri minItems: 1 maxItems: 3 description: 'Reference audio to guide video generation. Refer to them in the prompt as @Audio1, @Audio2, etc. Supported formats: MP3, WAV. Up to 3 files, combined duration must not exceed 15 seconds. If audio is provided, at least one reference image (image_urls) or video (video_urls) is required.' video_urls: type: array items: type: string format: uri minItems: 1 maxItems: 3 description: 'Reference videos to guide video generation. Refer to them in the prompt as @Video1, @Video2, etc. Supported formats: MP4, MOV. Each video must be between ~480p (640x640) and ~720p (834x1112) in resolution.' aspect_ratio: type: string enum: - '21:9' - '16:9' - '4:3' - '1:1' - '3:4' - '9:16' default: '16:9' description: 'The aspect ratio of the generated video. Defaults to 16:9 for text-to-video and reference-to-video requests. Must be omitted for first-frame / first-last-frame generation (`image_url`): the output aspect ratio follows the input image.' resolution: type: string enum: - 480p - 720p - 1080p - 4k default: 720p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The length of the output video in seconds. enum: - 4 - 5 - 6 - 7 - 8 - 9 - 10 - 11 - 12 - 13 - 14 - 15 default: '5' generate_audio: type: boolean default: true description: Whether to generate audio for the video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. watermark: type: boolean default: false description: Whether the video contains a watermark. required: - model - prompt title: bytedance/dreamina-seedance-2-0, bytedance/seedance-2-0, bytedance/seedance-2.0/image-to-video, bytedance/seedance-2.0/reference-to-video, bytedance/seedance-2.0/text-to-video - type: object properties: model: type: string enum: - bytedance/dreamina-seedance-2-0-fast - bytedance/dreamina-seedance-2-0-mini - bytedance/seedance-2-0-fast - bytedance/seedance-2.0/fast/image-to-video - bytedance/seedance-2.0/fast/reference-to-video - bytedance/seedance-2.0/fast/text-to-video - bytedance/seedance-2-0-mini provider: type: string description: Provider routing override, pinning the request to one provider and disabling fallback. `auto` (default) uses the full Tencent Cloud VOD -> native ByteDance -> fal.ai chain. `tencent_vod` is the only provider that accepts reference images containing real people (it registers them as assets first); pinning `bytedance` or `fal` sends the images straight to a provider that rejects them with "the input image may contain real person". Case-insensitive. example: auto prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. last_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. audio_url: type: string format: uri description: URL, Base64-encoded string or asset ID of the audio. The duration of the audio file specified in the parameter must not exceed 15.2 seconds. If audio is provided, at least one reference image or video is required. video_url: type: string format: uri description: The public URL of the video. Only video URLs are supported. image_urls: type: array items: type: string format: uri minItems: 1 maxItems: 9 description: Reference images to guide video generation. Refer to them in the prompt as @Image1, @Image2, etc. audio_urls: type: array items: type: string format: uri minItems: 1 maxItems: 3 description: 'Reference audio to guide video generation. Refer to them in the prompt as @Audio1, @Audio2, etc. Supported formats: MP3, WAV. Up to 3 files, combined duration must not exceed 15 seconds. If audio is provided, at least one reference image (image_urls) or video (video_urls) is required.' video_urls: type: array items: type: string format: uri minItems: 1 maxItems: 3 description: 'Reference videos to guide video generation. Refer to them in the prompt as @Video1, @Video2, etc. Supported formats: MP4, MOV. Each video must be between ~480p (640x640) and ~720p (834x1112) in resolution.' aspect_ratio: type: string enum: - '21:9' - '16:9' - '4:3' - '1:1' - '3:4' - '9:16' default: '16:9' description: 'The aspect ratio of the generated video. Defaults to 16:9 for text-to-video and reference-to-video requests. Must be omitted for first-frame / first-last-frame generation (`image_url`): the output aspect ratio follows the input image.' resolution: type: string enum: - 480p - 720p default: 720p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The length of the output video in seconds. enum: - 4 - 5 - 6 - 7 - 8 - 9 - 10 - 11 - 12 - 13 - 14 - 15 default: '5' generate_audio: type: boolean default: true description: Whether to generate audio for the video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. watermark: type: boolean default: false description: Whether the video contains a watermark. required: - model - prompt title: bytedance/dreamina-seedance-2-0-fast, bytedance/dreamina-seedance-2-0-mini, bytedance/seedance-2-0-fast, bytedance/seedance-2.0/fast/image-to-video, bytedance/seedance-2.0/fast/reference-to-video, bytedance/seedance-2.0/fast/text-to-video, bytedance/seedance-2-0-mini - type: object properties: model: type: string enum: - bytedance/dreamina-seedance-2-5 - bytedance/seedance-2-5 - bytedance/seedance-2.5 provider: type: string description: Provider routing override, pinning the request to one provider and disabling fallback. Seedance 2.5 runs on native ByteDance only, so `auto` (default) and `bytedance` behave identically today. Unlike Seedance 2.0 there is no Tencent Cloud VOD link, so reference images containing real people are rejected upstream. Case-insensitive. example: auto prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. last_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. audio_url: type: string format: uri description: URL, Base64-encoded string or asset ID of the audio. The duration of the audio file specified in the parameter must not exceed 15.2 seconds. If audio is provided, at least one reference image or video is required. video_url: type: string format: uri description: The public URL of the video. Only video URLs are supported. image_urls: type: array items: type: string format: uri minItems: 1 maxItems: 9 description: Reference images to guide video generation. Refer to them in the prompt as @Image1, @Image2, etc. audio_urls: type: array items: type: string format: uri minItems: 1 maxItems: 3 description: 'Reference audio to guide video generation. Refer to them in the prompt as @Audio1, @Audio2, etc. Supported formats: MP3, WAV. Up to 3 files, combined duration must not exceed 15 seconds. If audio is provided, at least one reference image (image_urls) or video (video_urls) is required.' video_urls: type: array items: type: string format: uri minItems: 1 maxItems: 3 description: 'Reference videos to guide video generation. Refer to them in the prompt as @Video1, @Video2, etc. Supported formats: MP4, MOV. Each video must be between ~480p (640x640) and ~720p (834x1112) in resolution.' aspect_ratio: type: string enum: - '21:9' - '16:9' - '4:3' - '1:1' - '3:4' - '9:16' default: '16:9' description: 'The aspect ratio of the generated video. Defaults to 16:9 for text-to-video and reference-to-video requests. Must be omitted for first-frame / first-last-frame generation (`image_url`): the output aspect ratio follows the input image.' resolution: type: string enum: - 480p - 720p default: 720p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The length of the output video in seconds. enum: - 4 - 5 - 6 - 7 - 8 - 9 - 10 - 11 - 12 - 13 - 14 - 15 - 16 - 17 - 18 - 19 - 20 - 21 - 22 - 23 - 24 - 25 - 26 - 27 - 28 - 29 - 30 default: '5' generate_audio: type: boolean default: true description: Whether to generate audio for the video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. watermark: type: boolean default: false description: Whether the video contains a watermark. required: - model - prompt title: bytedance/dreamina-seedance-2-5, bytedance/seedance-2-5, bytedance/seedance-2.5 - type: object properties: model: type: string enum: - veo-2.0-generate-001 - google/veo-2.0-generate-001 - veo2 - google/veo2 prompt: type: string description: The text description of the scene, subject, or action to generate in the video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 6 - 7 - 8 default: '5' aspect_ratio: type: string enum: - '16:9' - '9:16' description: The aspect ratio of the generated video. negative_prompt: type: string description: The description of elements to avoid in the generated video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. enhance_prompt: type: boolean default: true description: Whether to enhance the video generation. required: - model - prompt title: veo-2.0-generate-001, google/veo-2.0-generate-001, veo2, google/veo2 - type: object properties: model: type: string enum: - veo-3.0-fast-generate-001 - google/veo-3.0-fast-generate-001 - veo-3.0-generate-001 - google/veo-3.0-generate-001 - google/veo3 - google/veo-3.0-fast prompt: type: string description: The text description of the scene, subject, or action to generate in the video. duration: type: integer description: The length of the output video in seconds. enum: - 4 - 6 - 8 default: '8' aspect_ratio: type: string enum: - '16:9' - '9:16' description: The aspect ratio of the generated video. resolution: type: string enum: - 720P - 1080P default: 720P description: The resolution of the output video, where the number refers to the short side in pixels. negative_prompt: type: string description: The description of elements to avoid in the generated video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. enhance_prompt: type: boolean default: true description: Whether to enhance the video generation. generate_audio: type: boolean default: true description: Whether to generate audio for the video. required: - model - prompt title: veo-3.0-fast-generate-001, google/veo-3.0-fast-generate-001, veo-3.0-generate-001, google/veo-3.0-generate-001, google/veo3, google/veo-3.0-fast - type: object properties: model: type: string enum: - veo-3.1-lite-generate-001 - google/veo-3.1-lite-generate-001 - google/veo-3-1-lite-generate-preview provider: type: string description: Provider routing override. `google` runs native Google with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Google -> fal.ai fallback chain. Case-insensitive. example: auto prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the first frame of the video. Should be 720p or higher resolution. last_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. Should be 720p or higher resolution. aspect_ratio: type: string enum: - '16:9' - '9:16' default: '16:9' description: The aspect ratio of the generated video. resolution: type: string enum: - 720P - 1080P default: 720P description: The resolution of the output video, where the number refers to the short side in pixels. duration: type: integer description: The length of the output video in seconds. enum: - 4 - 6 - 8 default: '8' generate_audio: type: boolean default: true description: Whether to generate audio for the video. person_generation: type: string enum: - dont_allow - allow_adult default: allow_adult description: Allow generation of people. required: - model - prompt title: veo-3.1-lite-generate-001, google/veo-3.1-lite-generate-001, google/veo-3-1-lite-generate-preview - type: object properties: model: type: string enum: - veo-3.1-generate-001 - google/veo-3.1-generate-001 - veo-3.1-fast-generate-001 - google/veo-3.1-fast-generate-001 - google/veo-3.1-t2v - google/veo-3.1-i2v - google/veo-3.1-t2v-fast - google/veo-3.1-i2v-fast - google/veo-3.1-first-last-image-to-video - google/veo-3.1-first-last-image-to-video-fast - google/veo3-1-extend-video - google/veo3-1-fast-extend-video provider: type: string description: Provider routing override. `google` runs native Google with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Google -> fal.ai fallback chain. Case-insensitive. example: auto prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the first frame of the video. Should be 720p or higher resolution. last_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. Should be 720p or higher resolution. video_url: type: string format: uri description: URL of a Veo-generated MP4 to extend. Extension output is always 7 seconds; `duration` is forced to 7 when this field is set. aspect_ratio: type: string enum: - '16:9' - '9:16' default: '16:9' description: The aspect ratio of the generated video. resolution: type: string enum: - 720P - 1080P - 4K default: 720P description: The resolution of the output video, where the number refers to the short side in pixels. duration: type: integer description: The length of the output video in seconds. When `video_url` is set (video extension), Google always returns a fixed 7-second extension; `duration` is forced to 7 and other values are ignored. enum: - 4 - 6 - 8 default: '8' generate_audio: type: boolean default: true description: Whether to generate audio for the video. person_generation: type: string enum: - dont_allow - allow_adult default: allow_adult description: Allow generation of people. required: - model - prompt title: veo-3.1-generate-001, google/veo-3.1-generate-001, veo-3.1-fast-generate-001, google/veo-3.1-fast-generate-001, google/veo-3.1-t2v, google/veo-3.1-i2v, google/veo-3.1-t2v-fast, google/veo-3.1-i2v-fast, google/veo-3.1-first-last-image-to-video, google/veo-3.1-first-last-image-to-video-fast, google/veo3-1-extend-video, google/veo3-1-fast-extend-video - type: object properties: model: type: string enum: - gemini-omni-flash-preview - google/gemini-omni-flash-preview prompt: type: string minLength: 1 description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: Image URL, gs:// URI, or base64 data URI. Used as the first frame for image_to_video unless the prompt tags it as a reference. image_urls: type: array items: type: string format: uri maxItems: 6 description: Reference image URLs, gs:// URIs, or base64 data URIs. Use prompt tags such as to bind image roles. video_url: type: string format: uri description: Video URL, gs:// URI, or base64 data URI for edit workflows. audio_url: type: string format: uri description: Audio URL, gs:// URI, or base64 data URI for multimodal prompting. previous_interaction_id: type: string description: Gemini Interactions API id from a previous Omni generation, used for conversational edits. task: type: string enum: - text_to_video - image_to_video - reference_to_video - edit description: Gemini Omni video task. Defaults are inferred from the supplied media inputs. aspect_ratio: type: string enum: - '16:9' - '9:16' description: The aspect ratio of the generated video. delivery: type: string enum: - uri default: uri description: Gemini Omni video response delivery mode. required: - model - prompt title: gemini-omni-flash-preview, google/gemini-omni-flash-preview - type: object properties: model: type: string enum: - gemini-omni-1.1-flash - google/gemini-omni-1.1-flash prompt: type: string minLength: 1 description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: Image URL, gs:// URI, or base64 data URI. Used as the first frame for image_to_video unless the prompt tags it as a reference. image_urls: type: array items: type: string format: uri maxItems: 6 description: Reference image URLs, gs:// URIs, or base64 data URIs. Use prompt tags such as to bind image roles. Pass exactly two images with task=image_to_video for first/last frame interpolation. video_url: type: string format: uri description: Video URL, gs:// URI, or base64 data URI for edit and extend workflows. audio_url: type: string format: uri description: Audio URL, gs:// URI, or base64 data URI for multimodal prompting. previous_interaction_id: type: string description: Gemini Interactions API id from a previous Omni generation, used for conversational edits. task: type: string enum: - text_to_video - image_to_video - reference_to_video - edit - extend description: Gemini Omni video task. Defaults are inferred from the supplied media inputs. aspect_ratio: type: string enum: - '16:9' - '9:16' description: The aspect ratio of the generated video. resolution: type: string enum: - 360p - 720p - 1080p - 4k description: The resolution of the output video, where the number refers to the short side in pixels. duration: type: integer description: The length of the output video in seconds. enum: - 3 - 4 - 5 - 6 - 7 - 8 - 9 - 10 default: '8' delivery: type: string enum: - uri default: uri description: Gemini Omni video response delivery mode. required: - model - prompt title: gemini-omni-1.1-flash, google/gemini-omni-1.1-flash - type: object properties: model: type: string enum: - google/veo-3.1-reference-to-video - veo3.1/reference-to-video prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_urls: type: array items: type: string format: uri minItems: 1 maxItems: 3 description: URL of the input image to animate. Should be 720p or higher resolution. duration: type: integer description: The length of the output video in seconds. enum: - 8 provider: type: string description: Provider routing override. `google` runs native Google with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Google -> fal.ai fallback chain. Case-insensitive. example: auto resolution: type: string enum: - 720p - 1080p default: 1080p description: The resolution of the output video, where the number refers to the short side in pixels. generate_audio: type: boolean default: true description: Whether to generate audio for the video. required: - model - prompt - image_urls title: google/veo-3.1-reference-to-video, veo3.1/reference-to-video - type: object properties: model: type: string enum: - veo2/image-to-video - google/veo2-image-to-video prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. tail_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 6 - 7 - 8 default: '5' aspect_ratio: type: string enum: - '16:9' - '9:16' description: The aspect ratio of the generated video. negative_prompt: type: string description: The description of elements to avoid in the generated video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. enhance_prompt: type: boolean default: true description: Whether to enhance the video generation. required: - model - prompt - image_url title: veo2/image-to-video, google/veo2-image-to-video - type: object properties: model: type: string enum: - google/veo-3.0-i2v - google/veo-3.0-i2v-fast prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. duration: type: integer description: The length of the output video in seconds. enum: - 4 - 6 - 8 default: '8' aspect_ratio: type: string enum: - '16:9' - '9:16' description: The aspect ratio of the generated video. resolution: type: string enum: - 720P - 1080P default: 720P description: The resolution of the output video, where the number refers to the short side in pixels. negative_prompt: type: string description: The description of elements to avoid in the generated video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. enhance_prompt: type: boolean default: true description: Whether to enhance the video generation. generate_audio: type: boolean default: true description: Whether to generate audio for the video. required: - model - prompt - image_url title: google/veo-3.0-i2v, google/veo-3.0-i2v-fast - type: object properties: model: type: string enum: - wan2.1-t2v-plus - alibaba/wan2.1-t2v-plus prompt: type: string description: The text description of the scene, subject, or action to generate in the video. resolution: type: string enum: - 720P default: 720P description: An enumeration where the short side of the video frame determines the resolution. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' - '4:3' - '3:4' default: '16:9' description: The aspect ratio of the generated video. negative_prompt: type: string description: The description of elements to avoid in the generated video. watermark: type: boolean default: false description: Whether the video contains a watermark. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. enhance_prompt: type: boolean default: true description: Whether to enable prompt expansion. required: - model - prompt title: wan2.1-t2v-plus, alibaba/wan2.1-t2v-plus - type: object properties: model: type: string enum: - wan2.1-t2v-turbo - alibaba/wan2.1-t2v-turbo prompt: type: string description: The text description of the scene, subject, or action to generate in the video. resolution: type: string enum: - 480P - 720P default: 720P description: An enumeration where the short side of the video frame determines the resolution. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' - '4:3' - '3:4' default: '16:9' description: The aspect ratio of the generated video. negative_prompt: type: string description: The description of elements to avoid in the generated video. watermark: type: boolean default: false description: Whether the video contains a watermark. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. enhance_prompt: type: boolean default: true description: Whether to enable prompt expansion. required: - model - prompt title: wan2.1-t2v-turbo, alibaba/wan2.1-t2v-turbo - type: object properties: model: type: string enum: - wan2.2-i2v-plus - alibaba/wan2.2-i2v-plus prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. resolution: type: string enum: - 480P - 1080P default: 1080P description: An enumeration where the short side of the video frame determines the resolution. negative_prompt: type: string description: The description of elements to avoid in the generated video. watermark: type: boolean default: false description: Whether the video contains a watermark. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. enhance_prompt: type: boolean default: true description: Whether to enable prompt expansion. required: - model - prompt - image_url title: wan2.2-i2v-plus, alibaba/wan2.2-i2v-plus - type: object properties: model: type: string enum: - wan2.2-t2v-plus - alibaba/wan2.2-t2v-plus prompt: type: string description: The text description of the scene, subject, or action to generate in the video. resolution: type: string enum: - 480P - 1080P default: 1080P description: An enumeration where the short side of the video frame determines the resolution. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' - '4:3' - '3:4' default: '16:9' description: The aspect ratio of the generated video. negative_prompt: type: string description: The description of elements to avoid in the generated video. watermark: type: boolean default: false description: Whether the video contains a watermark. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. enhance_prompt: type: boolean default: true description: Whether to enable prompt expansion. required: - model - prompt title: wan2.2-t2v-plus, alibaba/wan2.2-t2v-plus - type: object properties: model: type: string enum: - wan2.5-t2v-preview - alibaba/wan2.5-t2v-preview prompt: type: string description: The text description of the scene, subject, or action to generate in the video. resolution: type: string enum: - 480p - 720p - 1080p default: 1080p description: An enumeration where the short side of the video frame determines the resolution. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' default: '16:9' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '10' negative_prompt: type: string description: The description of elements to avoid in the generated video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. enhance_prompt: type: boolean default: true description: Whether to enable prompt expansion. required: - model - prompt title: wan2.5-t2v-preview, alibaba/wan2.5-t2v-preview - type: object properties: model: type: string enum: - wan2.5-i2v-preview - alibaba/wan2.5-i2v-preview prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. resolution: type: string enum: - 480p - 720p - 1080p default: 1080p description: An enumeration where the short side of the video frame determines the resolution. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' default: '16:9' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '10' negative_prompt: type: string description: The description of elements to avoid in the generated video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. enhance_prompt: type: boolean default: true description: Whether to enable prompt expansion. required: - model - prompt - image_url title: wan2.5-i2v-preview, alibaba/wan2.5-i2v-preview - type: object properties: model: type: string enum: - wan2.6-t2v - alibaba/wan2.6-t2v - wan2.7-t2v - alibaba/wan2.7-t2v - alibaba/wan-2-6-t2v - alibaba/wan-2-7-t2v prompt: type: string description: The text description of the scene, subject, or action to generate in the video. audio_url: type: string format: uri description: The URL of the audio file. The model will use this audio to generate the video. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' - '4:3' - '3:4' default: '16:9' description: The aspect ratio of the generated video. resolution: type: string enum: - 720p - 1080p default: 1080p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 - 15 default: '10' negative_prompt: type: string description: The description of elements to avoid in the generated video. shot_type: type: string enum: - single - multi default: single description: 'Specifies the shot type of the generated video, that is, whether the video consists of a single continuous shot or multiple switched shots. This parameter takes effect only when "prompt_extend" is set to ''true'': - single: (default) Outputs a single-shot video. - multi: Outputs a multi-shot video.' generate_audio: type: boolean default: true description: 'Specifies whether to automatically add audio to the generated video. This parameter takes effect only when ''audio_url'' is not provided.' seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. enhance_prompt: type: boolean default: true description: Whether to enable prompt expansion. required: - model - prompt title: wan2.6-t2v, alibaba/wan2.6-t2v, wan2.7-t2v, alibaba/wan2.7-t2v, alibaba/wan-2-6-t2v, alibaba/wan-2-7-t2v - type: object properties: model: type: string enum: - wan2.6-i2v - alibaba/wan2.6-i2v - alibaba/wan-2-6-i2v prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. audio_url: type: string format: uri description: The URL of the audio file. The model will use this audio to generate the video. resolution: type: string enum: - 720p - 1080p default: 1080p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 - 15 default: '10' negative_prompt: type: string description: The description of elements to avoid in the generated video. shot_type: type: string enum: - single - multi default: single description: 'Specifies the shot type of the generated video, that is, whether the video consists of a single continuous shot or multiple switched shots. This parameter takes effect only when "prompt_extend" is set to ''true'': - single: (default) Outputs a single-shot video. - multi: Outputs a multi-shot video.' generate_audio: type: boolean default: true description: 'Specifies whether to automatically add audio to the generated video. This parameter takes effect only when ''audio_url'' is not provided.' seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. enhance_prompt: type: boolean default: true description: Whether to enable prompt expansion. required: - model - prompt - image_url title: wan2.6-i2v, alibaba/wan2.6-i2v, alibaba/wan-2-6-i2v - type: object properties: model: type: string enum: - wan2.6-i2v-flash - alibaba/wan2.6-i2v-flash - alibaba/wan-2-6-image-to-video-flash prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. audio_url: type: string format: uri description: The URL of the audio file. The model will use this audio to generate the video. resolution: type: string enum: - 720p - 1080p default: 720p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: Duration of the generated video in seconds (up to 15 seconds for Flash model). enum: - 5 - 10 - 15 default: '10' negative_prompt: type: string description: The description of elements to avoid in the generated video. shot_type: type: string enum: - single - multi default: single description: 'Specifies the shot type of the generated video, that is, whether the video consists of a single continuous shot or multiple switched shots. This parameter takes effect only when "prompt_extend" is set to ''true'': - single: (default) Outputs a single-shot video. - multi: Outputs a multi-shot video.' generate_audio: type: boolean default: true description: 'Specifies whether to automatically add audio to the generated video. This parameter takes effect only when ''audio_url'' is not provided.' seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. enhance_prompt: type: boolean default: true description: Whether to enable prompt expansion. required: - model - prompt - image_url title: wan2.6-i2v-flash, alibaba/wan2.6-i2v-flash, alibaba/wan-2-6-image-to-video-flash - type: object properties: model: type: string enum: - wan2.6-r2v - alibaba/wan2.6-r2v - alibaba/wan-2-6-r2v prompt: type: string description: The text description of the scene, subject, or action to generate in the video. video_urls: type: array items: type: string format: uri minItems: 1 maxItems: 3 description: 'An array of URLs for the uploaded reference video files. This parameter is used to extract the character''s appearance and voice (if any) to generate a video that matches the reference features. Each reference video must contain only one character. For example, character1 is a little girl and character2 is an alarm clock.' aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' - '4:3' - '3:4' default: '16:9' description: The aspect ratio of the generated video. resolution: type: string enum: - 720p - 1080p default: 1080p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 - 15 default: '10' negative_prompt: type: string description: The description of elements to avoid in the generated video. shot_type: type: string enum: - single - multi default: single description: 'Specifies the shot type of the generated video, that is, whether the video consists of a single continuous shot or multiple switched shots. This parameter takes effect only when "prompt_extend" is set to ''true'': - single: (default) Outputs a single-shot video. - multi: Outputs a multi-shot video.' seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. enhance_prompt: type: boolean default: true description: Whether to enable prompt expansion. required: - model - prompt - video_urls title: wan2.6-r2v, alibaba/wan2.6-r2v, alibaba/wan-2-6-r2v - type: object properties: model: type: string enum: - wan2.7-i2v - alibaba/wan2.7-i2v - alibaba/wan-2-7-i2v prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. audio_url: type: string format: uri description: The URL of the audio file. The model will use this audio to generate the video. resolution: type: string enum: - 720p - 1080p default: 1080p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 - 15 default: '10' enable_prompt_expansion: type: boolean default: true description: Whether to enable prompt expansion. enhance_prompt: type: boolean default: true description: Whether to enable prompt expansion. negative_prompt: type: string maxLength: 500 description: The description of elements to avoid in the generated video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. watermark: type: boolean default: false description: Whether the video contains a watermark. required: - model - prompt - image_url title: wan2.7-i2v, alibaba/wan2.7-i2v, alibaba/wan-2-7-i2v - type: object properties: model: type: string enum: - wan2.7-r2v - alibaba/wan2.7-r2v - alibaba/wan-2-7-r2v prompt: type: string description: The text description of the scene, subject, or action to generate in the video. first_frame_image: type: string format: uri description: Optional first frame image URL. image_urls: type: array items: type: string format: uri maxItems: 5 description: Reference image URLs. video_urls: type: array items: type: string format: uri maxItems: 5 description: Reference video URLs. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' - '4:3' - '3:4' default: '16:9' description: The aspect ratio of the generated video. resolution: type: string enum: - 720p - 1080p default: 1080p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 - 15 default: '10' enhance_prompt: type: boolean default: true description: Whether to enable prompt expansion. negative_prompt: type: string maxLength: 500 description: The description of elements to avoid in the generated video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. watermark: type: boolean default: false description: Whether the video contains a watermark. required: - model - prompt title: wan2.7-r2v, alibaba/wan2.7-r2v, alibaba/wan-2-7-r2v - type: object properties: model: type: string enum: - wan3.0-video - alibaba/wan3.0-video - alibaba/wan-3-0-video prompt: type: string maxLength: 20000 description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: The URL of the image strictly used as the first frame of the video. Cannot be used together with reference media, file_url, or link_url. last_image_url: type: string format: uri description: The URL of the image strictly used as the last frame of the video. Requires image_url. reference_image_urls: type: array items: type: string format: uri maxItems: 10 description: Array of image URLs for multi-image-to-video generation. video_urls: type: array items: type: string format: uri maxItems: 5 description: Reference video URLs. Up to 5 clips with a total duration of no more than 15 seconds. audio_urls: type: array items: type: string format: uri maxItems: 5 description: Reference audio URLs. Up to 5 clips with a total duration of no more than 15 seconds. file_url: type: string format: uri description: The URL of a document (docx, doc, xlsx, xls, pptx, ppt, pdf, txt, key, pages, numbers, md) the model uses to generate the video. Cannot be used together with link_url. link_url: type: string format: uri description: The URL of a publicly accessible web page the model uses to generate the video. Cannot be used together with file_url. aspect_ratio: type: string enum: - adaptive - '16:9' - '4:3' - '1:1' - '3:4' - '9:16' default: adaptive description: The aspect ratio of the generated video. resolution: type: string enum: - 480p - 720p - 1080p default: 1080p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The length of the output video in seconds, from 2 to 30. When reference videos are provided, the total input video duration plus the output video duration must not exceed 30 seconds. default: 5 generate_audio: type: boolean default: true description: Specifies whether the output video contains audio. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. watermark: type: boolean default: false description: Whether the video contains a watermark. required: - model - prompt title: wan3.0-video, alibaba/wan3.0-video, alibaba/wan-3-0-video - type: object properties: model: type: string enum: - custom:happyhorse-1.0 - alibaba/custom:happyhorse-1.0 - happyhorse-1.0-t2v - alibaba/happyhorse-1.0-t2v - happyhorse-1.0-i2v - alibaba/happyhorse-1.0-i2v - happyhorse-1.0-r2v - alibaba/happyhorse-1.0-r2v - happyhorse-1.0-video-edit - alibaba/happyhorse-1.0-video-edit - alibaba/happyhorse-1-0 prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. reference_image_urls: type: array items: type: string format: uri maxItems: 9 description: Array of image URLs for multi-image-to-video generation. video_url: type: string format: uri description: A HTTPS URL pointing to a video or a data URI containing a video. This video will be used as a reference during generation. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' - '4:3' - '3:4' default: '16:9' description: The aspect ratio of the generated video. resolution: type: string enum: - 720p - 1080p default: 1080p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: "The length of the output video in seconds. \nThe option does not work if a video reference is provided. \n- If the input video is 15 seconds or shorter, the output video has the same duration as the input.\n- If the input video is longer than 15 seconds, the system automatically uses only the first 15 seconds, so the maximum output duration is 15 seconds." enum: - 3 - 4 - 5 - 6 - 7 - 8 - 9 - 10 - 11 - 12 - 13 - 14 - 15 default: '10' audio_setting: type: string enum: - auto - origin default: auto description: 'Audio control. Works only if a video reference is provided. - auto (default): Determined by the model. - origin: Preserves the original audio from the input video.' seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. watermark: type: boolean default: false description: Whether the video contains a watermark. required: - model - prompt title: custom:happyhorse-1.0, alibaba/custom:happyhorse-1.0, happyhorse-1.0-t2v, alibaba/happyhorse-1.0-t2v, happyhorse-1.0-i2v, alibaba/happyhorse-1.0-i2v, happyhorse-1.0-r2v, alibaba/happyhorse-1.0-r2v, happyhorse-1.0-video-edit, alibaba/happyhorse-1.0-video-edit, alibaba/happyhorse-1-0 - type: object properties: model: type: string enum: - happyhorse-1.1-t2v - alibaba/happyhorse-1.1-t2v - alibaba/happyhorse-1-1 - alibaba/happyhorse-1-1-t2v prompt: type: string description: The text description of the scene, subject, or action to generate in the video. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' - '4:3' - '3:4' default: '16:9' description: The aspect ratio of the generated video. resolution: type: string enum: - 720p - 1080p default: 1080p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The length of the output video in seconds. enum: - 3 - 4 - 5 - 6 - 7 - 8 - 9 - 10 - 11 - 12 - 13 - 14 - 15 default: '10' seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. watermark: type: boolean default: false description: Whether the video contains a watermark. required: - model - prompt title: happyhorse-1.1-t2v, alibaba/happyhorse-1.1-t2v, alibaba/happyhorse-1-1, alibaba/happyhorse-1-1-t2v - type: object properties: model: type: string enum: - happyhorse-1.1-i2v - alibaba/happyhorse-1.1-i2v - alibaba/happyhorse-1-1-i2v prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' - '4:3' - '3:4' default: '16:9' description: The aspect ratio of the generated video. resolution: type: string enum: - 720p - 1080p default: 1080p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The length of the output video in seconds. enum: - 3 - 4 - 5 - 6 - 7 - 8 - 9 - 10 - 11 - 12 - 13 - 14 - 15 default: '10' seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. watermark: type: boolean default: false description: Whether the video contains a watermark. required: - model - prompt - image_url title: happyhorse-1.1-i2v, alibaba/happyhorse-1.1-i2v, alibaba/happyhorse-1-1-i2v - type: object properties: model: type: string enum: - happyhorse-1.1-r2v - alibaba/happyhorse-1.1-r2v - alibaba/happyhorse-1-1-r2v prompt: type: string description: The text description of the scene, subject, or action to generate in the video. reference_image_urls: type: array items: type: string format: uri minItems: 1 maxItems: 9 description: Array of image URLs for multi-image-to-video generation. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' - '4:3' - '3:4' default: '16:9' description: The aspect ratio of the generated video. resolution: type: string enum: - 720p - 1080p default: 1080p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The length of the output video in seconds. enum: - 3 - 4 - 5 - 6 - 7 - 8 - 9 - 10 - 11 - 12 - 13 - 14 - 15 default: '10' seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. watermark: type: boolean default: false description: Whether the video contains a watermark. required: - model - prompt - reference_image_urls title: happyhorse-1.1-r2v, alibaba/happyhorse-1.1-r2v, alibaba/happyhorse-1-1-r2v - type: object properties: model: type: string enum: - custom:happyhorse-1.1 - alibaba/custom:happyhorse-1.1 prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. reference_image_urls: type: array items: type: string format: uri maxItems: 9 description: Array of image URLs for multi-image-to-video generation. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' - '4:3' - '3:4' default: '16:9' description: The aspect ratio of the generated video. resolution: type: string enum: - 720p - 1080p default: 1080p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The length of the output video in seconds. enum: - 3 - 4 - 5 - 6 - 7 - 8 - 9 - 10 - 11 - 12 - 13 - 14 - 15 default: '10' seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. watermark: type: boolean default: false description: Whether the video contains a watermark. required: - model - prompt title: custom:happyhorse-1.1, alibaba/custom:happyhorse-1.1 - type: object properties: model: type: string enum: - alibaba/wan2.2-vace-fun-a14b-depth - alibaba/wan2.2-vace-fun-a14b-pose - wan-22-vace-fun-a14b/depth - wan-22-vace-fun-a14b/pose prompt: type: string description: The text description of the scene, subject, or action to generate in the video. video_url: type: string format: uri description: A HTTPS URL pointing to a video or a data URI containing a video. This video will be used as a reference during generation. negative_prompt: type: string default: letterboxing, borders, black bars, bright colors, overexposed, static, blurred details, subtitles, style, artwork, painting, picture, still, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, malformed limbs, fused fingers, still picture, cluttered background, three legs, many people in the background, walking backwards description: The description of elements to avoid in the generated video. match_input_num_frames: type: boolean num_frames: type: integer minimum: 17 maximum: 241 default: 81 description: Number of frames to generate. match_input_frames_per_second: type: boolean description: Whether to match the input video's frames per second (FPS). frames_per_second: type: integer minimum: 5 maximum: 30 default: 16 description: Frames per second of the generated video. resolution: type: string enum: - 480p - 580p - 720p default: 480p description: The resolution of the output video, where the number refers to the short side in pixels. aspect_ratio: type: string enum: - auto - '16:9' - '1:1' - '9:16' default: auto description: The aspect ratio of the generated video. num_inference_steps: type: integer default: 30 description: Number of inference steps for sampling. Higher values give better quality but take longer. guidance_scale: type: number default: 5 description: Classifier-free guidance scale. Controls prompt adherence / creativity. shift: type: number default: 5 description: Noise schedule shift parameter. Affects temporal dynamics. enable_safety_checker: type: boolean description: If set to true, the safety checker will be enabled. enable_prompt_expansion: type: boolean description: Whether to enable prompt expansion. preprocess: type: boolean description: Whether to preprocess the input video. acceleration: type: string enum: - none - regular default: regular description: Acceleration to use for inference. video_quality: type: string enum: - low - medium - high - maximum default: high description: The quality of the generated video. video_write_mode: type: string enum: - fast - balanced - small default: balanced description: The method used to write the video. num_interpolated_frames: type: integer description: Number of frames to interpolate between the original frames. temporal_downsample_factor: type: integer description: Temporal downsample factor for the video. enable_auto_downsample: type: boolean description: The minimum frames per second to downsample the video to. auto_downsample_min_fps: type: number default: 15 description: The minimum frames per second to downsample the video to. interpolator_model: type: string enum: - rife - film default: film description: The model to use for interpolation. Rife, or film are available. sync_mode: type: boolean description: The synchronization mode for audio and video. Loose or tight are available. required: - model - prompt - video_url title: alibaba/wan2.2-vace-fun-a14b-depth, alibaba/wan2.2-vace-fun-a14b-pose, wan-22-vace-fun-a14b/depth, wan-22-vace-fun-a14b/pose - type: object properties: model: type: string enum: - alibaba/wan2.2-vace-fun-a14b-inpainting - wan-22-vace-fun-a14b/inpainting prompt: type: string description: The text description of the scene, subject, or action to generate in the video. video_url: type: string format: uri description: A HTTPS URL pointing to a video or a data URI containing a video. This video will be used as a reference during generation. negative_prompt: type: string default: letterboxing, borders, black bars, bright colors, overexposed, static, blurred details, subtitles, style, artwork, painting, picture, still, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, malformed limbs, fused fingers, still picture, cluttered background, three legs, many people in the background, walking backwards description: The description of elements to avoid in the generated video. match_input_num_frames: type: boolean num_frames: type: integer minimum: 17 maximum: 241 default: 81 description: Number of frames to generate. match_input_frames_per_second: type: boolean description: Whether to match the input video's frames per second (FPS). frames_per_second: type: integer minimum: 5 maximum: 30 default: 16 description: Frames per second of the generated video. resolution: type: string enum: - 480p - 580p - 720p default: 480p description: The resolution of the output video, where the number refers to the short side in pixels. aspect_ratio: type: string enum: - auto - '16:9' - '1:1' - '9:16' default: auto description: The aspect ratio of the generated video. num_inference_steps: type: integer default: 30 description: Number of inference steps for sampling. Higher values give better quality but take longer. guidance_scale: type: number default: 5 description: Classifier-free guidance scale. Controls prompt adherence / creativity. shift: type: number default: 5 description: Noise schedule shift parameter. Affects temporal dynamics. enable_safety_checker: type: boolean description: If set to true, the safety checker will be enabled. enable_prompt_expansion: type: boolean description: Whether to enable prompt expansion. preprocess: type: boolean description: Whether to preprocess the input video. acceleration: type: string enum: - none - regular default: regular description: Acceleration to use for inference. video_quality: type: string enum: - low - medium - high - maximum default: high description: The quality of the generated video. video_write_mode: type: string enum: - fast - balanced - small default: balanced description: The method used to write the video. num_interpolated_frames: type: integer description: Number of frames to interpolate between the original frames. temporal_downsample_factor: type: integer description: Temporal downsample factor for the video. enable_auto_downsample: type: boolean description: The minimum frames per second to downsample the video to. auto_downsample_min_fps: type: number default: 15 description: The minimum frames per second to downsample the video to. interpolator_model: type: string enum: - rife - film default: film description: The model to use for interpolation. Rife, or film are available. sync_mode: type: boolean description: The synchronization mode for audio and video. Loose or tight are available. image_list: type: array items: type: string format: uri description: Array of image URLs for multi-image-to-video generation. mask_video_url: type: string format: uri description: URL to the source mask file required: - model - prompt - video_url title: alibaba/wan2.2-vace-fun-a14b-inpainting, wan-22-vace-fun-a14b/inpainting - type: object properties: model: type: string enum: - alibaba/wan2.2-vace-fun-a14b-outpainting - wan-22-vace-fun-a14b/outpainting prompt: type: string description: The text description of the scene, subject, or action to generate in the video. video_url: type: string format: uri description: A HTTPS URL pointing to a video or a data URI containing a video. This video will be used as a reference during generation. negative_prompt: type: string default: letterboxing, borders, black bars, bright colors, overexposed, static, blurred details, subtitles, style, artwork, painting, picture, still, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, malformed limbs, fused fingers, still picture, cluttered background, three legs, many people in the background, walking backwards description: The description of elements to avoid in the generated video. match_input_num_frames: type: boolean num_frames: type: integer minimum: 17 maximum: 241 default: 81 description: Number of frames to generate. match_input_frames_per_second: type: boolean description: Whether to match the input video's frames per second (FPS). frames_per_second: type: integer minimum: 5 maximum: 30 default: 16 description: Frames per second of the generated video. resolution: type: string enum: - 480p - 580p - 720p default: 480p description: The resolution of the output video, where the number refers to the short side in pixels. aspect_ratio: type: string enum: - auto - '16:9' - '1:1' - '9:16' default: auto description: The aspect ratio of the generated video. num_inference_steps: type: integer default: 30 description: Number of inference steps for sampling. Higher values give better quality but take longer. guidance_scale: type: number default: 5 description: Classifier-free guidance scale. Controls prompt adherence / creativity. shift: type: number default: 5 description: Noise schedule shift parameter. Affects temporal dynamics. enable_safety_checker: type: boolean description: If set to true, the safety checker will be enabled. enable_prompt_expansion: type: boolean description: Whether to enable prompt expansion. preprocess: type: boolean description: Whether to preprocess the input video. acceleration: type: string enum: - none - regular default: regular description: Acceleration to use for inference. video_quality: type: string enum: - low - medium - high - maximum default: high description: The quality of the generated video. video_write_mode: type: string enum: - fast - balanced - small default: balanced description: The method used to write the video. num_interpolated_frames: type: integer description: Number of frames to interpolate between the original frames. temporal_downsample_factor: type: integer description: Temporal downsample factor for the video. enable_auto_downsample: type: boolean description: The minimum frames per second to downsample the video to. auto_downsample_min_fps: type: number default: 15 description: The minimum frames per second to downsample the video to. interpolator_model: type: string enum: - rife - film default: film description: The model to use for interpolation. Rife, or film are available. sync_mode: type: boolean description: The synchronization mode for audio and video. Loose or tight are available. expand_left: type: boolean default: true description: Whether to expand the video to the left expand_right: type: boolean default: true description: Whether to expand the video to the right expand_top: type: boolean default: true description: Whether to expand the video to the top expand_bottom: type: boolean default: true description: Whether to expand the video to the bottom expand_ratio: type: number default: 0.25 description: Amount of expansion. This is a float value between 0 and 1, where 0.25 adds 25% to the original video size on the specified sides required: - model - prompt - video_url title: alibaba/wan2.2-vace-fun-a14b-outpainting, wan-22-vace-fun-a14b/outpainting - type: object properties: model: type: string enum: - alibaba/wan2.2-vace-fun-a14b-reframe - wan-22-vace-fun-a14b/reframe prompt: type: string description: The text description of the scene, subject, or action to generate in the video. video_url: type: string format: uri description: A HTTPS URL pointing to a video or a data URI containing a video. This video will be used as a reference during generation. negative_prompt: type: string default: letterboxing, borders, black bars, bright colors, overexposed, static, blurred details, subtitles, style, artwork, painting, picture, still, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, malformed limbs, fused fingers, still picture, cluttered background, three legs, many people in the background, walking backwards description: The description of elements to avoid in the generated video. match_input_num_frames: type: boolean num_frames: type: integer minimum: 17 maximum: 241 default: 81 description: Number of frames to generate. match_input_frames_per_second: type: boolean description: Whether to match the input video's frames per second (FPS). frames_per_second: type: integer minimum: 5 maximum: 30 default: 16 description: Frames per second of the generated video. resolution: type: string enum: - 480p - 580p - 720p default: 480p description: The resolution of the output video, where the number refers to the short side in pixels. aspect_ratio: type: string enum: - auto - '16:9' - '1:1' - '9:16' default: auto description: The aspect ratio of the generated video. num_inference_steps: type: integer default: 30 description: Number of inference steps for sampling. Higher values give better quality but take longer. guidance_scale: type: number default: 5 description: Classifier-free guidance scale. Controls prompt adherence / creativity. shift: type: number default: 5 description: Noise schedule shift parameter. Affects temporal dynamics. enable_safety_checker: type: boolean description: If set to true, the safety checker will be enabled. enable_prompt_expansion: type: boolean description: Whether to enable prompt expansion. preprocess: type: boolean description: Whether to preprocess the input video. acceleration: type: string enum: - none - regular default: regular description: Acceleration to use for inference. video_quality: type: string enum: - low - medium - high - maximum default: high description: The quality of the generated video. video_write_mode: type: string enum: - fast - balanced - small default: balanced description: The method used to write the video. num_interpolated_frames: type: integer description: Number of frames to interpolate between the original frames. temporal_downsample_factor: type: integer description: Temporal downsample factor for the video. enable_auto_downsample: type: boolean description: The minimum frames per second to downsample the video to. auto_downsample_min_fps: type: number default: 15 description: The minimum frames per second to downsample the video to. interpolator_model: type: string enum: - rife - film default: film description: The model to use for interpolation. Rife, or film are available. sync_mode: type: boolean description: The synchronization mode for audio and video. Loose or tight are available. zoom_factor: type: number description: Zoom factor for the video. When this value is greater than 0, the video will be zoomed in by this factor (in relation to the canvas size,) cutting off the edges of the video. A value of 0 means no zoom trim_borders: type: boolean default: true description: Whether to trim borders from the video required: - model - video_url title: alibaba/wan2.2-vace-fun-a14b-reframe, wan-22-vace-fun-a14b/reframe - type: object properties: model: type: string enum: - alibaba/wan2.2-14b-animate-move - alibaba/wan2.2-14b-animate-replace - wan/v2.2-14b/animate/move - wan/v2.2-14b/animate/replace video_url: type: string format: uri description: A HTTPS URL pointing to a video or a data URI containing a video. This video will be used as a reference during generation. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image. If the input image does not match the chosen aspect ratio, it is resized and center cropped resolution: type: string enum: - 480p - 580p - 720p default: 480p description: The resolution of the output video, where the number refers to the short side in pixels. num_inference_steps: type: integer default: 20 description: Number of inference steps for sampling. Higher values give better quality but take longer enable_safety_checker: type: boolean description: If set to true, the safety checker will be enabled. shift: type: number default: 5 description: Shift value for the video. video_quality: type: string enum: - low - medium - high - maximum default: high description: The quality of the generated video. video_write_mode: type: string enum: - fast - balanced - small default: balanced description: The write mode of the output video. Faster write mode means faster results but larger file size, balanced write mode is a good compromise between speed and quality, and small write mode is the slowest but produces the smallest file size required: - model - video_url - image_url title: alibaba/wan2.2-14b-animate-move, alibaba/wan2.2-14b-animate-replace, wan/v2.2-14b/animate/move, wan/v2.2-14b/animate/replace - type: object properties: model: type: string enum: - test/dummy-video prompt: type: string minLength: 1 duration: type: integer minimum: 1 maximum: 10 default: 5 test: type: object properties: delay: type: number runningPolls: type: number errorStatus: type: number submitErrorStatus: type: number required: - model - prompt title: test/dummy-video - type: object properties: model: type: string enum: - video-01 prompt: type: string maxLength: 2000 description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: "A direct link to an online image or a Base64-encoded local image that will serve as the first frame for the video.\nImage specifications: \n- format must be JPG, JPEG, or PNG; \n- aspect ratio should be greater than 2:5 and less than 5:2; \n- the shorter side must exceed 300 pixels; \n- file size must not exceed 20MB." enhance_prompt: type: boolean default: true description: If True, the incoming prompt will be automatically optimized to improve generation quality when needed. For more precise control, set it to False — the model will then follow the instructions more strictly. required: - model - prompt title: video-01 - type: object properties: model: type: string enum: - video-01-live2d prompt: type: string maxLength: 2000 description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: "A direct link to an online image or a Base64-encoded local image that will serve as the first frame for the video.\nImage specifications: \n- format must be JPG, JPEG, or PNG; \n- aspect ratio should be greater than 2:5 and less than 5:2; \n- the shorter side must exceed 300 pixels; \n- file size must not exceed 20MB." required: true enhance_prompt: type: boolean default: true description: If True, the incoming prompt will be automatically optimized to improve generation quality when needed. For more precise control, set it to False — the model will then follow the instructions more strictly. required: - model - prompt - image_url title: video-01-live2d - type: object properties: model: type: string enum: - minimax/hailuo-02 prompt: type: string maxLength: 2000 description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: "A direct link to an online image or a Base64-encoded local image that will serve as the first frame for the video.\nImage specifications: \n- format must be JPG, JPEG, or PNG; \n- aspect ratio should be greater than 2:5 and less than 5:2; \n- the shorter side must exceed 300 pixels; \n- file size must not exceed 20MB." last_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. resolution: type: string enum: - 768P - 1080P default: 768P description: The dimensions of the video display. 1080p corresponds to 1920 x 1080 pixels, 768p corresponds to 1366 x 768 pixels. duration: type: integer description: The length of the output video in seconds. For 1080p resolution, only a duration of 6 seconds is supported enum: - 6 - 10 enhance_prompt: type: boolean default: true description: If True, the incoming prompt will be automatically optimized to improve generation quality when needed. For more precise control, set it to False — the model will then follow the instructions more strictly. fast_pretreatment: type: boolean default: false description: Reduces optimization time when enhance_prompt is enabled. required: - model - prompt title: minimax/hailuo-02 - type: object properties: model: type: string enum: - minimax/hailuo-2.3 prompt: type: string maxLength: 2000 description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: "A direct link to an online image or a Base64-encoded local image that will serve as the first frame for the video.\nImage specifications: \n- format must be JPG, JPEG, or PNG; \n- aspect ratio should be greater than 2:5 and less than 5:2; \n- the shorter side must exceed 300 pixels; \n- file size must not exceed 20MB." resolution: type: string enum: - 768P - 1080P default: 768P description: The dimensions of the video display. 1080p corresponds to 1920 x 1080 pixels, 768p corresponds to 1366 x 768 pixels. duration: type: integer description: The length of the output video in seconds. For 1080p resolution, only a duration of 6 seconds is supported enum: - 6 - 10 enhance_prompt: type: boolean default: true description: If True, the incoming prompt will be automatically optimized to improve generation quality when needed. For more precise control, set it to False — the model will then follow the instructions more strictly. fast_pretreatment: type: boolean default: false description: Reduces optimization time when enhance_prompt is enabled. required: - model - prompt title: minimax/hailuo-2.3 - type: object properties: model: type: string enum: - minimax/hailuo-2.3-fast prompt: type: string maxLength: 2000 description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: "A direct link to an online image or a Base64-encoded local image that will serve as the first frame for the video.\nImage specifications: \n- format must be JPG, JPEG, or PNG; \n- aspect ratio should be greater than 2:5 and less than 5:2; \n- the shorter side must exceed 300 pixels; \n- file size must not exceed 20MB." resolution: type: string enum: - 768P - 1080P default: 768P description: The dimensions of the video display. 1080p corresponds to 1920 x 1080 pixels, 768p corresponds to 1366 x 768 pixels. duration: type: integer description: The length of the output video in seconds. For 1080p resolution, only a duration of 6 seconds is supported enum: - 6 - 10 enhance_prompt: type: boolean default: true description: If True, the incoming prompt will be automatically optimized to improve generation quality when needed. For more precise control, set it to False — the model will then follow the instructions more strictly. fast_pretreatment: type: boolean default: false description: Reduces optimization time when enhance_prompt is enabled. required: - model - prompt - image_url title: minimax/hailuo-2.3-fast - type: object properties: model: type: string enum: - minimax/h3 prompt: type: string maxLength: 7000 description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: "A direct link to an online image or a Base64-encoded local image that will serve as the first frame for the video.\nImage specifications: \n- format must be JPG, JPEG, or PNG; \n- aspect ratio should be greater than 2:5 and less than 5:2; \n- the shorter side must exceed 300 pixels; \n- file size must not exceed 20MB." last_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. reference_image_urls: type: array items: type: string format: uri maxItems: 9 description: Passing an image reference allows the model to emulate the style or content of the reference in the output. video_urls: type: array items: type: string format: uri maxItems: 3 description: Passing a video reference allows the model to emulate the style or content of the reference in the output. audio_urls: type: array items: type: string format: uri maxItems: 3 description: Reference audio clips whose voice timbre the generated speech follows. Cannot be the only reference — pass at least one reference image or video alongside. duration: type: integer minimum: 4 maximum: 15 default: 6 description: The length of the output video in seconds. resolution: type: string enum: - 2K default: 2K description: The resolution of the output video, where the number refers to the short side in pixels. ratio: type: string enum: - adaptive - '21:9' - '16:9' - '4:3' - '1:1' - '3:4' - '9:16' default: adaptive description: The aspect ratio of the generated video. Defaults to `adaptive`, which follows the first/last frame or reference input; text-only generation cannot be `adaptive` and defaults to `16:9` instead. required: - model - prompt title: minimax/h3 - type: object properties: model: type: string enum: - xai/grok-imagine-video - x-ai/grok-imagine-video prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: URL of the source image to animate (image-to-video mode). Omit for text-to-video. Cannot be combined with `reference_images`. reference_images: type: array items: type: string format: uri minItems: 1 maxItems: 7 description: Array of reference image URLs (reference-to-video mode). Up to 7 images. Cannot be combined with `image_url`. duration: type: integer description: The length of the output video in seconds. minimum: 1 maximum: 15 default: 10 aspect_ratio: type: string enum: - '1:1' - '16:9' - '9:16' - '4:3' - '3:4' - '3:2' - '2:3' default: '16:9' description: The aspect ratio of the generated video. resolution: type: string enum: - 480p - 720p description: 'Output video resolution. For image-to-video and reference-to-video modes, the value is selected automatically based on the resolution of the input images.' required: - model title: xai/grok-imagine-video, x-ai/grok-imagine-video - type: object properties: model: type: string enum: - xai/grok-imagine-video-1.5-preview - x-ai/grok-imagine-video-1.5-preview prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: URL of the source image to animate. Required — this model supports image-to-video only. duration: type: integer description: The length of the output video in seconds. minimum: 1 maximum: 15 default: 10 aspect_ratio: type: string enum: - '1:1' - '16:9' - '9:16' - '4:3' - '3:4' - '3:2' - '2:3' default: '16:9' description: The aspect ratio of the generated video. resolution: type: string enum: - 480p - 720p description: 'Output video resolution. The value is selected automatically based on the resolution of the input image.' required: - model - image_url title: xai/grok-imagine-video-1.5-preview, x-ai/grok-imagine-video-1.5-preview - type: object properties: model: type: string enum: - kling-video/v1/standard/image-to-video - kling-video/v1/pro/image-to-video - kling-video/v1.5/pro/image-to-video provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. prompt: type: string description: The text description of the scene, subject, or action to generate in the video. tail_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' negative_prompt: type: string description: The description of elements to avoid in the generated video. cfg_scale: type: number minimum: 0 maximum: 1 description: The CFG (Classifier Free Guidance) scale is a measure of how close you want the model to stick to your prompt. static_mask: type: string format: uri description: URL of the image for Static Brush Application Area (Mask image created by users using the motion brush). dynamic_masks: type: array items: type: object properties: mask: type: string format: uri trajectories: type: array items: type: object properties: x: type: integer y: type: integer required: - x - y minItems: 2 maxItems: 77 required: - mask - trajectories maxItems: 6 description: List of dynamic masks. required: - model - image_url title: kling-video/v1/standard/image-to-video, kling-video/v1/pro/image-to-video, kling-video/v1.5/pro/image-to-video - type: object properties: model: type: string enum: - kling-video/v1/standard/text-to-video - kling-video/v1/pro/text-to-video - kling-video/v1.5/pro/text-to-video - klingai/v2-master-text-to-video provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto prompt: type: string description: The text description of the scene, subject, or action to generate in the video. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' default: '16:9' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 negative_prompt: type: string description: The description of elements to avoid in the generated video. cfg_scale: type: number minimum: 0 maximum: 1 description: The CFG (Classifier Free Guidance) scale is a measure of how close you want the model to stick to your prompt. camera_control: type: string enum: - down_back - forward_up - right_turn_forward - left_turn_forward description: Camera control parameters. advanced_camera_control: type: object properties: movement_type: type: string enum: - horizontal - vertical - pan - tilt - roll - zoom default: horizontal description: The type of camera movement. movement_value: type: integer minimum: -10 maximum: 10 default: 0 description: The value of the camera movement. description: Advanced camera control parameters. required: - model - prompt title: kling-video/v1/standard/text-to-video, kling-video/v1/pro/text-to-video, kling-video/v1.5/pro/text-to-video, klingai/v2-master-text-to-video - type: object properties: model: type: string enum: - kling-video/v1.6/standard/text-to-video - kling-video/v1.6/pro/text-to-video - klingai/v2.1-master-text-to-video - klingai/v2.5-turbo/pro/text-to-video provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto prompt: type: string description: The text description of the scene, subject, or action to generate in the video. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' default: '16:9' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 negative_prompt: type: string description: The description of elements to avoid in the generated video. cfg_scale: type: number minimum: 0 maximum: 1 description: The CFG (Classifier Free Guidance) scale is a measure of how close you want the model to stick to your prompt. required: - model - prompt title: kling-video/v1.6/standard/text-to-video, kling-video/v1.6/pro/text-to-video, klingai/v2.1-master-text-to-video, klingai/v2.5-turbo/pro/text-to-video - type: object properties: model: type: string enum: - kling-video/v1.6/standard/image-to-video - kling-video/v2.1/standard/image-to-video provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. prompt: type: string description: The text description of the scene, subject, or action to generate in the video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' negative_prompt: type: string description: The description of elements to avoid in the generated video. required: - model - image_url title: kling-video/v1.6/standard/image-to-video, kling-video/v2.1/standard/image-to-video - type: object properties: model: type: string enum: - kling-video/v1.6/pro/image-to-video - kling-video/v2.1/pro/image-to-video provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. prompt: type: string description: The text description of the scene, subject, or action to generate in the video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' negative_prompt: type: string description: The description of elements to avoid in the generated video. tail_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. required: - model - image_url title: kling-video/v1.6/pro/image-to-video, kling-video/v2.1/pro/image-to-video - type: object properties: model: type: string enum: - klingai/kling-video-v1.6-pro-effects - klingai/kling-video-v1.6-standard-effects image_url: anyOf: - type: string format: uri - type: array items: type: string format: uri description: For hug, kiss, and heart_gesture effects, pass an array containing exactly two image URLs. For squish or expansion, only one image URL is required. effect_scene: type: string enum: - magic_fireball - pet_moto_rider - media_interview - pet_lion - pet_delivery - pet_chef - santa_gifts - santa_hug - girlfriend - boyfriend - heart_gesture_1 - pet_wizard - smoke_smoke - thumbs_up - instant_kid - dollar_rain - cry_cry - building_collapse - gun_shot - mushroom - double_gun - pet_warrior - lightning_power - jesus_hug - shark_alert - long_hair - lie_flat - polar_bear_hug - brown_bear_hug - jazz_jazz - office_escape_plow - fly_fly - watermelon_bomb - pet_dance - boss_coming - wool_curly - iron_warrior - pet_bee - marry_me - swing_swing - day_to_night - piggy_morph - wig_out - car_explosion - ski_ski - tiger_hug - siblings - construction_worker - let's_ride - snatched - magic_broom - felt_felt - jumpdrop - celebration - splashsplash - hula - surfsurf - fairy_wing - angel_wing - dark_wing - skateskate - plushcut - jelly_press - jelly_slice - jelly_squish - jelly_jiggle - pixelpixel - yearbook - instant_film - anime_figure - rocketrocket - bloombloom - dizzydizzy - fuzzyfuzzy - squish - expansion - hug - kiss - heart_gesture - fight description: Video effect scene type duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' required: - model - image_url - effect_scene title: klingai/kling-video-v1.6-pro-effects, klingai/kling-video-v1.6-standard-effects - type: object properties: model: type: string enum: - kling-video/v1.6/standard/multi-image-to-video image_list: type: array items: type: string format: uri minItems: 2 maxItems: 4 description: Array of image URLs for multi-image-to-video generation prompt: type: string description: The text description of the scene, subject, or action to generate in the video. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' default: '16:9' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' negative_prompt: type: string description: The description of elements to avoid in the generated video. required: - model - image_list title: kling-video/v1.6/standard/multi-image-to-video - type: object properties: model: type: string enum: - klingai/v2-master-image-to-video - klingai/v2.1-master-image-to-video provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. prompt: type: string description: The text description of the scene, subject, or action to generate in the video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' negative_prompt: type: string description: The description of elements to avoid in the generated video. cfg_scale: type: number minimum: 0 maximum: 1 description: The CFG (Classifier Free Guidance) scale is a measure of how close you want the model to stick to your prompt. required: - model - image_url title: klingai/v2-master-image-to-video, klingai/v2.1-master-image-to-video - type: object properties: model: type: string enum: - klingai/v2.5-turbo/pro/image-to-video provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. prompt: type: string description: The text description of the scene, subject, or action to generate in the video. tail_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' negative_prompt: type: string description: The description of elements to avoid in the generated video. cfg_scale: type: number minimum: 0 maximum: 1 description: The CFG (Classifier Free Guidance) scale is a measure of how close you want the model to stick to your prompt. required: - model - image_url title: klingai/v2.5-turbo/pro/image-to-video - type: object properties: model: type: string enum: - klingai/avatar-standard - klingai/avatar-pro - kling-video/v1/standard/ai-avatar - kling-video/v1/pro/ai-avatar provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. audio_url: type: string format: uri description: 'The URL of the audio file. Supported formats: MP3, WAV, M4A, AAC. Maximum file size: 5 MB.' prompt: type: string maxLength: 2500 description: The text description of the scene, subject, or action to generate in the video. required: - model - image_url - audio_url title: klingai/avatar-standard, klingai/avatar-pro, kling-video/v1/standard/ai-avatar, kling-video/v1/pro/ai-avatar - type: object properties: model: type: string enum: - klingai/video-o1-image-to-video - kling-video/o1/image-to-video provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto prompt: type: string maxLength: 2500 description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. last_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' required: - model - prompt - image_url title: klingai/video-o1-image-to-video, kling-video/o1/image-to-video - type: object properties: model: type: string enum: - klingai/video-o1-reference-to-video - kling-video/o1/reference-to-video provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto prompt: type: string maxLength: 2500 description: The text description of the scene, subject, or action to generate in the video. image_list: type: array items: type: string format: uri minItems: 1 maxItems: 7 description: Array of image URLs for multi-image-to-video generation. elements: type: array items: type: object properties: reference_image_urls: type: array items: type: string format: uri minItems: 1 maxItems: 4 description: Additional reference images from different angles. frontal_image_url: type: string format: uri description: The frontal image of the element (main view). required: - reference_image_urls - frontal_image_url maxItems: 4 description: Elements (characters/objects) to include in the video. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' default: '16:9' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' required: - model - prompt title: klingai/video-o1-reference-to-video, kling-video/o1/reference-to-video - type: object properties: model: type: string enum: - klingai/video-o1-video-to-video-edit - kling-video/o1/video-to-video/edit provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto prompt: type: string maxLength: 2500 description: The text description of the scene, subject, or action to generate in the video. video_url: type: string format: uri description: A HTTPS URL pointing to a video or a data URI containing a video. This video will be used as a reference during generation. image_list: type: array items: type: string format: uri minItems: 1 maxItems: 7 description: Array of image URLs for multi-image-to-video generation. elements: type: array items: type: object properties: reference_image_urls: type: array items: type: string format: uri minItems: 1 maxItems: 4 description: Additional reference images from different angles. frontal_image_url: type: string format: uri description: The frontal image of the element (main view). required: - reference_image_urls - frontal_image_url maxItems: 4 description: Elements (characters/objects) to include in the video. keep_audio: type: boolean default: false description: Whether to keep the original audio from the video. required: - model - prompt - video_url title: klingai/video-o1-video-to-video-edit, kling-video/o1/video-to-video/edit - type: object properties: model: type: string enum: - klingai/video-o1-video-to-video-reference - kling-video/o1/video-to-video/reference provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto prompt: type: string maxLength: 2500 description: The text description of the scene, subject, or action to generate in the video. video_url: type: string format: uri description: A HTTPS URL pointing to a video or a data URI containing a video. This video will be used as a reference during generation. image_list: type: array items: type: string format: uri minItems: 1 maxItems: 4 description: Array of image URLs for multi-image-to-video generation. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' default: '16:9' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' elements: type: array items: type: object properties: reference_image_urls: type: array items: type: string format: uri minItems: 1 maxItems: 4 description: Additional reference images from different angles. frontal_image_url: type: string format: uri description: The frontal image of the element (main view). required: - reference_image_urls - frontal_image_url maxItems: 4 description: Elements (characters/objects) to include in the video. keep_audio: type: boolean default: false description: Whether to keep the original audio from the video. required: - model - prompt - video_url title: klingai/video-o1-video-to-video-reference, kling-video/o1/video-to-video/reference - type: object properties: model: type: string enum: - klingai/video-v2-6-pro-image-to-video - kling-video/v2.6/pro/image-to-video provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto prompt: type: string maxLength: 2500 description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. tail_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' negative_prompt: type: string description: The description of elements to avoid in the generated video. generate_audio: type: boolean default: true description: Whether to generate audio for the video. required: - model - prompt - image_url title: klingai/video-v2-6-pro-image-to-video, kling-video/v2.6/pro/image-to-video - type: object properties: model: type: string enum: - klingai/video-v2-6-pro-text-to-video - kling-video/v2.6/pro/text-to-video provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto prompt: type: string maxLength: 2500 description: The text description of the scene, subject, or action to generate in the video. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' negative_prompt: type: string description: The description of elements to avoid in the generated video. cfg_scale: type: number minimum: 0 maximum: 1 description: The CFG (Classifier Free Guidance) scale is a measure of how close you want the model to stick to your prompt. generate_audio: type: boolean default: true description: Whether to generate audio for the video. required: - model - prompt title: klingai/video-v2-6-pro-text-to-video, kling-video/v2.6/pro/text-to-video - type: object properties: model: type: string enum: - klingai/video-v2-6-pro-motion-control prompt: type: string description: Optional instructions that define the background elements, including their appearance, timing in the frame, and behavior, and can also subtly adjust the character’s animation. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that serves as the character reference for animation. The image must contain exactly one clearly visible character, who will be animated using the motion from the reference video provided in the video_url parameter. For optimal results, be sure the character’s proportions in the image match those in the video. video_url: type: string format: uri description: A HTTPS URL pointing to a video or a data URI containing a video. The character’s movements from this video will be applied to the character from the image provided in the image_url parameter. For best results, use a video with a single clearly visible character. If the video contains two or more characters, the motion of the character occupying the largest portion of the frame will be used for generation. character_orientation: type: string enum: - image - video default: video description: 'Generate the orientation of the character in the video, which can be selected to match the image or the video: - image: has the same orientation as the person in the picture; At this time, the reference video duration should not exceed 10 seconds; - video: consistent with the orientation of the characters in the video; At this time, the reference video duration should not exceed 30 seconds;' keep_audio: type: boolean default: true description: Whether to keep the original audio from the video. required: - model - image_url - video_url title: klingai/video-v2-6-pro-motion-control - type: object properties: model: type: string enum: - klingai/video-v2-6-motion-control - klingai/video-v3-motion-control prompt: type: string description: Optional instructions that define the background elements, including their appearance, timing in the frame, and behavior, and can also subtly adjust the character’s animation. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that serves as the character reference for animation. The image must contain exactly one clearly visible character, who will be animated using the motion from the reference video provided in the video_url parameter. For optimal results, be sure the character’s proportions in the image match those in the video. video_url: type: string format: uri description: A HTTPS URL pointing to a video or a data URI containing a video. The character’s movements from this video will be applied to the character from the image provided in the image_url parameter. For best results, use a video with a single clearly visible character. If the video contains two or more characters, the motion of the character occupying the largest portion of the frame will be used for generation. character_orientation: type: string enum: - image - video default: video description: 'Generate the orientation of the character in the video, which can be selected to match the image or the video: - image: has the same orientation as the person in the picture; At this time, the reference video duration should not exceed 10 seconds; - video: consistent with the orientation of the characters in the video; At this time, the reference video duration should not exceed 30 seconds;' keep_audio: type: boolean default: true description: Whether to keep the original audio from the video. mode: type: string enum: - std - pro default: std description: 'Video generation mode: - std: Standard Mode — basic, cost-effective. - pro: Professional Mode — higher quality, higher cost.' required: - model - image_url - video_url title: klingai/video-v2-6-motion-control, klingai/video-v3-motion-control - type: object properties: model: type: string enum: - klingai/video-v3-standard-text-to-video - klingai/video-v3-pro-text-to-video - klingai/video-v3-omni-720p-text-to-video - klingai/video-v3-omni-1080p-text-to-video - kling-video/v3/standard/text-to-video - kling-video/v3/pro/text-to-video prompt: type: string description: Text prompt for video generation. Either prompt or multi_prompt must be provided, but not both. provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto multi_prompt: type: array items: type: string description: List of prompts for multi-shot video generation. If provided, overrides the single prompt and divides the video into multiple shots with specified prompts and durations. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' default: '16:9' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 3 - 4 - 5 - 6 - 7 - 8 - 9 - 10 - 11 - 12 - 13 - 14 - 15 default: '5' shot_type: type: string enum: - customize - intelligent - intelligence default: customize description: The type of multi-shot video generation generate_audio: type: boolean default: true description: Whether to generate audio for the video. negative_prompt: type: string description: The description of elements to avoid in the generated video. cfg_scale: type: number minimum: 0 maximum: 1 description: The CFG (Classifier Free Guidance) scale is a measure of how close you want the model to stick to your prompt. required: - model title: klingai/video-v3-standard-text-to-video, klingai/video-v3-pro-text-to-video, klingai/video-v3-omni-720p-text-to-video, klingai/video-v3-omni-1080p-text-to-video, kling-video/v3/standard/text-to-video, kling-video/v3/pro/text-to-video - type: object properties: model: type: string enum: - klingai/video-v3-standard-image-to-video - klingai/video-v3-pro-image-to-video - klingai/video-v3-omni-720p-image-to-video - klingai/video-v3-omni-1080p-image-to-video - kling-video/v3/standard/image-to-video - kling-video/v3/pro/image-to-video prompt: type: string description: Text prompt for video generation. Either prompt or multi_prompt must be provided, but not both. provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto multi_prompt: type: array items: type: string description: List of prompts for multi-shot video generation. If provided, overrides the single prompt and divides the video into multiple shots with specified prompts and durations. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. tail_image_url: type: string format: uri description: Last frame image URL. Not supported when generate_audio is enabled — set generate_audio to false to use a tail frame. duration: type: integer description: The length of the output video in seconds. enum: - 3 - 4 - 5 - 6 - 7 - 8 - 9 - 10 - 11 - 12 - 13 - 14 - 15 default: '5' elements: type: array items: type: object properties: reference_image_urls: type: array items: type: string format: uri minItems: 1 maxItems: 4 description: Additional reference images from different angles. frontal_image_url: type: string format: uri description: The frontal image of the element (main view). video_url: type: string format: uri description: The video URL of the element. A request can only have one element with a video. required: - reference_image_urls maxItems: 4 description: Elements (characters/objects) to include in the video. Each example can either be an image set (frontal + reference images) or a video shot_type: type: string enum: - customize - intelligent - intelligence default: customize description: The type of multi-shot video generation generate_audio: type: boolean default: true description: Whether to generate audio for the video. negative_prompt: type: string description: The description of elements to avoid in the generated video. cfg_scale: type: number minimum: 0 maximum: 1 description: The CFG (Classifier Free Guidance) scale is a measure of how close you want the model to stick to your prompt. required: - model - image_url title: klingai/video-v3-standard-image-to-video, klingai/video-v3-pro-image-to-video, klingai/video-v3-omni-720p-image-to-video, klingai/video-v3-omni-1080p-image-to-video, kling-video/v3/standard/image-to-video, kling-video/v3/pro/image-to-video - type: object properties: model: type: string enum: - klingai/video-v3-standard-turbo-text-to-video - klingai/video-v3-turbo-pro-text-to-video - kling-video/v3/turbo/standard/text-to-video - kling-video/v3/turbo/pro/text-to-video prompt: type: string maxLength: 3072 description: Optional text prompt. For best results keep the prompt under 2500 characters. Mutually exclusive with multi_prompt. multi_prompt: type: - array - 'null' items: type: object properties: prompt: type: string maxLength: 3072 description: Text prompt for this storyboard shot. duration: type: integer description: The length of the output video in seconds. enum: - 1 - 2 - 3 - 4 - 5 - 6 - 7 - 8 - 9 - 10 - 11 - 12 - 13 - 14 - 15 default: '5' required: - prompt minItems: 1 maxItems: 6 description: Multi-shot storyboard with 1 to 6 shots. Mutually exclusive with prompt. duration: type: integer description: The length of the output video in seconds. enum: - 3 - 4 - 5 - 6 - 7 - 8 - 9 - 10 - 11 - 12 - 13 - 14 - 15 default: '5' aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' default: '16:9' description: The aspect ratio of the generated video. required: - model title: klingai/video-v3-standard-turbo-text-to-video, klingai/video-v3-turbo-pro-text-to-video, kling-video/v3/turbo/standard/text-to-video, kling-video/v3/turbo/pro/text-to-video - type: object properties: model: type: string enum: - klingai/video-v3-standard-turbo-image-to-video - klingai/video-v3-turbo-pro-image-to-video - kling-video/v3/turbo/standard/image-to-video - kling-video/v3/turbo/pro/image-to-video prompt: type: string maxLength: 3072 description: Optional text prompt. For best results keep the prompt under 2500 characters. Mutually exclusive with multi_prompt. multi_prompt: type: - array - 'null' items: type: object properties: prompt: type: string maxLength: 3072 description: Text prompt for this storyboard shot. duration: type: integer description: The length of the output video in seconds. enum: - 1 - 2 - 3 - 4 - 5 - 6 - 7 - 8 - 9 - 10 - 11 - 12 - 13 - 14 - 15 default: '5' required: - prompt minItems: 1 maxItems: 6 description: Multi-shot storyboard with 1 to 6 shots. Mutually exclusive with prompt. duration: type: integer description: The length of the output video in seconds. enum: - 3 - 4 - 5 - 6 - 7 - 8 - 9 - 10 - 11 - 12 - 13 - 14 - 15 default: '5' image_url: type: string format: uri description: 'First-frame reference image. Formats: .jpg/.jpeg/.png; max 50MB; min 300px per side; aspect ratio within 1:2.5 to 2.5:1.' required: - model - image_url title: klingai/video-v3-standard-turbo-image-to-video, klingai/video-v3-turbo-pro-image-to-video, kling-video/v3/turbo/standard/image-to-video, kling-video/v3/turbo/pro/image-to-video - type: object properties: model: type: string enum: - pixverse/v5/text-to-video prompt: type: string description: The text description of the scene, subject, or action to generate in the video. aspect_ratio: type: string enum: - '16:9' - '4:3' - '1:1' - '3:4' - '9:16' default: '16:9' description: The aspect ratio of the generated video. resolution: type: string enum: - 360p - 540p - 720p - 1080p default: 720p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The output video length in seconds. The 1080p quality option does not support 8-second videos. enum: - 5 - 8 default: '5' negative_prompt: type: string description: The description of elements to avoid in the generated video. style: type: string enum: - anime - 3d_animation - clay - comic - cyberpunk description: The style of the generated video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. lip_sync_tts_content: type: string description: The text content to be lip-synced in the video. lip_sync_tts_speaker: type: string enum: - Harper - Ava - Isabella - Sophia - Emily - Chloe - Julia - Mason - Jack - Liam - James - Oliver - Adrian - Ethan - Auto description: A predefined system voice used for generating speech in the video. required: - model - prompt title: pixverse/v5/text-to-video - type: object properties: model: type: string enum: - pixverse/v5/image-to-video prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: URL of the image to be used as the first frame of the video. resolution: type: string enum: - 360p - 540p - 720p - 1080p default: 720p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The output video length in seconds. The 1080p quality option does not support 8-second videos. enum: - 5 - 8 default: '5' negative_prompt: type: string description: The description of elements to avoid in the generated video. style: type: string enum: - anime - 3d_animation - clay - comic - cyberpunk description: The style of the generated video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. lip_sync_tts_content: type: string description: The text content to be lip-synced in the video. lip_sync_tts_speaker: type: string enum: - Harper - Ava - Isabella - Sophia - Emily - Chloe - Julia - Mason - Jack - Liam - James - Oliver - Adrian - Ethan - Auto description: A predefined system voice used for generating speech in the video. required: - model - prompt - image_url title: pixverse/v5/image-to-video - type: object properties: model: type: string enum: - pixverse/v5/transition prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: URL of the image to be used as the first frame of the video. tail_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. resolution: type: string enum: - 360p - 540p - 720p - 1080p default: 720p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The output video length in seconds. The 1080p quality option does not support 8-second videos. enum: - 5 - 8 default: '5' negative_prompt: type: string description: The description of elements to avoid in the generated video. style: type: string enum: - anime - 3d_animation - clay - comic - cyberpunk description: The style of the generated video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. lip_sync_tts_content: type: string description: The text content to be lip-synced in the video. lip_sync_tts_speaker: type: string enum: - Harper - Ava - Isabella - Sophia - Emily - Chloe - Julia - Mason - Jack - Liam - James - Oliver - Adrian - Ethan - Auto description: A predefined system voice used for generating speech in the video. required: - model - prompt - image_url - tail_image_url title: pixverse/v5/transition - type: object properties: model: type: string enum: - pixverse/lip-sync video_url: type: string format: uri description: A HTTPS URL pointing to a video or a data URI containing a video. This video will be used as a reference during generation. audio_url: type: string format: uri description: A direct link to an online audio file or a Base64-encoded local to an audio file used for lip-syncing in the video. Use either audio_url or (lip_sync_tts_speaker together with lip_sync_tts_content), but not both. lip_sync_tts_content: type: string description: The text content to be lip-synced in the video. Use either audio_url or (lip_sync_tts_speaker together with lip_sync_tts_content), but not both. lip_sync_tts_speaker: type: string enum: - Harper - Ava - Isabella - Sophia - Emily - Chloe - Julia - Mason - Jack - Liam - James - Oliver - Adrian - Ethan - Auto description: A predefined system voice used for generating speech in the video. Use either audio_url or (lip_sync_tts_speaker together with lip_sync_tts_content), but not both. required: - model - video_url title: pixverse/lip-sync - type: object properties: model: type: string enum: - pixverse/v5.5/text-to-video - pixverse/v5-5-text-to-video prompt: type: string description: The text description of the scene, subject, or action to generate in the video. aspect_ratio: type: string enum: - '16:9' - '4:3' - '1:1' - '3:4' - '9:16' default: '16:9' description: The aspect ratio of the generated video. resolution: type: string enum: - 360p - 540p - 720p - 1080p default: 720p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The output video length in seconds. The 1080p quality option does not support 8-second videos. enum: - 5 - 8 - 10 default: '5' negative_prompt: type: string description: The description of elements to avoid in the generated video. style: type: string enum: - anime - 3d_animation - clay - comic - cyberpunk description: The style of the generated video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. generate_audio_switch: type: boolean default: false description: 'Enable audio generation. - true: Audio on. - false: Audio off.' generate_multi_clip_switch: type: boolean default: false description: "Enable multi-clip generation with dynamic camera changes.\n- true: Multi-clip. \n- false: Single-clip." thinking_type: type: string enum: - enabled - disabled - auto default: enabled description: "Prompt reasoning enhancement mode. \n- \"enabled\": Turn on prompt optimization. \n- \"disabled\": Turn off prompt optimization. \n- \"auto\" or omitted: Let the model decide automatically." required: - model - prompt title: pixverse/v5.5/text-to-video, pixverse/v5-5-text-to-video - type: object properties: model: type: string enum: - pixverse/v5.5/image-to-video - pixverse/v5-5-image-to-video prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: URL of the image to be used as the first frame of the video. resolution: type: string enum: - 360p - 540p - 720p - 1080p default: 720p description: An enumeration where the short side of the video frame determines the resolution. duration: type: integer description: The output video length in seconds. The 1080p quality option does not support 8-second videos. enum: - 5 - 8 - 10 default: '5' negative_prompt: type: string description: The description of elements to avoid in the generated video. style: type: string enum: - anime - 3d_animation - clay - comic - cyberpunk description: The style of the generated video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. generate_audio_switch: type: boolean default: false description: 'Enable audio generation. - true: Audio on. - false: Audio off.' generate_multi_clip_switch: type: boolean default: false description: "Enable multi-clip generation with dynamic camera changes.\n- true: Multi-clip. \n- false: Single-clip." thinking_type: type: string enum: - enabled - disabled - auto default: enabled description: "Prompt reasoning enhancement mode. \n- \"enabled\": Turn on prompt optimization. \n- \"disabled\": Turn off prompt optimization. \n- \"auto\" or omitted: Let the model decide automatically." required: - model - prompt - image_url title: pixverse/v5.5/image-to-video, pixverse/v5-5-image-to-video - type: object properties: model: type: string enum: - ray-2 - luma/ray-2 - ray-flash-2 - luma/ray-flash-2 prompt: type: string description: The text description of the scene, subject, or action to generate in the video. resolution: type: string enum: - 540p - 720p - 1080p - 4k default: 1080p description: The resolution of the output video, where the number refers to the short side in pixels. aspect_ratio: type: string enum: - '1:1' - '16:9' - '9:16' - '4:3' - '3:4' - '21:9' - '9:21' default: '16:9' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 9 default: '5' keyframes: type: object properties: frame0: anyOf: - type: object properties: type: type: string enum: - image url: type: string format: uri required: - type - url - type: object properties: type: type: string enum: - generation id: type: string format: uuid required: - type - id - {} frame1: anyOf: - type: object properties: type: type: string enum: - image url: type: string format: uri required: - type - url - type: object properties: type: type: string enum: - generation id: type: string format: uuid required: - type - id - {} description: Keyframes for image-to-video, extend, or interpolate loop: type: boolean default: false description: Whether to loop the video required: - model - prompt title: ray-2, luma/ray-2, ray-flash-2, luma/ray-flash-2 - type: object properties: model: type: string enum: - luma/ray-3.2 prompt: type: string description: Text description of the video to generate. type: type: string enum: - video - video_edit - video_reframe default: video description: 'Generation kind: "video" (generate / extend), "video_edit", or "video_reframe".' resolution: type: string enum: - 360p - 540p - 720p - 1080p default: 720p description: Output resolution. duration: type: integer description: Clip duration in seconds. enum: - 5 - 10 default: '5' aspect_ratio: type: string enum: - '9:16' - '3:4' - '1:1' - '4:3' - '16:9' - '21:9' default: '16:9' description: Output aspect ratio. loop: type: boolean default: false description: Seamlessly loop the video (creation only). hdr: type: boolean default: false description: Generate HDR output (5s, 720p/1080p only). exr_export: type: boolean default: false description: Export EXR frames (requires hdr). image_url: type: string format: uri description: Start-frame image URL for image-to-video / extend. last_image_url: type: string format: uri description: End-frame image URL (first-last-frame). video_url: type: string format: uri description: Source video URL for editing or reframing. required: - model - prompt title: luma/ray-3.2 - type: object properties: model: type: string enum: - gen3a_turbo - runway/gen3a_turbo prompt: type: string maxLength: 1000 description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A HTTPS URL or data URI containing an encoded image to be used as the first frame of the generated video. tail_image_url: type: string format: uri description: A HTTPS URL or data URI containing an encoded image to be used as the last frame of the generated video. aspect_ratio: type: string enum: - '16:9' - '9:16' default: '16:9' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' seed: type: integer minimum: 0 maximum: 4294967295 description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. required: - model - image_url title: gen3a_turbo, runway/gen3a_turbo - type: object properties: model: type: string enum: - gen4_turbo - runway/gen4_turbo prompt: type: string maxLength: 1000 description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A HTTPS URL or data URI containing an encoded image to be used as the first frame of the generated video. tail_image_url: type: string format: uri description: A HTTPS URL or data URI containing an encoded image to be used as the last frame of the generated video. aspect_ratio: type: string enum: - '16:9' - '9:16' - '4:3' - '3:4' - '1:1' - '21:9' default: '16:9' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' seed: type: integer minimum: 0 maximum: 4294967295 description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. required: - model - image_url title: gen4_turbo, runway/gen4_turbo - type: object properties: model: type: string enum: - gen4_aleph - runway/gen4_aleph prompt: type: string maxLength: 1000 description: The text description of the scene, subject, or action to generate in the video. video_url: type: string format: uri description: A HTTPS URL pointing to a video or a data URI containing a video. This video will be used as a reference during generation. references: type: array items: type: object properties: type: type: string enum: - image url: type: string format: uri required: - type - url description: Passing an image reference allows the model to emulate the style or content of the reference in the output. frame_size: type: string enum: - 1280:720 - 720:1280 - 1104:832 - 832:1104 - 960:960 - 1584:672 - 848:480 - 640:480 default: 1280:720 description: The width and height of the video. duration: type: number enum: - 5 default: 5 description: The length of the output video in seconds. seed: type: integer minimum: 0 maximum: 4294967295 description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. required: - model - prompt - video_url title: gen4_aleph, runway/gen4_aleph - type: object properties: model: type: string enum: - act_two - runway/act_two character: oneOf: - type: object properties: type: type: string enum: - video url: type: string format: uri required: - type - url description: A video of your character. In the output, the character will use the reference video performance in its original animated environment and some of the character's own movements. - type: object properties: type: type: string enum: - image url: type: string format: uri required: - type - url description: An image of your character. In the output, the character will use the reference video performance in its original static environment. description: The character to control. You can either provide a video or an image. A visually recognizable face must be visible and stay within the frame. reference: type: object properties: type: type: string enum: - video url: type: string format: uri required: - type - url description: Passing a video reference allows the model to emulate the style or content of the reference in the output. frame_size: type: string enum: - 1280:720 - 720:1280 - 1104:832 - 832:1104 - 960:960 - 1584:672 - 848:480 - 640:480 default: 1280:720 description: The width and height of the video. body_control: type: boolean description: A boolean indicating whether to enable body control. When enabled, non-facial movements and gestures will be applied to the character in addition to facial expressions. expression_intensity: type: integer minimum: 1 maximum: 5 default: 3 description: An integer between 1 and 5 (inclusive). A larger value increases the intensity of the character's expression. seed: type: integer minimum: 0 maximum: 4294967295 description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. required: - model - character - reference title: act_two, runway/act_two - type: object properties: model: type: string enum: - magic/text-to-video prompt: type: string description: Text that will appear in the video. template: type: string enum: - Shanghai Drone Show default: Shanghai Drone Show description: Video design template. required: - model - prompt title: magic/text-to-video - type: object properties: model: type: string enum: - magic/image-to-video image_url: type: string format: uri description: An image (supplied via URL or Base64) that will be inserted into the selected video template as the embedded ad content. template: type: string enum: - Thailand Street - Times Square Billboard - New York Times Square (77) - Phone Social - Art Gallery - New York Times Square (66) - Dubai Museum - Digital Float - Rotating Cards - Desktop Reveal - Egypt Pyramid - Frames Drop - Cappadocia Balloons - Times Square Round Screen - Stockholm Metro - Tokyo Billboard - San Francisco Skyscrapers - Malaysia Shop - Las Vegas LED - Phone App - Paris Eiffel Tower default: Thailand Street description: Video design template. required: - model - image_url title: magic/image-to-video - type: object properties: model: type: string enum: - magic/video-to-video video_url: type: string format: uri description: A video (supplied via URL or Base64) that will be inserted into the selected video template as the embedded ad content. template: type: string enum: - Thailand Street - Times Square Billboard - New York Times Square (78) - Phone Social - Art Gallery - New York Times Square (67) - Dubai Museum - Rotating Cards - Desktop Reveal - Egypt Pyramid - Cappadocia Balloons - Times Square Round Screen - Stockholm Metro - Tokyo Billboard - San Francisco Skyscrapers - Malaysia Shop - Las Vegas LED - Phone App - Paris Eiffel Tower default: Thailand Street description: Video design template. required: - model - video_url title: magic/video-to-video - type: object properties: model: type: string enum: - kling-video/v2.5-turbo/pro/image-to-video - kling-video/v2/master/image-to-video - kling-video/v2.1/master/image-to-video image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto prompt: type: string description: The text description of the scene, subject, or action to generate in the video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' negative_prompt: type: string description: The description of elements to avoid in the generated video. cfg_scale: type: number minimum: 0 maximum: 1 description: The CFG (Classifier Free Guidance) scale is a measure of how close you want the model to stick to your prompt. required: - model - image_url - prompt title: kling-video/v2.5-turbo/pro/image-to-video, kling-video/v2/master/image-to-video, kling-video/v2.1/master/image-to-video - type: object properties: model: type: string enum: - kling-video/v2.5-turbo/pro/text-to-video - kling-video/v2/master/text-to-video - kling-video/v2.1/master/text-to-video provider: type: string description: Provider routing override. `kling` (alias `klingai`) runs native Kling with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Kling -> fal.ai fallback chain. Case-insensitive. example: auto prompt: type: string description: The text description of the scene, subject, or action to generate in the video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' negative_prompt: type: string description: The description of elements to avoid in the generated video. cfg_scale: type: number minimum: 0 maximum: 1 description: The CFG (Classifier Free Guidance) scale is a measure of how close you want the model to stick to your prompt. aspect_ratio: type: string enum: - '16:9' - '9:16' - '1:1' description: The aspect ratio of the generated video. required: - model - prompt title: kling-video/v2.5-turbo/pro/text-to-video, kling-video/v2/master/text-to-video, kling-video/v2.1/master/text-to-video - type: object properties: model: type: string enum: - veo3.1 - veo3.1/fast prompt: type: string description: The text description of the scene, subject, or action to generate in the video. provider: type: string description: Provider routing override. `google` runs native Google with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Google -> fal.ai fallback chain. Case-insensitive. example: auto aspect_ratio: type: string enum: - '16:9' - '9:16' description: The aspect ratio of the generated video. resolution: type: string enum: - 720p - 1080p default: 1080p description: The resolution of the output video, where the number refers to the short side in pixels. duration: type: integer description: The length of the output video in seconds. enum: - 4 - 6 - 8 default: '8' generate_audio: type: boolean default: true description: Whether to generate audio for the video. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. auto_fix: type: boolean default: true description: Whether to automatically attempt to fix prompts that fail content policy or other validation checks by rewriting them. negative_prompt: type: string description: The description of elements to avoid in the generated video. enhance_prompt: type: boolean default: true description: Whether to enhance the video generation. required: - model - prompt title: veo3.1, veo3.1/fast - type: object properties: model: type: string enum: - veo3.1/image-to-video - veo3.1/fast/image-to-video prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: URL of the input image to animate. Should be 720p or higher resolution. provider: type: string description: Provider routing override. `google` runs native Google with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Google -> fal.ai fallback chain. Case-insensitive. example: auto aspect_ratio: type: string enum: - '16:9' - '9:16' description: The aspect ratio of the generated video. resolution: type: string enum: - 720p - 1080p default: 1080p description: The resolution of the output video, where the number refers to the short side in pixels. duration: type: integer description: The length of the output video in seconds. enum: - 4 - 6 - 8 default: '8' generate_audio: type: boolean default: true description: Whether to generate audio for the video. required: - model - prompt - image_url title: veo3.1/image-to-video, veo3.1/fast/image-to-video - type: object properties: model: type: string enum: - veo3.1/first-last-frame-to-video - veo3.1/fast/first-last-frame-to-video prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: URL of the input image to animate. Should be 720p or higher resolution. last_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. When this parameter is provided, the video duration can only be 8s. provider: type: string description: Provider routing override. `google` runs native Google with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Google -> fal.ai fallback chain. Case-insensitive. example: auto aspect_ratio: type: string enum: - '16:9' - '9:16' description: The aspect ratio of the generated video. resolution: type: string enum: - 720p - 1080p default: 1080p description: The resolution of the output video, where the number refers to the short side in pixels. duration: type: integer description: The length of the output video in seconds. enum: - 4 - 6 - 8 default: '8' generate_audio: type: boolean default: true description: Whether to generate audio for the video. required: - model - prompt - image_url - last_image_url title: veo3.1/first-last-frame-to-video, veo3.1/fast/first-last-frame-to-video - type: object properties: model: type: string enum: - veo3.1/extend-video - veo3.1/fast/extend-video prompt: type: string description: The text description of the scene, subject, or action to generate in the video. video_url: type: string format: uri description: A HTTPS URL pointing to a video or a data URI containing a video. This video will be used as a reference during generation. provider: type: string description: Provider routing override. `google` runs native Google with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Google -> fal.ai fallback chain. Case-insensitive. example: auto aspect_ratio: type: string enum: - auto - '16:9' - '9:16' default: auto description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 7 default: '7' resolution: type: string enum: - 720p default: 720p description: The resolution of the output video, where the number refers to the short side in pixels. generate_audio: type: boolean default: true description: Whether to generate audio for the video. auto_fix: type: boolean default: false description: Whether to automatically attempt to fix prompts that fail content policy or other validation checks by rewriting them. required: - model - prompt - video_url title: veo3.1/extend-video, veo3.1/fast/extend-video - type: object properties: model: type: string enum: - veo3.1/lite/image-to-video - veo3.1/lite/first-last-frame-to-video - veo3.1/lite prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the first frame of the video. Should be 720p or higher resolution. last_image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image to be used as the last frame of the video. Should be 720p or higher resolution. When this parameter is provided, the video duration can only be 8s. provider: type: string description: Provider routing override. `google` runs native Google with no fallback; `fal` runs the fal.ai mirror; `auto` (default) uses the Google -> fal.ai fallback chain. Case-insensitive. example: auto aspect_ratio: type: string enum: - '16:9' - '9:16' description: The aspect ratio of the generated video. resolution: type: string enum: - 720p - 1080p default: 1080p description: The resolution of the output video, where the number refers to the short side in pixels. duration: type: integer description: The length of the output video in seconds. enum: - 4 - 6 - 8 default: '8' generate_audio: type: boolean default: true description: Whether to generate audio for the video. required: - model - prompt title: veo3.1/lite/image-to-video, veo3.1/lite/first-last-frame-to-video, veo3.1/lite - type: object properties: model: type: string enum: - bytedance/omnihuman - bytedance/omnihuman/v1.5 image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. audio_url: type: string format: uri description: The URL of the audio file for lip-sync animation. The model detects spoken parts and syncs the character's mouth to them. Audio must be under 30s long. required: - model - image_url - audio_url title: bytedance/omnihuman, bytedance/omnihuman/v1.5 - type: object properties: model: type: string enum: - hunyuan-video-foley - tencent/hunyuan-video-foley video_url: type: string format: uri description: A HTTPS URL pointing to a video or a data URI containing a video. This video will be used as a reference during generation. prompt: type: string description: The text description of the scene, subject, or action to generate in the video. negative_prompt: type: string default: noisy, harsh description: The description of elements to avoid in the generated video. guidance_scale: type: number default: 4.5 description: Classifier-free guidance scale. Controls prompt adherence / creativity. num_inference_steps: type: integer default: 50 description: Number of inference steps for sampling. Higher values give better quality but take longer. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. required: - model - video_url title: hunyuan-video-foley, tencent/hunyuan-video-foley - type: object properties: model: type: string enum: - kandinsky5/text-to-video - kandinsky5/text-to-video/distill - sber-ai/kandinsky5-t2v - sber-ai/kandinsky5-distill-t2v prompt: type: string description: The text description of the scene, subject, or action to generate in the video. aspect_ratio: type: string enum: - '3:2' - '1:1' - '2:3' default: '3:2' description: The aspect ratio of the generated video. duration: type: integer description: The length of the output video in seconds. enum: - 5 - 10 default: '5' num_inference_steps: type: integer default: 30 description: Number of inference steps for sampling. Higher values give better quality but take longer. required: - model - prompt title: kandinsky5/text-to-video, kandinsky5/text-to-video/distill, sber-ai/kandinsky5-t2v, sber-ai/kandinsky5-distill-t2v - type: object properties: model: type: string enum: - krea-wan-14b/text-to-video - krea/krea-wan-14b/text-to-video prompt: type: string description: The text description of the scene, subject, or action to generate in the video. num_frames: type: integer minimum: 18 maximum: 162 default: 78 description: Number of frames to generate. Must be a multiple of 12 plus 6, for example 18, 30, 42, etc. enable_prompt_expansion: type: boolean default: true description: Whether to enable prompt expansion. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. required: - model - prompt title: krea-wan-14b/text-to-video, krea/krea-wan-14b/text-to-video - type: object properties: model: type: string enum: - krea-wan-14b/video-to-video - krea/krea-wan-14b/video-to-video prompt: type: string description: The text description of the scene, subject, or action to generate in the video. video_url: type: string format: uri description: A HTTPS URL pointing to a video or a data URI containing a video. This video will be used as a reference during generation. strength: type: number minimum: 0 maximum: 1 default: 0.85 description: Denoising strength for the video-to-video generation. 0.0 preserves the original, 1.0 completely remakes the video. enable_prompt_expansion: type: boolean default: true description: Whether to enable prompt expansion. seed: type: integer description: Varying the seed integer is a way to get different results for the same other request parameters. Using the same value for an identical request will produce similar results. If unspecified, a random number is chosen. required: - model - prompt - video_url title: krea-wan-14b/video-to-video, krea/krea-wan-14b/video-to-video - type: object properties: model: type: string enum: - ltxv-2 - ltxv-2/image-to-video - ltxv-2/text-to-video - ltxv-2/fast - ltxv-2/image-to-video/fast - ltxv-2/text-to-video/fast - ltxv/ltxv-2 - ltxv/ltxv-2-fast prompt: type: string description: The text description of the scene, subject, or action to generate in the video. image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. duration: type: integer description: The length of the output video in seconds. enum: - 6 - 8 - 10 resolution: type: string enum: - 1080p - 1440p - 2160p default: 1080p description: The resolution of the output video, where the number refers to the short side in pixels. aspect_ratio: type: string enum: - '16:9' default: '16:9' description: The aspect ratio of the generated video. fps: type: integer description: Frames per second of the generated video. enum: - 25 - 50 generate_audio: type: boolean default: true description: Whether to generate audio for the video. required: - model - prompt title: ltxv-2, ltxv-2/image-to-video, ltxv-2/text-to-video, ltxv-2/fast, ltxv-2/image-to-video/fast, ltxv-2/text-to-video/fast, ltxv/ltxv-2, ltxv/ltxv-2-fast - type: object properties: model: type: string enum: - veed/fabric-1.0 - veed/fabric-1.0/fast image_url: type: string format: uri description: A direct link to an online image or a Base64-encoded local image that will serve as the visual base or the first frame for the video. audio_url: type: string format: uri description: The URL of the audio file for lip-sync animation. The model detects spoken parts and syncs the character's mouth to them. Audio must be under 30s long. resolution: type: string enum: - 480p - 720p default: 480p description: The resolution of the generated video. required: - model - image_url - audio_url title: veed/fabric-1.0, veed/fabric-1.0/fast - type: object properties: model: type: string enum: - blackforestlabs/flux-3-video-t2v prompt: type: string minLength: 1 description: The text prompt describing the video to generate. aspect_ratio: type: string enum: - auto - '21:9' - '2:1' - '16:9' - '4:3' - '1:1' - '3:4' - '9:16' default: auto description: The aspect ratio of the generated video. `auto` lets the model pick, and image-to-video follows the first keyframe. duration: type: integer minimum: 5 maximum: 20 default: 5 description: The video length in whole seconds, from 5 to 20. generate_audio: type: boolean default: true description: Whether to generate a synchronized audio track. Set to `false` for a silent clip. safety_tolerance: type: integer minimum: 0 maximum: 4 default: 2 description: The moderation strictness, from 0 (strictest) to 4 (most permissive). resolution: type: string enum: - hd - fhd default: hd description: The output resolution of the generated video. required: - model - prompt title: blackforestlabs/flux-3-video-t2v - type: object properties: model: type: string enum: - blackforestlabs/flux-3-video-i2v prompt: type: string minLength: 1 description: The text prompt describing the video to generate. aspect_ratio: type: string enum: - auto - '21:9' - '2:1' - '16:9' - '4:3' - '1:1' - '3:4' - '9:16' default: auto description: The aspect ratio of the generated video. `auto` lets the model pick, and image-to-video follows the first keyframe. duration: type: integer minimum: 5 maximum: 20 default: 5 description: The video length in whole seconds, from 5 to 20. generate_audio: type: boolean default: true description: Whether to generate a synchronized audio track. Set to `false` for a silent clip. safety_tolerance: type: integer minimum: 0 maximum: 4 default: 2 description: The moderation strictness, from 0 (strictest) to 4 (most permissive). image_url: type: string format: uri description: The image URL that starts the clip. last_image_url: type: string format: uri description: An optional image URL that pins the last frame of the clip. Provide it together with `image_url` to interpolate between the two. image_urls: type: array items: type: string format: uri maxItems: 10 description: Optional additional image URLs placed between the first and the last keyframe. resolution: type: string enum: - hd - fhd default: hd description: The output resolution of the generated video. required: - model - prompt - image_url title: blackforestlabs/flux-3-video-i2v - type: object properties: model: type: string enum: - blackforestlabs/flux-3-video-v2v prompt: type: string minLength: 1 description: The text prompt describing the video to generate. aspect_ratio: type: string enum: - auto - '21:9' - '2:1' - '16:9' - '4:3' - '1:1' - '3:4' - '9:16' default: auto description: The aspect ratio of the generated video. `auto` lets the model pick, and image-to-video follows the first keyframe. duration: type: integer minimum: 5 maximum: 20 default: 5 description: The video length in whole seconds, from 5 to 20. generate_audio: type: boolean default: true description: Whether to generate a synchronized audio track. Set to `false` for a silent clip. safety_tolerance: type: integer minimum: 0 maximum: 4 default: 2 description: The moderation strictness, from 0 (strictest) to 4 (most permissive). video_url: type: string format: uri description: The URL of the video to continue. The generated clip carries on from its final frames. resolution: type: string enum: - hd - fhd default: hd description: The output resolution of the generated video. required: - model - prompt - video_url title: blackforestlabs/flux-3-video-v2v - type: object properties: model: type: string enum: - blackforestlabs/flux-3-video-draft-t2v prompt: type: string minLength: 1 description: The text prompt describing the video to generate. aspect_ratio: type: string enum: - auto - '21:9' - '2:1' - '16:9' - '4:3' - '1:1' - '3:4' - '9:16' default: auto description: The aspect ratio of the generated video. `auto` lets the model pick, and image-to-video follows the first keyframe. duration: type: integer minimum: 5 maximum: 20 default: 5 description: The video length in whole seconds, from 5 to 20. generate_audio: type: boolean default: true description: Whether to generate a synchronized audio track. Set to `false` for a silent clip. safety_tolerance: type: integer minimum: 0 maximum: 4 default: 2 description: The moderation strictness, from 0 (strictest) to 4 (most permissive). resolution: type: string enum: - hd default: hd description: The output resolution of the generated video. Draft generations are hd-only; use the non-draft model for `fhd`. required: - model - prompt title: blackforestlabs/flux-3-video-draft-t2v - type: object properties: model: type: string enum: - blackforestlabs/flux-3-video-draft-i2v prompt: type: string minLength: 1 description: The text prompt describing the video to generate. aspect_ratio: type: string enum: - auto - '21:9' - '2:1' - '16:9' - '4:3' - '1:1' - '3:4' - '9:16' default: auto description: The aspect ratio of the generated video. `auto` lets the model pick, and image-to-video follows the first keyframe. duration: type: integer minimum: 5 maximum: 20 default: 5 description: The video length in whole seconds, from 5 to 20. generate_audio: type: boolean default: true description: Whether to generate a synchronized audio track. Set to `false` for a silent clip. safety_tolerance: type: integer minimum: 0 maximum: 4 default: 2 description: The moderation strictness, from 0 (strictest) to 4 (most permissive). image_url: type: string format: uri description: The image URL that starts the clip. last_image_url: type: string format: uri description: An optional image URL that pins the last frame of the clip. Provide it together with `image_url` to interpolate between the two. image_urls: type: array items: type: string format: uri maxItems: 10 description: Optional additional image URLs placed between the first and the last keyframe. resolution: type: string enum: - hd default: hd description: The output resolution of the generated video. Draft generations are hd-only; use the non-draft model for `fhd`. required: - model - prompt - image_url title: blackforestlabs/flux-3-video-draft-i2v - type: object properties: model: type: string enum: - blackforestlabs/flux-3-video-draft-v2v prompt: type: string minLength: 1 description: The text prompt describing the video to generate. aspect_ratio: type: string enum: - auto - '21:9' - '2:1' - '16:9' - '4:3' - '1:1' - '3:4' - '9:16' default: auto description: The aspect ratio of the generated video. `auto` lets the model pick, and image-to-video follows the first keyframe. duration: type: integer minimum: 5 maximum: 20 default: 5 description: The video length in whole seconds, from 5 to 20. generate_audio: type: boolean default: true description: Whether to generate a synchronized audio track. Set to `false` for a silent clip. safety_tolerance: type: integer minimum: 0 maximum: 4 default: 2 description: The moderation strictness, from 0 (strictest) to 4 (most permissive). video_url: type: string format: uri description: The URL of the video to continue. The generated clip carries on from its final frames. resolution: type: string enum: - hd default: hd description: The output resolution of the generated video. Draft generations are hd-only; use the non-draft model for `fhd`. required: - model - prompt - video_url title: blackforestlabs/flux-3-video-draft-v2v - type: object properties: model: type: string enum: - beeble/switchx-video-to-video - switchx-video-to-video video_url: type: string format: uri description: 'URL of the source video to recomposite. Allowed formats: MP4, MOV (H.264 or HEVC), up to 240 frames. Source must not exceed 2,770,000 total pixels.' alpha_url: type: string format: uri description: URL of the alpha mask. Required when alpha_mode is "custom" or "select"; ignored for "auto" and "fill". A mask video for video-to-video, a mask image for image-to-image. reference_image_url: type: string format: uri description: 'URL of the reference image defining the target look and lighting for the replaced region. Allowed formats: JPEG, PNG, WEBP. At least one of reference_image_url or prompt must be provided.' alpha_mode: type: string enum: - auto - fill - custom - select default: auto description: 'Subject masking strategy: "auto" (AI auto-detects the subject), "fill" (no masking), "select" (propagate a mask from one keyframe), or "custom" (frame-by-frame mask supplied via alpha_url). Defaults to "auto".' max_resolution: type: integer description: 'Maximum output resolution (longer side) in pixels: 720 or 1080. Defaults to 1080. 1080 costs more than 720.' enum: - 720 - 1080 default: '1080' alpha_keyframe_index: type: integer minimum: 0 description: Frame index (0-based) whose mask is propagated when alpha_mode is "select" on a video. Defaults to the first frame. Ignored for image generation and for the auto, fill, and custom modes. seed: type: integer minimum: 0 maximum: 4294967295 description: Random seed (0–4294967295) for reproducibility. Omit for a random seed. The seed used is always returned in the response. prompt: type: string maxLength: 2000 description: Text description of the desired output to guide the generation (max 2000 characters). At least one of prompt or reference_image_url must be provided. required: - model - video_url title: beeble/switchx-video-to-video, switchx-video-to-video responses: '200': content: application/json: schema: type: object properties: id: type: string description: The ID of the generated video. example: 60ac7c34-3224-4b14-8e7d-0aa0db708325 status: type: string enum: - queued - generating - completed - error description: The current status of the generation task. example: completed video: anyOf: - type: object properties: url: type: string format: uri description: The URL where the file can be downloaded from. example: https://cdn.aimlapi.com/generations/hedgehog/1759866285599-0cdfb138-c03a-49d4-a601-4f6413e27b15.mp4 required: - url - type: array items: type: object properties: url: type: string format: uri description: The URL where the file can be downloaded from. example: https://cdn.aimlapi.com/generations/hedgehog/1759866285599-0cdfb138-c03a-49d4-a601-4f6413e27b15.mp4 required: - url - {} error: type: - object - 'null' properties: name: type: string message: type: string required: - name - message description: Description of the error, if any. meta: type: - object - 'null' properties: usage: type: - object - 'null' properties: credits_used: type: number description: The number of tokens consumed during generation. example: 120000 usd_spent: type: number description: The total amount of money spent by the user in USD. example: 0.06 required: - credits_used - usd_spent description: Additional details about the generation. required: - id - status tags: - Video summary: V2 video generations x-summary-source: derived get: operationId: getV2VideoGenerations requestBody: required: true content: application/json: schema: type: object properties: model: type: string enum: - sora-2-t2v - openai/sora-2-t2v - sora-2 - openai/sora-2 - sora-2-i2v - openai/sora-2-i2v - sora-2-pro-t2v - openai/sora-2-pro-t2v - sora-2-pro - openai/sora-2-pro - sora-2-pro-i2v - openai/sora-2-pro-i2v - bytedance/seedance-1-0-pro - bytedance/seedance-1-0-pro-fast - bytedance/seedance-1-5-pro - bytedance/dreamina-seedance-2-0 - bytedance/dreamina-seedance-2-0-fast - bytedance/dreamina-seedance-2-0-mini - bytedance/dreamina-seedance-2-5 - bytedance/seedance-1-0-pro-t2v - bytedance/seedance-1-0-pro-i2v - bytedance/seedance-2-0-fast - bytedance/seedance-2.0/fast/image-to-video - bytedance/seedance-2.0/fast/reference-to-video - bytedance/seedance-2.0/fast/text-to-video - bytedance/seedance-2-0 - bytedance/seedance-2.0/image-to-video - bytedance/seedance-2.0/reference-to-video - bytedance/seedance-2.0/text-to-video - bytedance/seedance-2-0-mini - bytedance/seedance-2-5 - bytedance/seedance-2.5 - veo-2.0-generate-001 - google/veo-2.0-generate-001 - veo-3.0-fast-generate-001 - google/veo-3.0-fast-generate-001 - veo-3.0-generate-001 - google/veo-3.0-generate-001 - veo-3.1-lite-generate-001 - google/veo-3.1-lite-generate-001 - veo-3.1-generate-001 - google/veo-3.1-generate-001 - veo-3.1-fast-generate-001 - google/veo-3.1-fast-generate-001 - gemini-omni-flash-preview - google/gemini-omni-flash-preview - gemini-omni-1.1-flash - google/gemini-omni-1.1-flash - google/veo-3.1-t2v - google/veo-3.1-i2v - google/veo-3.1-t2v-fast - google/veo-3.1-i2v-fast - google/veo-3.1-first-last-image-to-video - google/veo-3.1-reference-to-video - google/veo-3.1-first-last-image-to-video-fast - google/veo3-1-extend-video - google/veo3-1-fast-extend-video - veo2 - google/veo2 - veo2/image-to-video - google/veo2-image-to-video - google/veo3 - google/veo-3.0-i2v - google/veo-3.0-fast - google/veo-3.0-i2v-fast - google/veo-3-1-lite-generate-preview - wan2.1-t2v-plus - alibaba/wan2.1-t2v-plus - wan2.1-t2v-turbo - alibaba/wan2.1-t2v-turbo - wan2.2-i2v-plus - alibaba/wan2.2-i2v-plus - wan2.2-t2v-plus - alibaba/wan2.2-t2v-plus - wan2.5-t2v-preview - alibaba/wan2.5-t2v-preview - wan2.5-i2v-preview - alibaba/wan2.5-i2v-preview - wan2.6-t2v - alibaba/wan2.6-t2v - wan2.6-i2v - alibaba/wan2.6-i2v - wan2.6-i2v-flash - alibaba/wan2.6-i2v-flash - wan2.6-r2v - alibaba/wan2.6-r2v - wan2.7-t2v - alibaba/wan2.7-t2v - wan2.7-i2v - alibaba/wan2.7-i2v - wan2.7-r2v - alibaba/wan2.7-r2v - wan3.0-video - alibaba/wan3.0-video - custom:happyhorse-1.0 - alibaba/custom:happyhorse-1.0 - happyhorse-1.0-t2v - alibaba/happyhorse-1.0-t2v - happyhorse-1.0-i2v - alibaba/happyhorse-1.0-i2v - happyhorse-1.0-r2v - alibaba/happyhorse-1.0-r2v - happyhorse-1.0-video-edit - alibaba/happyhorse-1.0-video-edit - happyhorse-1.1-t2v - alibaba/happyhorse-1.1-t2v - happyhorse-1.1-i2v - alibaba/happyhorse-1.1-i2v - happyhorse-1.1-r2v - alibaba/happyhorse-1.1-r2v - custom:happyhorse-1.1 - alibaba/custom:happyhorse-1.1 - alibaba/wan2.2-vace-fun-a14b-depth - alibaba/wan2.2-vace-fun-a14b-pose - alibaba/wan2.2-vace-fun-a14b-inpainting - alibaba/wan2.2-vace-fun-a14b-outpainting - alibaba/wan2.2-vace-fun-a14b-reframe - alibaba/wan2.2-14b-animate-move - alibaba/wan2.2-14b-animate-replace - alibaba/wan-2-6-t2v - alibaba/wan-2-6-i2v - alibaba/wan-2-6-image-to-video-flash - alibaba/wan-2-6-r2v - alibaba/wan-2-7-t2v - alibaba/wan-2-7-i2v - alibaba/wan-2-7-r2v - alibaba/wan-3-0-video - alibaba/happyhorse-1-0 - alibaba/happyhorse-1-1 - alibaba/happyhorse-1-1-t2v - alibaba/happyhorse-1-1-i2v - alibaba/happyhorse-1-1-r2v - test/dummy-video - video-01 - video-01-live2d - minimax/hailuo-02 - minimax/hailuo-2.3 - minimax/hailuo-2.3-fast - minimax/h3 - xai/grok-imagine-video - x-ai/grok-imagine-video - xai/grok-imagine-video-1.5-preview - x-ai/grok-imagine-video-1.5-preview - kling-video/v1/standard/image-to-video - kling-video/v1/standard/text-to-video - kling-video/v1/pro/image-to-video - kling-video/v1/pro/text-to-video - kling-video/v1.5/pro/image-to-video - kling-video/v1.5/pro/text-to-video - kling-video/v1.6/standard/text-to-video - kling-video/v1.6/standard/image-to-video - kling-video/v1.6/pro/image-to-video - kling-video/v1.6/pro/text-to-video - klingai/kling-video-v1.6-pro-effects - klingai/kling-video-v1.6-standard-effects - kling-video/v1.6/standard/multi-image-to-video - klingai/v2-master-image-to-video - klingai/v2-master-text-to-video - kling-video/v2.1/standard/image-to-video - klingai/v2.1-master-image-to-video - klingai/v2.1-master-text-to-video - kling-video/v2.1/pro/image-to-video - klingai/v2.5-turbo/pro/image-to-video - klingai/v2.5-turbo/pro/text-to-video - klingai/avatar-standard - klingai/avatar-pro - klingai/video-o1-image-to-video - klingai/video-o1-reference-to-video - klingai/video-o1-video-to-video-edit - klingai/video-o1-video-to-video-reference - klingai/video-v2-6-pro-image-to-video - klingai/video-v2-6-pro-text-to-video - klingai/video-v2-6-pro-motion-control - klingai/video-v2-6-motion-control - klingai/video-v3-standard-text-to-video - klingai/video-v3-standard-image-to-video - klingai/video-v3-pro-text-to-video - klingai/video-v3-pro-image-to-video - klingai/video-v3-omni-720p-text-to-video - klingai/video-v3-omni-720p-image-to-video - klingai/video-v3-omni-1080p-text-to-video - klingai/video-v3-omni-1080p-image-to-video - klingai/video-v3-standard-turbo-text-to-video - klingai/video-v3-standard-turbo-image-to-video - klingai/video-v3-turbo-pro-text-to-video - klingai/video-v3-turbo-pro-image-to-video - klingai/video-v3-motion-control - pixverse/v5/text-to-video - pixverse/v5/image-to-video - pixverse/v5/transition - pixverse/lip-sync - pixverse/v5.5/text-to-video - pixverse/v5-5-text-to-video - pixverse/v5.5/image-to-video - pixverse/v5-5-image-to-video - ray-2 - luma/ray-2 - ray-flash-2 - luma/ray-flash-2 - luma/ray-3.2 - gen3a_turbo - runway/gen3a_turbo - gen4_turbo - runway/gen4_turbo - gen4_aleph - runway/gen4_aleph - act_two - runway/act_two - magic/text-to-video - magic/image-to-video - magic/video-to-video - kling-video/v2.5-turbo/pro/image-to-video - kling-video/v2.5-turbo/pro/text-to-video - kling-video/v1/standard/ai-avatar - kling-video/v1/pro/ai-avatar - kling-video/o1/image-to-video - kling-video/o1/reference-to-video - kling-video/o1/video-to-video/edit - kling-video/o1/video-to-video/reference - kling-video/v2.6/pro/text-to-video - kling-video/v2.6/pro/image-to-video - kling-video/v2/master/image-to-video - kling-video/v2/master/text-to-video - kling-video/v2.1/master/image-to-video - kling-video/v2.1/master/text-to-video - kling-video/v3/standard/text-to-video - kling-video/v3/standard/image-to-video - kling-video/v3/pro/text-to-video - kling-video/v3/pro/image-to-video - kling-video/v3/turbo/standard/text-to-video - kling-video/v3/turbo/standard/image-to-video - kling-video/v3/turbo/pro/text-to-video - kling-video/v3/turbo/pro/image-to-video - veo3.1 - veo3.1/image-to-video - veo3.1/first-last-frame-to-video - veo3.1/reference-to-video - veo3.1/fast - veo3.1/fast/image-to-video - veo3.1/fast/first-last-frame-to-video - veo3.1/extend-video - veo3.1/fast/extend-video - veo3.1/lite/image-to-video - veo3.1/lite/first-last-frame-to-video - veo3.1/lite - bytedance/omnihuman - bytedance/omnihuman/v1.5 - hunyuan-video-foley - wan-22-vace-fun-a14b/depth - wan-22-vace-fun-a14b/pose - wan-22-vace-fun-a14b/inpainting - wan-22-vace-fun-a14b/outpainting - wan-22-vace-fun-a14b/reframe - wan/v2.2-14b/animate/move - wan/v2.2-14b/animate/replace - kandinsky5/text-to-video - kandinsky5/text-to-video/distill - krea-wan-14b/text-to-video - krea-wan-14b/video-to-video - ltxv-2 - ltxv-2/image-to-video - ltxv-2/text-to-video - ltxv-2/fast - ltxv-2/image-to-video/fast - ltxv-2/text-to-video/fast - veed/fabric-1.0 - veed/fabric-1.0/fast - tencent/hunyuan-video-foley - sber-ai/kandinsky5-t2v - sber-ai/kandinsky5-distill-t2v - krea/krea-wan-14b/text-to-video - krea/krea-wan-14b/video-to-video - ltxv/ltxv-2 - ltxv/ltxv-2-fast - blackforestlabs/flux-3-video-t2v - blackforestlabs/flux-3-video-i2v - blackforestlabs/flux-3-video-v2v - blackforestlabs/flux-3-video-draft-t2v - blackforestlabs/flux-3-video-draft-i2v - blackforestlabs/flux-3-video-draft-v2v - beeble/switchx-video-to-video - switchx-video-to-video id: type: string required: - model - id title: sora-2-t2v, openai/sora-2-t2v, sora-2, openai/sora-2, sora-2-i2v, openai/sora-2-i2v, sora-2-pro-t2v, openai/sora-2-pro-t2v, sora-2-pro, openai/sora-2-pro, sora-2-pro-i2v, openai/sora-2-pro-i2v, bytedance/seedance-1-0-pro, bytedance/seedance-1-0-pro-fast, bytedance/seedance-1-5-pro, bytedance/dreamina-seedance-2-0, bytedance/dreamina-seedance-2-0-fast, bytedance/dreamina-seedance-2-0-mini, bytedance/dreamina-seedance-2-5, bytedance/seedance-1-0-pro-t2v, bytedance/seedance-1-0-pro-i2v, bytedance/seedance-2-0-fast, bytedance/seedance-2.0/fast/image-to-video, bytedance/seedance-2.0/fast/reference-to-video, bytedance/seedance-2.0/fast/text-to-video, bytedance/seedance-2-0, bytedance/seedance-2.0/image-to-video, bytedance/seedance-2.0/reference-to-video, bytedance/seedance-2.0/text-to-video, bytedance/seedance-2-0-mini, bytedance/seedance-2-5, bytedance/seedance-2.5, veo-2.0-generate-001, google/veo-2.0-generate-001, veo-3.0-fast-generate-001, google/veo-3.0-fast-generate-001, veo-3.0-generate-001, google/veo-3.0-generate-001, veo-3.1-lite-generate-001, google/veo-3.1-lite-generate-001, veo-3.1-generate-001, google/veo-3.1-generate-001, veo-3.1-fast-generate-001, google/veo-3.1-fast-generate-001, gemini-omni-flash-preview, google/gemini-omni-flash-preview, gemini-omni-1.1-flash, google/gemini-omni-1.1-flash, google/veo-3.1-t2v, google/veo-3.1-i2v, google/veo-3.1-t2v-fast, google/veo-3.1-i2v-fast, google/veo-3.1-first-last-image-to-video, google/veo-3.1-reference-to-video, google/veo-3.1-first-last-image-to-video-fast, google/veo3-1-extend-video, google/veo3-1-fast-extend-video, veo2, google/veo2, veo2/image-to-video, google/veo2-image-to-video, google/veo3, google/veo-3.0-i2v, google/veo-3.0-fast, google/veo-3.0-i2v-fast, google/veo-3-1-lite-generate-preview, wan2.1-t2v-plus, alibaba/wan2.1-t2v-plus, wan2.1-t2v-turbo, alibaba/wan2.1-t2v-turbo, wan2.2-i2v-plus, alibaba/wan2.2-i2v-plus, wan2.2-t2v-plus, alibaba/wan2.2-t2v-plus, wan2.5-t2v-preview, alibaba/wan2.5-t2v-preview, wan2.5-i2v-preview, alibaba/wan2.5-i2v-preview, wan2.6-t2v, alibaba/wan2.6-t2v, wan2.6-i2v, alibaba/wan2.6-i2v, wan2.6-i2v-flash, alibaba/wan2.6-i2v-flash, wan2.6-r2v, alibaba/wan2.6-r2v, wan2.7-t2v, alibaba/wan2.7-t2v, wan2.7-i2v, alibaba/wan2.7-i2v, wan2.7-r2v, alibaba/wan2.7-r2v, wan3.0-video, alibaba/wan3.0-video, custom:happyhorse-1.0, alibaba/custom:happyhorse-1.0, happyhorse-1.0-t2v, alibaba/happyhorse-1.0-t2v, happyhorse-1.0-i2v, alibaba/happyhorse-1.0-i2v, happyhorse-1.0-r2v, alibaba/happyhorse-1.0-r2v, happyhorse-1.0-video-edit, alibaba/happyhorse-1.0-video-edit, happyhorse-1.1-t2v, alibaba/happyhorse-1.1-t2v, happyhorse-1.1-i2v, alibaba/happyhorse-1.1-i2v, happyhorse-1.1-r2v, alibaba/happyhorse-1.1-r2v, custom:happyhorse-1.1, alibaba/custom:happyhorse-1.1, alibaba/wan2.2-vace-fun-a14b-depth, alibaba/wan2.2-vace-fun-a14b-pose, alibaba/wan2.2-vace-fun-a14b-inpainting, alibaba/wan2.2-vace-fun-a14b-outpainting, alibaba/wan2.2-vace-fun-a14b-reframe, alibaba/wan2.2-14b-animate-move, alibaba/wan2.2-14b-animate-replace, alibaba/wan-2-6-t2v, alibaba/wan-2-6-i2v, alibaba/wan-2-6-image-to-video-flash, alibaba/wan-2-6-r2v, alibaba/wan-2-7-t2v, alibaba/wan-2-7-i2v, alibaba/wan-2-7-r2v, alibaba/wan-3-0-video, alibaba/happyhorse-1-0, alibaba/happyhorse-1-1, alibaba/happyhorse-1-1-t2v, alibaba/happyhorse-1-1-i2v, alibaba/happyhorse-1-1-r2v, test/dummy-video, video-01, video-01-live2d, minimax/hailuo-02, minimax/hailuo-2.3, minimax/hailuo-2.3-fast, minimax/h3, xai/grok-imagine-video, x-ai/grok-imagine-video, xai/grok-imagine-video-1.5-preview, x-ai/grok-imagine-video-1.5-preview, kling-video/v1/standard/image-to-video, kling-video/v1/standard/text-to-video, kling-video/v1/pro/image-to-video, kling-video/v1/pro/text-to-video, kling-video/v1.5/pro/image-to-video, kling-video/v1.5/pro/text-to-video, kling-video/v1.6/standard/text-to-video, kling-video/v1.6/standard/image-to-video, kling-video/v1.6/pro/image-to-video, kling-video/v1.6/pro/text-to-video, klingai/kling-video-v1.6-pro-effects, klingai/kling-video-v1.6-standard-effects, kling-video/v1.6/standard/multi-image-to-video, klingai/v2-master-image-to-video, klingai/v2-master-text-to-video, kling-video/v2.1/standard/image-to-video, klingai/v2.1-master-image-to-video, klingai/v2.1-master-text-to-video, kling-video/v2.1/pro/image-to-video, klingai/v2.5-turbo/pro/image-to-video, klingai/v2.5-turbo/pro/text-to-video, klingai/avatar-standard, klingai/avatar-pro, klingai/video-o1-image-to-video, klingai/video-o1-reference-to-video, klingai/video-o1-video-to-video-edit, klingai/video-o1-video-to-video-reference, klingai/video-v2-6-pro-image-to-video, klingai/video-v2-6-pro-text-to-video, klingai/video-v2-6-pro-motion-control, klingai/video-v2-6-motion-control, klingai/video-v3-standard-text-to-video, klingai/video-v3-standard-image-to-video, klingai/video-v3-pro-text-to-video, klingai/video-v3-pro-image-to-video, klingai/video-v3-omni-720p-text-to-video, klingai/video-v3-omni-720p-image-to-video, klingai/video-v3-omni-1080p-text-to-video, klingai/video-v3-omni-1080p-image-to-video, klingai/video-v3-standard-turbo-text-to-video, klingai/video-v3-standard-turbo-image-to-video, klingai/video-v3-turbo-pro-text-to-video, klingai/video-v3-turbo-pro-image-to-video, klingai/video-v3-motion-control, pixverse/v5/text-to-video, pixverse/v5/image-to-video, pixverse/v5/transition, pixverse/lip-sync, pixverse/v5.5/text-to-video, pixverse/v5-5-text-to-video, pixverse/v5.5/image-to-video, pixverse/v5-5-image-to-video, ray-2, luma/ray-2, ray-flash-2, luma/ray-flash-2, luma/ray-3.2, gen3a_turbo, runway/gen3a_turbo, gen4_turbo, runway/gen4_turbo, gen4_aleph, runway/gen4_aleph, act_two, runway/act_two, magic/text-to-video, magic/image-to-video, magic/video-to-video, kling-video/v2.5-turbo/pro/image-to-video, kling-video/v2.5-turbo/pro/text-to-video, kling-video/v1/standard/ai-avatar, kling-video/v1/pro/ai-avatar, kling-video/o1/image-to-video, kling-video/o1/reference-to-video, kling-video/o1/video-to-video/edit, kling-video/o1/video-to-video/reference, kling-video/v2.6/pro/text-to-video, kling-video/v2.6/pro/image-to-video, kling-video/v2/master/image-to-video, kling-video/v2/master/text-to-video, kling-video/v2.1/master/image-to-video, kling-video/v2.1/master/text-to-video, kling-video/v3/standard/text-to-video, kling-video/v3/standard/image-to-video, kling-video/v3/pro/text-to-video, kling-video/v3/pro/image-to-video, kling-video/v3/turbo/standard/text-to-video, kling-video/v3/turbo/standard/image-to-video, kling-video/v3/turbo/pro/text-to-video, kling-video/v3/turbo/pro/image-to-video, veo3.1, veo3.1/image-to-video, veo3.1/first-last-frame-to-video, veo3.1/reference-to-video, veo3.1/fast, veo3.1/fast/image-to-video, veo3.1/fast/first-last-frame-to-video, veo3.1/extend-video, veo3.1/fast/extend-video, veo3.1/lite/image-to-video, veo3.1/lite/first-last-frame-to-video, veo3.1/lite, bytedance/omnihuman, bytedance/omnihuman/v1.5, hunyuan-video-foley, wan-22-vace-fun-a14b/depth, wan-22-vace-fun-a14b/pose, wan-22-vace-fun-a14b/inpainting, wan-22-vace-fun-a14b/outpainting, wan-22-vace-fun-a14b/reframe, wan/v2.2-14b/animate/move, wan/v2.2-14b/animate/replace, kandinsky5/text-to-video, kandinsky5/text-to-video/distill, krea-wan-14b/text-to-video, krea-wan-14b/video-to-video, ltxv-2, ltxv-2/image-to-video, ltxv-2/text-to-video, ltxv-2/fast, ltxv-2/image-to-video/fast, ltxv-2/text-to-video/fast, veed/fabric-1.0, veed/fabric-1.0/fast, tencent/hunyuan-video-foley, sber-ai/kandinsky5-t2v, sber-ai/kandinsky5-distill-t2v, krea/krea-wan-14b/text-to-video, krea/krea-wan-14b/video-to-video, ltxv/ltxv-2, ltxv/ltxv-2-fast, blackforestlabs/flux-3-video-t2v, blackforestlabs/flux-3-video-i2v, blackforestlabs/flux-3-video-v2v, blackforestlabs/flux-3-video-draft-t2v, blackforestlabs/flux-3-video-draft-i2v, blackforestlabs/flux-3-video-draft-v2v, beeble/switchx-video-to-video, switchx-video-to-video responses: '200': content: application/json: schema: type: object properties: id: type: string description: The ID of the generated video. example: 60ac7c34-3224-4b14-8e7d-0aa0db708325 status: type: string enum: - queued - generating - completed - error description: The current status of the generation task. example: completed video: anyOf: - type: object properties: url: type: string format: uri description: The URL where the file can be downloaded from. example: https://cdn.aimlapi.com/generations/hedgehog/1759866285599-0cdfb138-c03a-49d4-a601-4f6413e27b15.mp4 required: - url - type: array items: type: object properties: url: type: string format: uri description: The URL where the file can be downloaded from. example: https://cdn.aimlapi.com/generations/hedgehog/1759866285599-0cdfb138-c03a-49d4-a601-4f6413e27b15.mp4 required: - url - {} error: type: - object - 'null' properties: name: type: string message: type: string required: - name - message description: Description of the error, if any. meta: type: - object - 'null' properties: usage: type: - object - 'null' properties: credits_used: type: number description: The number of tokens consumed during generation. example: 120000 usd_spent: type: number description: The total amount of money spent by the user in USD. example: 0.06 required: - credits_used - usd_spent description: Additional details about the generation. required: - id - status tags: - Video summary: V2 video generations x-summary-source: derived x-operation-id-source: normalized x-operation-id-original: _v2_video_generations