openapi: 3.2.0 info: title: VideoGen Tools API version: 1.0.0 description: Programmatically generate images, videos, voiceovers, sound effects, and avatar clips. servers: - url: https://api.videogen.io description: Production security: - bearerAuth: [] tags: - name: Tools description: Generate images, videos, audio, and more. All tool endpoints are asynchronous. paths: /v1/tools/generate-image: post: tags: - Tools operationId: generateImage x-fern-audiences: - rest summary: Generate image description: Generate an image from a text prompt, optionally guided by reference images and actor, product, or visual-style entity ids. When reference images are provided, the prompt describes the desired transformation. VideoGen automatically routes each request to the most effective state-of-the-art image model for your prompt, reference images, entities, and quality tier, so you don't pick a model. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/GenerateImageRequest' responses: '202': description: Execution accepted; poll until complete. content: application/json: schema: $ref: '#/components/schemas/StartToolExecutionResponse' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' /v1/tools/generate-video-clip: post: tags: - Tools operationId: generateVideoClip x-fern-audiences: - rest summary: Generate video clip description: Generate a single short video clip (up to 30 seconds) from a text prompt, optionally guided by an opening-frame still, reference images, videos, and audio. At least one of `prompt`, `startFrameFileId`, `imageFileIds`, `videoFileIds`, `audioFileIds`, or `spokenDialogue` must be provided. VideoGen automatically routes each request to the most effective state-of-the-art video model for your inputs and settings, so you don't pick a model. This endpoint returns one standalone clip. For longer, higher-quality, professionally edited videos with narration, captions, music, and multiple scenes, use a video workflow such as Script to video (`POST /v1/workflows/script-to-video`) instead. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/GenerateVideoClipRequest' responses: '202': description: Execution accepted; poll until complete. content: application/json: schema: $ref: '#/components/schemas/StartToolExecutionResponse' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' /v1/tools/generate-motion-graphic: post: tags: - Tools operationId: generateMotionGraphic x-fern-audiences: - rest summary: Generate motion graphic description: 'Generate an animated motion graphic video from a text prompt. This is an experimental, fully agentic alternative to a video workflow: VideoGen plans the animation, optionally generates or fetches supporting media, and renders a self-contained animated clip. It is especially well suited to precise text animations (e.g. a typing effect, animated captions, kinetic typography, lower thirds) that are hard to express with stock or generated footage. Optionally pass uploaded `fileIds` for reference media and `entityIds` for actors, products, or visual styles the animation should use. This endpoint returns one standalone video. For longer, narrated, multi-scene videos, use a video workflow such as Script to video (`POST /v1/workflows/script-to-video`) instead.' requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/GenerateMotionGraphicRequest' responses: '202': description: Execution accepted; poll until complete. content: application/json: schema: $ref: '#/components/schemas/StartToolExecutionResponse' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' /v1/tools/text-to-speech: post: tags: - Tools operationId: textToSpeech x-fern-audiences: - rest summary: Text to speech description: Convert text into a spoken audio file. Only voices with `supportsDirectToolExecution` set to true can be used. Optionally choose a voice, language, speed, and pronunciation overrides. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/TextToSpeechRequest' responses: '202': description: Execution accepted; poll until complete. content: application/json: schema: $ref: '#/components/schemas/StartToolExecutionResponse' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' /v1/tools/generate-sound-effect: post: tags: - Tools operationId: generateSoundEffect x-fern-audiences: - rest summary: Generate sound effect description: Generate a sound effect from a text description. Optionally control the duration and prompt influence. VideoGen automatically routes each request to the most effective state-of-the-art sound effect model for your prompt and settings, so you don't pick a model. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/GenerateSoundEffectRequest' responses: '202': description: Execution accepted; poll until complete. content: application/json: schema: $ref: '#/components/schemas/StartToolExecutionResponse' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' /v1/tools/generate-music: post: tags: - Tools operationId: generateMusic x-fern-audiences: - rest summary: Generate music description: Generate an instrumental music track from a text description. The returned track is approximately 30 seconds long. VideoGen automatically routes each request to the most effective state-of-the-art music model for your prompt, so you don't pick a model. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/GenerateMusicRequest' responses: '202': description: Execution accepted; poll until complete. content: application/json: schema: $ref: '#/components/schemas/StartToolExecutionResponse' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' /v1/tools/generate-avatar: post: tags: - Tools operationId: generateAvatar x-fern-audiences: - rest summary: Generate avatar clip description: Generate a talking-head avatar video by pairing an ACTOR entity with an audio file, typically from a prior text-to-speech result. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/GenerateAvatarRequest' responses: '202': description: Execution accepted; poll until complete. content: application/json: schema: $ref: '#/components/schemas/StartToolExecutionResponse' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' /v1/tools/vectorize-image: post: tags: - Tools operationId: vectorizeImage x-fern-audiences: - rest summary: Vectorize image description: Convert any raster image into a scalable vector graphic (SVG). The output traces the shapes and colors of the input image. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/ImageAssetRequest' responses: '202': description: Execution accepted; poll until complete. content: application/json: schema: $ref: '#/components/schemas/StartToolExecutionResponse' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' /v1/tools/remove-image-background: post: tags: - Tools operationId: removeImageBackground x-fern-audiences: - rest summary: Remove background from an image description: Remove the background from an image, returning a transparent-background PNG of the foreground subject. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/ImageAssetRequest' responses: '202': description: Execution accepted; poll until complete. content: application/json: schema: $ref: '#/components/schemas/StartToolExecutionResponse' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' /v1/tools/remove-video-background: post: tags: - Tools operationId: removeVideoBackground x-fern-audiences: - rest summary: Remove background from a video description: Remove the background from a video, producing a transparent-background video of the foreground subject. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/VideoAssetRequest' responses: '202': description: Execution accepted; poll until complete. content: application/json: schema: $ref: '#/components/schemas/StartToolExecutionResponse' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' /v1/tools/upscale-image: post: tags: - Tools operationId: upscaleImage x-fern-audiences: - rest summary: Upscale an image description: Increase the resolution of an image while preserving detail and sharpness. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/ImageAssetRequest' responses: '202': description: Execution accepted; poll until complete. content: application/json: schema: $ref: '#/components/schemas/StartToolExecutionResponse' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' /v1/tools/upscale-video: post: tags: - Tools operationId: upscaleVideo x-fern-audiences: - rest summary: Upscale a video description: Increase the resolution of a video while preserving detail and sharpness. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/VideoAssetRequest' responses: '202': description: Execution accepted; poll until complete. content: application/json: schema: $ref: '#/components/schemas/StartToolExecutionResponse' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' /v1/tools/image-3d-effect: post: tags: - Tools operationId: image3dEffect x-fern-audiences: - rest summary: Add 3D motion to an image description: Turn a still image into a short video clip with a 3D parallax motion effect, simulating camera movement through the scene. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/ImageAssetRequest' responses: '202': description: Execution accepted; poll until complete. content: application/json: schema: $ref: '#/components/schemas/StartToolExecutionResponse' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' /v1/tools/executions/{toolExecutionId}/cancel: post: tags: - Tools operationId: cancelToolExecution x-fern-audiences: - rest summary: Cancel tool execution description: Request cancellation of a running tool execution. The execution transitions to `cancelled` if it has not already completed. parameters: - $ref: '#/components/parameters/ToolExecutionIdPath' responses: '202': description: Cancellation request accepted content: application/json: schema: $ref: '#/components/schemas/StartToolExecutionResponse' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' /v1/tools/executions: get: tags: - Tools operationId: listToolExecutions x-fern-audiences: - rest summary: List tool executions description: List tool executions started via the API, most recently created first. Use `selfOnly=true` to restrict results to the calling API key's user; otherwise all executions for the team are returned. Cursor-paginated; see the Pagination guide. Executions remain listable indefinitely (including those older than 7 days). For efficiency this list does not re-sign result download URLs, so `downloadUrl`/`thumbnailUrl` reflect the last time they were signed and may be expired (always check `downloadUrlExpiresAt`). To obtain a fresh signed URL, GET the individual execution (`GET /v1/tools/executions/{toolExecutionId}`) or hydrate the file (`GET /v1/files/{fileId}` / `POST /v1/files/{fileId}/hydrate`). parameters: - $ref: '#/components/parameters/PaginationLimit' - $ref: '#/components/parameters/PaginationCursor' - $ref: '#/components/parameters/SelfOnlyQuery' responses: '200': description: Paginated list of tool executions. content: application/json: schema: $ref: '#/components/schemas/ToolExecutionListResponse' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' /v1/tools/executions/{toolExecutionId}: get: tags: - Tools operationId: getToolExecutionInfo x-fern-audiences: - rest summary: Get tool execution info description: 'Retrieve the current status and result of a tool execution. Poll this endpoint until `status` is `succeeded`, `failed`, or `cancelled`. Each succeeded result carries signed `downloadUrl`/`thumbnailUrl` (with matching `*ExpiresAt` timestamps) alongside the full hydrated `file`. The signed URLs are private and valid for 7 days; this single-execution endpoint automatically re-signs them when they are within an hour of expiring, so a caller always receives a URL valid long enough to use. Results remain retrievable indefinitely, including executions older than 7 days. Note that list endpoints (e.g. `GET /v1/tools/executions`) do not re-sign URLs: they return the last-signed URL, which may already be expired (check `downloadUrlExpiresAt`). To refresh a URL, GET the individual execution here, or hydrate the file directly via `GET /v1/files/{fileId}` / `POST /v1/files/{fileId}/hydrate`.' parameters: - $ref: '#/components/parameters/ToolExecutionIdPath' responses: '200': description: Current execution state content: application/json: schema: $ref: '#/components/schemas/ExecutedTool' default: description: Error content: application/json: schema: $ref: '#/components/schemas/ApiError' components: schemas: GenerateVideoClipRequest: type: object description: At least one of `prompt`, `startFrameFileId`, `imageFileIds`, `videoFileIds`, `audioFileIds`, or `spokenDialogue` must be provided. properties: prompt: type: string example: A golden retriever running through a sunlit meadow in slow motion, cinematic description: Text prompt describing the video to generate. Optional when reference media, `startFrameFileId`, or `spokenDialogue` is provided. Describe the video in plain language; any reference media you provide is incorporated automatically. startFrameFileId: type: string description: Optional file id of the opening-frame still (e.g. `vg_file_...`). Upload first via `POST /v1/files/upload`. When set, this image is the first frame of the clip. If the same id also appears in `imageFileIds`, it is used only as the opening frame and dropped from the reference list. Can be the only input (prompt optional). imageFileIds: type: array items: type: string maxItems: 4 description: Optional file ids of reference images (e.g. `["vg_file_..."]`). Upload files first via `POST /v1/files/upload`, then pass the returned ids here. When provided, the images are used as visual guidance. To animate a specific still as the opening frame, pass it as `startFrameFileId` instead. videoFileIds: type: array items: type: string maxItems: 4 description: Optional file ids of reference videos (e.g. `["vg_file_..."]`). Upload files first via `POST /v1/files/upload`, then pass the returned ids here. They are used as motion or style guidance for the generated video. audioFileIds: type: array items: type: string maxItems: 4 description: Optional file ids of reference audio clips (e.g. `["vg_file_..."]`) used for lip-sync from that recording. Upload files first via `POST /v1/files/upload`, then pass the returned ids here. To have the model speak a line it generates itself, pass `spokenDialogue` instead (or in addition). spokenDialogue: type: string description: Optional exact line the subject should speak as native, lip-synced speech in the generated clip. The model synthesizes the voice from this text. Can be the only input. Combine with a visual `prompt`, `startFrameFileId`, or reference media. Combine with `audioFileIds` when you also have a reference recording. voiceDescription: type: string description: Optional natural-language description of the voice that speaks `spokenDialogue` (for example, a warm, confident young man's voice). Used when `spokenDialogue` is set. When omitted, a clear natural voice is used. generateAudio: type: boolean default: false description: When true, the generated video is guaranteed to include audio. When false, audio may still be present. Defaults to false. suppressBackgroundMusic: type: boolean default: false description: When true, the generated clip will not include a musical soundtrack. Spoken dialogue and environmental sound are still allowed. Use this when you will add background music separately (for example at the project level). Defaults to false. durationSeconds: type: - integer - 'null' minimum: 1 maximum: 30 description: Optional clip length in whole seconds (1 to 30). Omit or pass null for Auto (duration is estimated at generate time so spoken text or the visual beat fits, clamped to the selected quality's supported range). When set, that length is used as-is. This endpoint produces a single short clip. For longer, multi-scene, professionally edited videos, use a video workflow such as `POST /v1/workflows/script-to-video`. aspectRatio: $ref: '#/components/schemas/AspectRatio' description: Aspect ratio for the generated video. Defaults to 16:9 when omitted. quality: $ref: '#/components/schemas/ModelQuality' description: Video generation quality tier (`LOW`, `STANDARD`, `HIGH`, or `MAX`). Optional; when omitted, your account's Default AI quality for video is used (change it at https://app.videogen.io/settings/account). contentPolicyConfig: $ref: '#/components/schemas/ContentPolicyConfig' watermarkMode: $ref: '#/components/schemas/WatermarkMode' numResults: type: integer minimum: 1 maximum: 100 default: 1 description: Number of output results to generate. Defaults to 1. isOutputTemporary: type: boolean default: false description: When true, generated files are temporary. Temporary files are guaranteed to be available for 24 hours, after which they may be archived at any time. Temporary files are not analyzed (no description, transcript, or embedding will be generated), so they will not appear in search results. Defaults to false. hideFromUi: type: boolean default: false description: When true, generated files are hidden from the VideoGen Media page by default. They remain accessible through the API. Defaults to false. AspectRatio: type: object required: - width - height description: Aspect ratio as a width:height pair (e.g. 16 and 9 for 16:9). Not pixel dimensions. properties: width: type: integer minimum: 1 height: type: integer minimum: 1 GenerateAvatarRequest: type: object required: - actorEntityId - audioFileId description: Generate a talking-head avatar from an ACTOR entity and audio. properties: actorEntityId: type: string description: The id of a built-in stock actor or an ACTOR entity (e.g. `vg_enti_...`) with an image reference. avatarQuality: $ref: '#/components/schemas/ModelQuality' description: Avatar generation quality tier. Optional; when omitted, your account's Default AI quality for avatars is used. audioFileId: type: string description: File id of an AUDIO file (e.g. `vg_file_...`), typically from a prior text-to-speech result. Upload a file first via `POST /v1/files/upload` or generate one with `POST /v1/tools/text-to-speech`, then pass the returned id here. watermarkMode: $ref: '#/components/schemas/WatermarkMode' numResults: type: integer minimum: 1 maximum: 100 default: 1 description: Number of output results to generate. Defaults to 1. isOutputTemporary: type: boolean default: false description: When true, generated files are temporary. Temporary files are guaranteed to be available for 24 hours, after which they may be archived at any time. Temporary files are not analyzed (no description, transcript, or embedding will be generated), so they will not appear in search results. Defaults to false. hideFromUi: type: boolean default: false description: When true, generated files are hidden from the VideoGen Media page by default. They remain accessible through the API. Defaults to false. PronunciationReplacement: type: object required: - original - replacement properties: original: type: string replacement: type: string WatermarkMode: type: string enum: - NONE - VIDEO_GEN - AUTO default: AUTO description: Controls whether the VideoGen watermark is applied to the output. `AUTO` applies the watermark unless you have a Pro plan. `VIDEO_GEN` always applies it. `NONE` removes the watermark (requires Pro; returns an error if you don't have it). FileInfo: type: object description: Metadata for a generated file. Obtain ids from tool results or `GET /v1/files`. required: - fileId - scope properties: fileId: type: string description: File id (e.g. `vg_file_...`). type: description: File type. Null when the file is still being processed and the type has not yet been determined. anyOf: - $ref: '#/components/schemas/FileType' - type: 'null' scope: type: string enum: - GLOBAL - PROJECT - EXPORT - TEMPORARY - ENTITY description: 'File scope. - `GLOBAL`: user-uploaded or standalone generated files that persist indefinitely. - `PROJECT`: project-specific files (e.g. text-to-speech clips in a generated project). - `EXPORT`: project exports. - `TEMPORARY`: short-lived files guaranteed to be available for 24 hours, after which they may be archived at any time. Not analyzed (no description, transcript, or embedding). - `ENTITY`: files attached to a reusable entity (e.g. a voice sample for an actor), shared across your team. ' displayName: type: string description: Display name for the file. description: type: - string - 'null' durationSeconds: type: - number - 'null' description: Duration in seconds for video and audio files. Null for images. transcript: description: Timed transcript for video and audio files, when available, as a `Transcript` object with timed `words`. Null for images or when no transcript has been generated. For plain transcript text, use `transcriptText`. anyOf: - $ref: '#/components/schemas/Transcript' - type: 'null' transcriptText: type: - string - 'null' description: Plain transcript text for video and audio files, when available. Null for images or when no transcript has been generated. downloadUrl: type: - string - 'null' format: uri description: Private signed URL for the highest-quality downloadable rendition, provided at the top level for convenience. Valid for 7 days from when it was signed. `null` when the rendition is still processing or the URL has not been signed yet. See `downloadUrlExpiresAt` for the exact expiry and `downloadSource` for the full rendition metadata; call `POST /v1/files/{fileId}/hydrate` to refresh it. downloadUrlExpiresAt: type: - integer - 'null' description: Seconds since epoch (Unix timestamp) when `downloadUrl` expires. `null` when `downloadUrl` is null. thumbnailUrl: type: - string - 'null' format: uri description: Private signed URL for the thumbnail rendition, provided at the top level for convenience. Valid for 7 days from when it was signed. `null` for file types that have no thumbnail (e.g. audio) or when it has not been signed yet. See `thumbnailSource` for the full rendition metadata. thumbnailUrlExpiresAt: type: - integer - 'null' description: Seconds since epoch (Unix timestamp) when `thumbnailUrl` expires. `null` when `thumbnailUrl` is null. thumbnailSource: description: Thumbnail image source. Populated after hydration. anyOf: - $ref: '#/components/schemas/FileSource' - type: 'null' previewSource: description: Preview rendition source (720p for video, resized for images). Populated after hydration. anyOf: - $ref: '#/components/schemas/FileSource' - type: 'null' downloadSource: description: Highest-quality downloadable rendition. Populated after hydration. anyOf: - $ref: '#/components/schemas/FileSource' - type: 'null' hlsSource: description: Private HLS streaming source. Populated for video and audio files once streaming renditions are ready. Uses a signed token; treat like other signed sources. anyOf: - $ref: '#/components/schemas/FileSource' - type: 'null' isPublicPreviewEnabled: type: boolean description: Whether public preview is enabled for this file. When true, `staticPublicPreviewSource` is populated for all file types. For video and audio, `publicHlsUrl` and `publicPlaybackId` are also populated once embed streaming is ready. staticPublicPreviewSource: description: Permanent public URL for the file's highest-quality rendition. Populated when `isPublicPreviewEnabled` is true. Does not expire (`expiresAt` is null). Use for direct links to images, downloads, or any file type. For embedded video or audio players, prefer `publicPlaybackId`. anyOf: - $ref: '#/components/schemas/FileSource' - type: 'null' publicHlsUrl: type: - string - 'null' description: Public HLS streaming URL for video and audio. Only present when `isPublicPreviewEnabled` is true and embed streaming is ready. Prefer `publicPlaybackId` with `@videogen/player` for embeds. publicPlaybackId: type: - string - 'null' description: Encoded public playback id (e.g. `vg_play_...`) for video and audio embeds. Pass this to `@videogen/player` or `@videogen/player-react`. Only present when `isPublicPreviewEnabled` is true and embed streaming is ready. For a permanent direct file URL (any type), use `staticPublicPreviewSource` instead. sourceToolType: type: string description: Tool type that generated this file (e.g. `GENERATE_IMAGE`, `TEXT_TO_SPEECH`). Only present when the file was created by a tool execution. sourceToolExecutionId: type: string description: Execution id of the tool call that generated this file (e.g. `vg_tool_...`). Only present when the file was created by a tool execution. fileAnalysisMetadata: description: Background analysis state for the file (used to populate `description`, `transcript`, `durationSeconds`, and the search embedding). Omitted when the file was returned via a path that does not check analysis progress (e.g. tool-result inline files and webhook payloads). $ref: '#/components/schemas/FileAnalysisMetadata' MotionGraphicSubToolModes: type: object description: Per-capability access modes for a motion graphic. Any omitted capability defaults to AUTO. properties: generateImages: $ref: '#/components/schemas/MotionGraphicSubToolMode' description: Whether the motion graphic may generate images. Defaults to AUTO. generateVideoClips: $ref: '#/components/schemas/MotionGraphicSubToolMode' description: Whether the motion graphic may generate video clips. Requires a paid plan. Defaults to AUTO. generateVoiceover: $ref: '#/components/schemas/MotionGraphicSubToolMode' description: Whether the motion graphic may generate voiceover audio. Defaults to AUTO. searchStockMedia: $ref: '#/components/schemas/MotionGraphicSubToolMode' description: Whether the motion graphic may search stock media. Defaults to AUTO. StartToolExecutionResponse: type: object description: Returned when a tool execution is started. Use `toolExecutionId` to poll for results or cancel. required: - toolExecutionId properties: toolExecutionId: type: string description: Execution id (e.g. `vg_tool_...`). FileSource: type: object description: A rendition source for a file (e.g. thumbnail, preview, download). Contains a signed URL and metadata. required: - status properties: status: type: string enum: - pending - ready - failed - skipped description: '`pending`: asset is still processing or has not been hydrated yet. `ready`: signed URL is available. `failed`: rendition generation failed. `skipped`: rendition does not apply to this file type (e.g. thumbnail for audio).' url: type: - string - 'null' description: Signed URL. Present when status is `ready` and file has been recently hydrated. If missing, call the hydrate endpoint. expiresAt: type: - integer - 'null' description: Seconds since epoch (Unix timestamp) when the signed URL expires. width: type: - integer - 'null' description: Rendition width in pixels, when known. height: type: - integer - 'null' description: Rendition height in pixels, when known. fileBytes: type: - integer - 'null' description: File size in bytes, when known. GenerateSoundEffectRequest: type: object required: - prompt properties: prompt: type: string description: A text description of the sound effect to generate. durationSeconds: type: - number - 'null' minimum: 1 maximum: 30 description: Desired length of the sound effect in seconds, between 1 and 30. Defaults to about 10 seconds when omitted. promptInfluence: type: - number - 'null' minimum: 0 maximum: 1 description: How closely the generated sound effect follows the prompt, between 0 (more creative, more variation) and 1 (more literal, less variation). Defaults to a balanced value when omitted. numResults: type: integer minimum: 1 maximum: 100 default: 1 description: Number of output results to generate. Defaults to 1. isOutputTemporary: type: boolean default: false description: When true, generated files are temporary. Temporary files are guaranteed to be available for 24 hours, after which they may be archived at any time. Temporary files are not analyzed (no description, transcript, or embedding will be generated), so they will not appear in search results. Defaults to false. hideFromUi: type: boolean default: false description: When true, generated files are hidden from the VideoGen Media page by default. They remain accessible through the API. Defaults to false. ImageAssetRequest: type: object required: - imageFileId properties: imageFileId: type: string description: File id of the source image (e.g. `vg_file_...`). Upload a file first via `POST /v1/files/upload`, then pass the returned id here. watermarkMode: $ref: '#/components/schemas/WatermarkMode' numResults: type: integer minimum: 1 maximum: 100 default: 1 description: Number of output results to generate. Defaults to 1. isOutputTemporary: type: boolean default: false description: When true, generated files are temporary. Temporary files are guaranteed to be available for 24 hours, after which they may be archived at any time. Temporary files are not analyzed (no description, transcript, or embedding will be generated), so they will not appear in search results. Defaults to false. hideFromUi: type: boolean default: false description: When true, generated files are hidden from the VideoGen Media page by default. They remain accessible through the API. Defaults to false. ApiError: type: object description: 'Standard error body returned with every non-2xx response (the `default` response of every operation). The HTTP status code conveys the error class; this body carries the details: - `400` invalid request, `401` missing or invalid API key, `403` not permitted (e.g. plan or add-on required, see `requirement`), `404` not found, `409` conflict, `429` rate limited or out of credits, `5xx` server error. Common `code` values include `invalid_request`, `invalid_api_key`, `not_authorized`, `not_found`, `insufficient_credits`, and `rate_limited`. Always branch on `code` (and `requirement.type` when present) rather than parsing `message`. ' required: - message properties: message: type: string description: Human-readable error description. For display and logging only; do not branch on its exact text. code: type: - string - 'null' description: Machine-readable error code in snake_case (e.g. `invalid_api_key`, `insufficient_credits`). `null` when no specific code applies. requirement: description: What is needed to resolve the error. Present when the error can be fixed by fulfilling a specific requirement (e.g. purchasing an add-on); `null` otherwise. anyOf: - $ref: '#/components/schemas/ErrorRequirement' - type: 'null' internalErrorCode: type: - string - 'null' description: Opaque internal error code for debugging. Include this when contacting support. `null` when not applicable. ErrorRequirement: type: object description: What is needed to resolve an error, when it can be fixed by fulfilling a specific requirement (e.g. purchasing an add-on or upgrading the plan). required: - type properties: type: type: string description: Machine-readable requirement type in snake_case (e.g. `purchase_add_on`, `upgrade_plan`). details: type: object additionalProperties: type: string description: Key-value pairs with requirement-specific context (e.g. the add-on id to purchase). VideoAssetRequest: type: object required: - videoFileId properties: videoFileId: type: string description: File id of the source video (e.g. `vg_file_...`). Upload a file first via `POST /v1/files/upload`, then pass the returned id here. watermarkMode: $ref: '#/components/schemas/WatermarkMode' numResults: type: integer minimum: 1 maximum: 100 default: 1 description: Number of output results to generate. Defaults to 1. isOutputTemporary: type: boolean default: false description: When true, generated files are temporary. Temporary files are guaranteed to be available for 24 hours, after which they may be archived at any time. Temporary files are not analyzed (no description, transcript, or embedding will be generated), so they will not appear in search results. Defaults to false. hideFromUi: type: boolean default: false description: When true, generated files are hidden from the VideoGen Media page by default. They remain accessible through the API. Defaults to false. TranscriptWord: type: object required: - startSeconds - endSeconds - word description: A single timed word of a transcript. properties: startSeconds: type: number minimum: 0 description: Start time of the word in seconds from the beginning of the audio. endSeconds: type: number description: End time of the word in seconds from the beginning of the audio. Must be greater than `startSeconds`. word: type: string description: The spoken word, used verbatim for narration timing and captions. Transcript: type: object required: - words description: A transcript of an audio file, as timed words in order. properties: languageCode: type: - string - 'null' description: Optional BCP-47 language code of the spoken audio (e.g. `en`, `es`). Used to tag the transcript's language; omit if unknown. words: type: array items: $ref: '#/components/schemas/TranscriptWord' description: The transcript words, sorted by `startSeconds` and non-overlapping. Must contain at least one word. GenerateMotionGraphicRequest: type: object required: - prompt properties: prompt: type: string example: A dark terminal window that types out the command `npm run build` character by character, then shows a green success checkmark description: Text prompt describing the animated motion graphic to generate. Describe the on-screen elements, any text and how it should animate, and the overall motion in plain language. fileIds: type: array items: type: string maxItems: 8 description: Optional file ids of uploaded reference media (images, videos, or audio) the motion graphic may display or animate (e.g. `["vg_file_..."]`). Upload files first via `POST /v1/files/upload`, then pass the returned ids here. entityIds: type: array items: type: string description: Optional actor, product, or visual-style entity ids (e.g. `["vg_enti_..."]`). The motion graphic uses each entity as identity/reference the same way in-app motion graphic generation does. Can be combined with `fileIds`. Mentions in `prompt` are also collected. A missing id returns not found; an inaccessible id returns a permission error. durationSeconds: type: - integer - 'null' minimum: 1 maximum: 300 description: Desired length of the motion graphic in seconds, a whole number between 1 and 300. When omitted, the duration is chosen automatically to fit the prompt (recommended). aspectRatio: $ref: '#/components/schemas/AspectRatio' description: Aspect ratio for the generated motion graphic. Defaults to 16:9 when omitted. transparentBackground: type: boolean default: true description: When true, renders the motion graphic with a transparent background as a WebM video suitable for overlaying on other video or images. Set to false for an opaque MP4. Defaults to true. subToolModes: $ref: '#/components/schemas/MotionGraphicSubToolModes' description: Optional per-capability controls for the media the motion graphic may generate or fetch (generated images, generated video clips, generated voiceover, and stock media search). Each capability is AUTO, ENABLED, or DISABLED. Omit to use AUTO for every capability. numResults: type: integer minimum: 1 maximum: 100 default: 1 description: Number of output results to generate. Defaults to 1. isOutputTemporary: type: boolean default: false description: When true, generated files are temporary. Temporary files are guaranteed to be available for 24 hours, after which they may be archived at any time. Temporary files are not analyzed (no description, transcript, or embedding will be generated), so they will not appear in search results. Defaults to false. hideFromUi: type: boolean default: false description: When true, generated files are hidden from the VideoGen Media page by default. They remain accessible through the API. Defaults to false. GenerateImageRequest: type: object required: - prompt properties: prompt: type: string example: A serene Japanese garden with cherry blossoms at golden hour description: Text prompt describing the image to generate. When reference images are provided, the prompt describes the desired transformation. imageFileIds: type: array items: type: string maxItems: 4 description: Optional file ids of reference images (e.g. `["vg_file_..."]`). Upload files first via `POST /v1/files/upload`, then pass the returned ids here. Maximum 4 images. When provided, the model uses these as guidance for generation. entityIds: type: array items: type: string description: Optional actor, product, or visual-style entity ids (e.g. `["vg_enti_..."]`). The model uses each entity as identity/reference the same way in-app image generation does. Can be combined with `imageFileIds`. A missing id returns not found; an inaccessible id returns a permission error. aspectRatio: $ref: '#/components/schemas/AspectRatio' description: Aspect ratio for the generated image. Defaults to 16:9 when omitted. quality: $ref: '#/components/schemas/ModelQuality' description: Image generation quality tier. Optional; when omitted, your account's Default AI quality for images is used (change it at https://app.videogen.io/settings/account). contentPolicyConfig: $ref: '#/components/schemas/ContentPolicyConfig' watermarkMode: $ref: '#/components/schemas/WatermarkMode' numResults: type: integer minimum: 1 maximum: 100 default: 1 description: Number of output results to generate. Defaults to 1. isOutputTemporary: type: boolean default: false description: When true, generated files are temporary. Temporary files are guaranteed to be available for 24 hours, after which they may be archived at any time. Temporary files are not analyzed (no description, transcript, or embedding will be generated), so they will not appear in search results. Defaults to false. hideFromUi: type: boolean default: false description: When true, generated files are hidden from the VideoGen Media page by default. They remain accessible through the API. Defaults to false. MotionGraphicSubToolMode: type: string enum: - AUTO - ENABLED - DISABLED description: Access mode for a motion graphic capability. AUTO uses the capability only when it is available on your plan (generated video clips require a paid plan). ENABLED forces the capability on and returns an upgrade error if your plan lacks it. DISABLED prevents the motion graphic from using the capability. TextToSpeechRequest: type: object required: - ttsText - voiceId properties: ttsText: type: string voiceId: type: string description: Catalog `displayName` (e.g. `Matilda`) or voice id from `GET /v1/resources/tts-voices` (e.g. `vg_voic_...`). Only voices with `supportsDirectToolExecution` set to true are accepted. speechLanguageCode: type: - string - 'null' description: ISO-639-1 language hint for pronunciation (e.g. `en`, `es`, `zh`). pronunciationReplacements: type: array items: $ref: '#/components/schemas/PronunciationReplacement' autoExpandPronunciationReplacements: type: boolean description: When true, automatically expands numbers, symbols, acronyms, and other non-word tokens into their spoken forms before synthesis so the voice pronounces them correctly (e.g. `$100` → `one hundred dollars`, `NASA` → `nasa`, `3rd` → `third`). Defaults to false when omitted. voiceSpeed: type: number minimum: 0.5 maximum: 2 description: Speech rate multiplier, between 0.5 (half speed) and 2 (double speed). Defaults to the voice's default speed. numResults: type: integer minimum: 1 maximum: 100 default: 1 description: Number of output results to generate. Defaults to 1. isOutputTemporary: type: boolean default: false description: When true, generated files are temporary. Temporary files are guaranteed to be available for 24 hours, after which they may be archived at any time. Temporary files are not analyzed (no description, transcript, or embedding will be generated), so they will not appear in search results. Defaults to false. hideFromUi: type: boolean default: false description: When true, generated files are hidden from the VideoGen Media page by default. They remain accessible through the API. Defaults to false. ToolSuccessResult: type: object description: Result for a single generated file. Only appears inside a succeeded execution's `results`, so every field below is always present. required: - fileId - type - downloadUrl - downloadUrlExpiresAt - thumbnailUrl - thumbnailUrlExpiresAt - file properties: fileId: type: string description: File id for the generated asset. type: $ref: '#/components/schemas/FileType' description: File type. downloadUrl: type: - string - 'null' format: uri description: Private signed download URL for the generated file, valid for 7 days from when it was signed. Provided at the top level for convenience so you don't have to read it out of `file`. When you GET a single execution it is automatically re-signed if within an hour of expiring; list endpoints do not re-sign, so there it may be expired (check `downloadUrlExpiresAt`). See `downloadUrlExpiresAt` for the exact expiry. Null only in the rare case that the highest-quality rendition is still finalizing. downloadUrlExpiresAt: type: - integer - 'null' description: Seconds since epoch (Unix timestamp) when `downloadUrl` expires. Null only when `downloadUrl` is null. thumbnailUrl: type: - string - 'null' format: uri description: Private signed thumbnail URL for the generated file, valid for 7 days from when it was signed. Provided at the top level for convenience so you don't have to read it out of `file`. Re-signed on the same terms as `downloadUrl` (single-execution GET re-signs when near expiry; list endpoints do not). Null for file types that have no thumbnail (e.g. audio). thumbnailUrlExpiresAt: type: - integer - 'null' description: Seconds since epoch (Unix timestamp) when `thumbnailUrl` expires. Null when there is no thumbnail URL. file: description: Hydrated file metadata with signed download URLs (always present and hydrated for a succeeded result). Its signed URLs follow the same 24-hour validity and automatic re-signing as `downloadUrl`. $ref: '#/components/schemas/FileInfo' FileAnalysisMetadata: type: object description: 'Background analysis state for a file. Background analysis populates `description`, `transcript`, `durationSeconds`, and the search embedding after a file is uploaded or generated; this object lets you render a progress indicator while it runs (and skip rendering once it''s done). ' required: - analysisLoadingState - analysisProgressPercentage properties: analysisLoadingState: type: string enum: - UNATTEMPTED - LOADING - FULFILLED - REJECTED description: 'Coarse-grained analysis state. - `UNATTEMPTED`: analysis has not started yet. - `LOADING`: analysis is in progress. - `FULFILLED`: analysis completed successfully. `description`, `transcript`, and `durationSeconds` are now populated where applicable for the file''s type. - `REJECTED`: analysis failed permanently and will not be retried. ' analysisProgressPercentage: type: number description: Progress in `[0, 100]`. Always `100` when `analysisLoadingState` is `FULFILLED`. Otherwise the most recent in-flight progress reported by the analysis task (or `0` if no progress has been reported yet). analysisAttemptIndex: type: integer description: Zero-based index of the current analysis task attempt. Only present while analysis is still loading (`UNATTEMPTED` or `LOADING`); omitted once analysis reaches a terminal state. ExecutedTool: type: object required: - toolExecutionId - status - toolType - progressPercentage - attemptIndex - results - error properties: toolExecutionId: type: string description: Execution id matching the original request. status: $ref: '#/components/schemas/JobStatus' toolType: type: string description: Tool name (e.g. `GENERATE_IMAGE`, `TEXT_TO_SPEECH`). progressPercentage: type: number minimum: 0 maximum: 100 description: Completion progress for the current attempt (0-100). Always `100` when `status` is `succeeded`. attemptIndex: type: integer minimum: 0 description: Zero-based index of the current or most recent execution attempt. results: type: array description: One entry per generated result. Always present; empty until `status` is `succeeded`, then one entry per generated file (each with signed URLs and a hydrated `file`). items: $ref: '#/components/schemas/ToolSuccessResult' error: description: Error details. Always present; `null` unless `status` is `failed`. anyOf: - $ref: '#/components/schemas/ApiError' - type: 'null' ModelQuality: type: string enum: - LOW - STANDARD - HIGH - MAX description: 'AI generation quality tier, shared across every generative feature (image, video, text, and so on). `LOW` is fastest and cheapest, `STANDARD` balances quality and cost, `HIGH` is higher quality, and `MAX` is the highest quality. When a request omits the quality field, VideoGen falls back to your account''s **Default AI quality** for that feature, which you can change at [Account settings](https://app.videogen.io/settings/account). Not every feature supports every tier; unsupported tiers are rejected with an error (see each field''s description). ' JobStatus: type: string description: Lifecycle status shared by every asynchronous job (tool executions, workflow runs, remix actions, project exports, and timeline interchange jobs). `pending` and `running` are in-progress; `succeeded`, `failed`, and `cancelled` are terminal. enum: - pending - running - succeeded - failed - cancelled ContentPolicyConfig: type: object description: Controls how content-policy rejections are handled during generation. properties: maxPromptRewrites: type: integer minimum: 0 maximum: 4 default: 0 description: Maximum number of automatic prompt rewrites to attempt after a content-policy rejection before failing. Must be an integer between 0 and 4. 0 (the default) fails on the first rejection and returns the moderation error so you can revise the prompt yourself. Higher values let the request automatically rephrase and retry the prompt. ToolExecutionListResponse: type: object description: Paginated list of API-started tool executions, most recently created first. required: - toolExecutions - hasMore - nextCursor properties: toolExecutions: type: array items: $ref: '#/components/schemas/ExecutedTool' hasMore: type: boolean description: When true, there are more executions available. Pass `nextCursor` as the `cursor` query param to fetch the next page. nextCursor: type: - string - 'null' description: Opaque cursor to fetch the next page. `null` when `hasMore` is false. GenerateMusicRequest: type: object required: - prompt properties: prompt: type: string description: A text description of the music to generate. Include genre, mood, instrumentation, and tempo for best results. numResults: type: integer minimum: 1 maximum: 100 default: 1 description: Number of output results to generate. Defaults to 1. isOutputTemporary: type: boolean default: false description: When true, generated files are temporary. Temporary files are guaranteed to be available for 24 hours, after which they may be archived at any time. Temporary files are not analyzed (no description, transcript, or embedding will be generated), so they will not appear in search results. Defaults to false. hideFromUi: type: boolean default: false description: When true, generated files are hidden from the VideoGen Media page by default. They remain accessible through the API. Defaults to false. FileType: type: string enum: - IMAGE - VIDEO - AUDIO - PDF - SLIDESHOW - TEXT - LOTTIE description: File type. `TEXT` covers plain-text and editor-interchange documents; `LOTTIE` is a JSON animation. parameters: ToolExecutionIdPath: name: toolExecutionId in: path required: true schema: type: string description: The tool execution id returned when the tool was started. PaginationCursor: name: cursor in: query required: false schema: type: string description: Opaque pagination cursor returned as `nextCursor` by the previous page. Omit on the first request. Cursors are tied to the endpoint that produced them and must be passed unmodified. See [Pagination](/pagination). PaginationLimit: name: limit in: query required: false schema: type: integer minimum: 1 maximum: 200 default: 50 description: Maximum number of items to return in the page. Defaults to 50; capped at 200. See [Pagination](/pagination). SelfOnlyQuery: name: selfOnly in: query required: false schema: type: boolean default: false description: When true, returns only items created by the API key's owner. When false (default), returns all items accessible to the team. securitySchemes: bearerAuth: type: http scheme: bearer bearerFormat: opaque description: API key from [app.videogen.io/api](https://app.videogen.io/api). The full key is only shown once when you create it.