--- name: edit-video-and-transfer-motion description: "Change a generated clip with a prompt (swap the outfit from reference images, restyle the scene, change lighting or props) or make the person in an image perform the moves of a motion clip such as a dance, with Spicy API's video edit models, priced per second with the cost approved before each run. Use when a user asks to edit an AI video, change what someone wears in a clip, restyle a video, make a character dance, or do motion transfer from an API or an agent." tags: [video-editing, motion-transfer, ai-dance, outfit-swap, nsfw-video-generation, uncensored, mcp] license: MIT metadata: version: "1.4.0" homepage: https://www.spicyapi.com/skills/edit-video-and-transfer-motion/SKILL.md mcp: https://www.spicyapi.com/api/mcp openapi: https://www.spicyapi.com/openapi.json vendor: Spicy API (spicyapi.com) --- # Edit video and transfer motion ## When to use The user already has a clip (or a still of a person) and wants it changed: a different outfit, a new look or setting, or the person performing a dance or gesture. Every input must be the account's own: clips from its finished video tasks, images from `list_images`. ## Instruction edits 1. Pick the model (`list_models` with kind `video-edit`): `spicy-video-edit-1` for 2 to 10 second clips (output keeps the clip's length, `aspect_ratio` can reshape it, up to 4 reference images, `enhance_prompt`), or `spicy-cinema-1-edit` for 3 to 30 second clips (the first 15 seconds are edited and returned, up to 5 reference images). 2. `estimate_cost` with the model, `resolution` (720P or 1080P) and `input_seconds` (the clip's length). Both bill input plus output seconds: a 5-second 720P clip on `spicy-video-edit-1` is $2.00. 3. `edit_video` (guarded) with `video_url` (the clip's `output.video_url`), a `prompt` naming the change ("replace her dress with the red lingerie from Image 1, keep everything else"), optional `reference_image_urls` and `keep_audio: true` to keep the soundtrack. It is debited at the clip's length and settled to what is rendered. 4. `get_job` every 10 to 15 seconds until `succeeded` (`output.video_url`) or `failed` (refunded). ## Motion transfer 1. `list_motions`: the library (`mo_slow_sway` slow sway, `mo_hip_dance` hip dance, `mo_hair_turn` hair flip and turn, `mo_catwalk` catwalk, `mo_wave` wave hello), all 9:16, full body, fixed camera, generated with SpicyAPI video models. A clip the account generated (2 to 30 seconds) works too, as `video_url`. 2. The performer: one of the account's images of a dressed person, full body, framed like the motion clip (9:16). Motion transfer refuses explicit or nude images and clips (refunded, `error_code` `blocked`), so generate a clothed full-body still first if needed. 3. `estimate_cost` with `model: "spicy-animate-1"`, `motion` and `quality` (`standard` at $0.24 or `pro` at $0.36 per output second); the output is as long as the motion clip. 4. `edit_video` (guarded) with `model: "spicy-animate-1"`, `image_url` and `motion` (or `video_url`), then `get_job` as above. ## What to tell the user up front - Edits and motion transfer run on finished outputs of this account only; there are no uploads. - Instruction edits bill the input clip too, so trim long clips by generating shorter ones. ## Video edit models | Model | Kind | Price | Limits | Inputs (audio mode, valid combinations) | |---|---|---|---|---| | `spicy-video-edit-1` | video-edit | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4; prompt up to 5,000 characters; the input clip's seconds are billed too, at the same rate; edits your own 2 to 10 second clips (output up to 10 seconds) with up to 4 reference images, aspect_ratio, keep_audio | | | `spicy-cinema-1-edit` | video-edit | $0.28 per second at 720P, $0.48 per second at 1080P | resolutions 720P, 1080P; prompt up to 5,000 characters; the input clip's seconds are billed too, at the same rate; edits your own 3 to 30 second clips (output up to 15 seconds) with up to 5 reference images, keep_audio | | | `spicy-animate-1` | video-edit | $0.24 per second at standard, $0.36 per second at pro | resolutions standard, pro; motion transfer: image_url plus a motion id (GET /v1/videos/motions) or your own 2 to 30 second clip, quality standard or pro, billed per output second; needs image_url (one of your generated images) | | ## Tools used by this skill - `list_models`: Every model with its kind (image, image-edit, video, video-edit, chat, speech, realtime, transcription, embedding), USD price (per image, per second by resolution or quality, per 1M tokens, per 10,000 characters of speech input, or per minute of audio), limits (image sizes, outputs per call, resolutions, clip lengths, aspect ratios, prompt character caps), whether it needs or accepts an input image or references, and example clips. Video rows carry `inputs`: which of first frame, last frame, reference images, clip and audio the model takes, `inputs.audio.mode` (driving: the clip is the soundtrack and the mouth follows it; reference: the model generates sound and uses the clip for voice, tone and beat; none) and `inputs.combinations`, the valid input combinations in plain sentences. Read those before generate_video. `silent: true` marks a video model that renders without sound; `bills_input_video_seconds: true` one that also bills the seconds of an input clip. Video-edit rows carry `video_edit` (instruction or motion mode, clip lengths, reference images, whether input seconds are billed) for edit_video. Speech rows carry `speech`: preset voices, whether `instructions` are accepted, custom voices, languages, inline `tags` and voice creation fees. Realtime rows (spicy-live-1) are live voice calls over a WebSocket: a server or app opens a session with POST /v1/realtime/sessions and connects the returned ticket URL; they cannot be driven from a tool call. Prices here are exactly what the API bills. Works with a $0 balance. Arguments: `kind`?: image | image-edit | video | video-edit | chat | speech | realtime | transcription | embedding (Filter by kind.). - `estimate_cost`: USD cost of a request before making it, from the same price table the gateway bills: images by count, video by seconds and resolution (plus the input clip or action clip on models that bill it), video edits by clip length (input plus output seconds for instruction edits, output seconds for motion transfer), chat by tokens, speech by characters of input (plus instructions), transcription by seconds of audio, embeddings by tokens, live calls by minutes. Use it to quote the user and to check against get_account before a spend. Works with a $0 balance. Arguments: `model`: string (A model id from list_models, e.g. spicy-image-1, spicy-image-1-pro, spicy-image-photo-1.); `n`?: integer (Images per call (image models), clamped to the model's maximum.); `resolution`?: string (Video or video-edit resolution, e.g. 720P (default) or 1080P.); `quality`?: standard | pro (Motion transfer tier on spicy-animate-1, default standard.); `duration`?: integer (Video seconds, default 5; -1 for a model-chosen length where supported.); `input_seconds`?: number (Length of the input clip in seconds: the clip to edit or motion clip on video-edit models (default 5), or a video_url on video models that bill input seconds.); `motion`?: string (A library motion id (list_motions) on spicy-animate-1; its length sets the output length.); `action`?: string (An action id (list_actions) on spicy-motion-3 and -fast: its reference clip seconds (4 on most) are billed as input, and duration defaults to 8.); `fps`?: any (Video frame rate: 30 (default) or 60. 60 smooths motion with frame interpolation for 20% on top of the video price (input or action clip seconds included); audio is kept. If that step fails the 30 fps clip is delivered and the fee refunded.); `tokens`?: integer (Chat: total tokens, default 1000. Embeddings: text tokens, default 1000.); `image_tokens`?: integer (Embeddings on spicy-embed-vision-1: image tokens, default 0.); `characters`?: integer (Speech models: characters of input plus instructions (CJK ideographs and inline tags count too), default 100.); `audio_seconds`?: number (Transcription: seconds of audio, default 60.); `minutes`?: number (Live calls (spicy-live-1): minutes of conversation, default 1.); `count`?: integer (How many such requests, default 1.). - `get_account`: Balance in USD, spend this calendar month, the monthly spend limit if set, whether approvals are skipped for agents, whether the account has ever topped up, the webhook URL, and whether this key is a sandbox key. Call it first and before any spend. Arguments: no arguments. - `list_images`: The account's generated images, newest first. These are the only URLs accepted as image inputs by edit_image, generate_video, edit_video and create_embeddings (no uploads, no third-party URLs). Use it to pick a first frame to animate or a performer for motion transfer. Arguments: `limit`?: integer (Default 20.). - `generate_image` (guarded): Text to image, synchronous (10 to 30 seconds), returns durable CDN URLs reusable as image_url for edits and video. Or an image action: model spicy-image-action-1 with an action from list_image_actions and image_url of the woman (one of list_images) puts her face, hair and skin tone into that ready-made explicit scene, no prompt (a minute or two); style studio gives the finished picture a warm glamour colour grade, and anime, 3d or cartoon put her into the drawn version of the scene, for a drawn woman. Or a reference: model spicy-image-reference-1 (businesses approved for character import only) with reference_image (a picture the user supplies, https or data URL) recreates it as new people, every face replaced, body and outfit changed a little, room redesigned, and returns its description to reuse as a prompt (30 to 60 seconds, one image). One to six images per call, billed per image at the model's rate; the exact amount is cost_usd in the result. Prompts are screened first and outputs after (withheld images are refunded and counted in `withheld`); blocked prompts cost nothing. Every request depicts adults only. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token. Arguments: `model`: string (An image model id: spicy-image-1 (fast), spicy-image-1-pro (best prompt adherence), spicy-image-photo-1 (candid natural-light photo look), spicy-image-action-1 (image actions only) or spicy-image-reference-1 (reference_image only; businesses approved for character import).); `prompt`?: string (What to generate, up to 2000 characters. Required, except with action or reference_image where it is not accepted.); `reference_image`?: string (On spicy-image-reference-1 only, and required there: the picture to recreate as new people, a public https URL or a data URL (JPEG, PNG or WebP). No prompt, size, style, character or enhance_prompt; n is 1. Pictures of anyone who may be under 18 or of known people are refused at no charge.); `action`?: string (An image action id from list_image_actions, on spicy-image-action-1 only, e.g. pov-missionary: the still sets the act, pose, room and shape; image_url (or character) supplies the woman's face, hair and skin tone. No prompt, negative_prompt, size or enhance_prompt. Her body stays the still's. style studio adds a warm glamour colour grade to the result; anime, 3d or cartoon picks the drawn version of the still (see styles in list_image_actions); match it to her image.); `image_url`?: string (With action: the woman, one of this account's own generated images (list_images). Photo women take no style (or studio); anime, 3D or cartoon women take the matching style.); `negative_prompt`?: string; `size`?: string (width*height from the model's limits.sizes, one price per model whatever the size: 1024*1024, 832*1216, 1216*832, 1280*1280, 1440*1440, 1024*1536, 1536*1024, 1080*1920, 1920*1080, 1152*2048, 2048*1152. Default 1024*1024.); `n`?: integer (Images to return, 1 to 6, default 1. Billed per image delivered.); `seed`?: integer; `enhance_prompt`?: boolean (Let the model expand the prompt with lighting, camera and detail cues. Default false: the prompt goes as written.); `enhance_mode`?: direct | agent (With enhance_prompt: direct (default) rewrites once, agent plans the shot then writes.); `character`?: string (A saved character id (list_characters): the same person is rendered in the new scene, billed at the image-edit rate.); `style`?: photorealistic | studio | anime | 3d | cartoon (Art style (list_styles): photorealistic, studio (glossy glamour), anime, 3d (animated-film CGI) or cartoon (2D adult cartoon). Prompt text only, same price.); `approval_token`?: string (Approval token from a previous approval_required response, after the user said yes.). - `edit_video` (guarded): Change a finished clip, or make a person perform a motion. Asynchronous like generate_video: returns a task with the price debited; poll get_job every 10 to 15 seconds (refunded if it fails). Instruction edits (spicy-video-edit-1: 2 to 10 second clips, output the clip's length; spicy-cinema-1-edit: 3 to 30 second clips, the first 15 seconds are edited and returned): video_url (one of the account's finished video outputs) plus a prompt describing the change, optionally reference_image_urls for an outfit, prop or style (the account's own images), keep_audio to keep the soundtrack. Billed per second of input plus output at the resolution's rate, debited at the clip's length and settled to what is rendered. Motion transfer (spicy-animate-1): image_url (the performer, one of the account's images, full body and dressed works best) plus a motion id from list_motions or a video_url of the account's own clip (2 to 30 s); the output is as long as the motion clip, billed per output second by quality. Motion transfer refuses explicit or nude images and clips (refunded, error_code blocked). Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token. Arguments: `model`: string (A video-edit model id: spicy-video-edit-1, spicy-cinema-1-edit or spicy-animate-1.); `video_url`?: string (The clip: output.video_url of one of the account's finished video tasks. Required for instruction edits; on spicy-animate-1 an alternative to motion.); `prompt`?: string (Instruction edits (required): the change to make, up to 5000 characters. Not needed for motion transfer.); `reference_image_urls`?: array (Instruction edits: outfit, prop or style images from list_images, up to 4 on spicy-video-edit-1 and 5 on spicy-cinema-1-edit.); `resolution`?: 720P | 1080P (Instruction edits: 720P (default) or 1080P.); `aspect_ratio`?: string (spicy-video-edit-1 only: reshape to 16:9, 9:16, 1:1, 4:3 or 3:4.); `keep_audio`?: boolean (Instruction edits: true keeps the clip's original soundtrack.); `negative_prompt`?: string; `seed`?: integer; `enhance_prompt`?: boolean (spicy-video-edit-1 only: let the model expand the prompt.); `image_url`?: string (spicy-animate-1 (required): the person who performs the motion, one of the account's images.); `motion`?: string (spicy-animate-1: a library motion id from list_motions, e.g. mo_hip_dance. Send this or video_url.); `quality`?: standard | pro (spicy-animate-1: standard (default) or pro, priced separately.); `user`?: string (Your own id for the end user making this request.); `approval_token`?: string (Approval token from a previous approval_required response, after the user said yes.). - `list_motions`: The motion library for motion transfer on edit_video (spicy-animate-1): id, label, description, preview url, duration_s and aspect_ratio of each clip. Every clip is 9:16, full body, fixed camera and was generated with SpicyAPI video models (no real performer footage). The output is as long as the clip; match the performer image to its framing for best results. Works with a $0 balance. Arguments: no arguments. - `get_job`: Status of a video task from generate_video or edit_video (queued, processing, finalizing, succeeded with output.video_url and usage, or failed and refunded) or of a logged image, chat, speech, transcription or embedding request by id. Arguments: `id`: string (Task or request id, e.g. sj_...). - `list_jobs`: Recent requests on the account (images, edits, video tasks including edits and motion transfer, chat, speech, voice creation, transcription, embeddings, live calls), newest first: id, type, model, status, cost_usd, output URL when finished, error when failed, and whether it came from the playground, the API or an agent. Arguments: `limit`?: integer (Default 20.); `type`?: image | image-edit | video | chat | speech | voice | transcription | embedding | realtime. - `get_topup_link`: Checkout URL to add funds by card (hosted checkout) or crypto for a preset amount ($50, $100, $250, $500, $1000; minimum $50). The USER opens it and pays in the browser; never enter payment details yourself. The balance is credited by the payment webhook; call get_account afterwards. If it returns commitment_required, the user must first sign the one-page Customer Commitment Letter on the Billing page; an agent cannot sign it. Works with a $0 balance. Arguments: `amount_usd`?: integer (One of 50, 100, 250, 500, 1000. Default 50.); `method`?: card | crypto (Default card.). - `read_acceptable_use`: The acceptable use policy as markdown: prohibited content (minors in any form, images or video of real people with or without consent, non-consensual scenarios and the rest), age-verification duties for the customer's product, enforcement. Read it before the first generation and tell the user what their product must do. Works with a $0 balance. Arguments: no arguments. ## Setup (once per user) 1. Key: the user signs in once at https://www.spicyapi.com/auth (Google or email). The account is live at once with a $0 balance and a Default key shown once at https://www.spicyapi.com/dashboard/api-keys. Keep the key in an environment variable (`SPICYAPI_KEY`), never in chat, never in client-side code. With one key you can mint more with `create_api_key`. 2. Connect. MCP (Streamable HTTP): `https://www.spicyapi.com/api/mcp` with header `Authorization: Bearer `. Claude Code: `claude mcp add --transport http spicyapi https://www.spicyapi.com/api/mcp --header "Authorization: Bearer sk-spicy-..."`. Cursor, Codex and any URL-plus-headers client: same URL and header. REST instead of MCP: `POST https://www.spicyapi.com/api/v1/tools/` with the same header and a JSON body of arguments; `GET https://www.spicyapi.com/api/v1/tools` lists them; OpenAPI 3.1 at https://www.spicyapi.com/openapi.json. OAuth 2.1 clients (Claude custom connectors, ChatGPT, directory scanners) need only the URL: the endpoint advertises its authorization server, the user signs in and consents in the browser, and the token it returns is an API key they can revoke in the dashboard. 3. Dry run first: a key created with Sandbox ticked (or `create_api_key` with `sandbox: true`) answers every endpoint and every tool from fixture output flagged `sandbox: true`, needs no balance and bills nothing. Build against it, then swap the key. Never present sandbox output to the user as a real generation. 4. Money: `get_account` shows the balance; `estimate_cost` prices a request from the same table the API bills; `get_topup_link` returns a card or crypto checkout URL (presets $50, $100, $250, $500, $1000, minimum $50) that the USER opens and pays in the browser. Never enter card details yourself. A 402 means the balance ran out or the monthly limit was hit. 5. Rate limit: 60 requests per minute per key on the agent surfaces. ## Rules that always apply - Guarded tools (`generate_image`, `edit_image`, `generate_video`, `edit_video`, `transcribe_audio`, `create_embeddings`, `upload_audio`, `delete_audio`, `generate_speech`, `create_voice`, `delete_voice`, `create_character`, `delete_character`, `create_api_key`, `revoke_api_key`, and `set_spend_limit` when raising or removing a limit) first return `{error: "approval_required", summary, approval_token}`. Show the summary to the user verbatim (it carries the cost), get an explicit yes in the conversation, then call the tool again with the same arguments plus `approval_token` (15-minute expiry, bound to those exact arguments). Never approve in bulk or reuse a token for a different action. The account can turn approvals off in the dashboard; moderation, provenance and the spend limit still apply. - Input images for `edit_image`, `generate_video`, `edit_video` and `create_embeddings` must be images this account generated (`list_images`); input clips (`video_url`) must be the account's own finished tasks. Uploads and third-party URLs are rejected before any charge; do not try to work around it. The external inputs are audio: `upload_audio` takes a WAV or MP3 (2 to 30 seconds) from a public URL, transcribes and screens it for $0.01, and its returned `url`, or the `url` of a `generate_speech` clip, is what `audio_url` accepts; `transcribe_audio` reads any public audio URL into text. - The same person across many outputs: `create_character` from 1 to 3 generated images of them, then pass the returned id as `character` to `generate_image`, `edit_image` or `generate_video` (`spicy-character-video-1` needs no first frame). This is the supported way to get consistency; never ask the user for a photo. - Read `read_acceptable_use` before the first generation: adults only, no minors in any form (including "young-looking" or youth-coded content), no images or video of real people (with or without consent; voice clones need the speaker's written consent), no non-consensual scenarios. Blocked prompts return 422 and cost nothing; repeated attempts terminate the account. Tell the user their product must age-verify its end users and disclose that content is AI-generated. - When the user's product serves other people, pass each person's own stable id as `user` on every call. Screening history, strikes and suspensions then apply to that one person (403 `end_user_suspended`) instead of the whole account, and the id is kept with each generation's record, so a report about one person's output is traced to them and not to everyone on the account. - Everything is scoped to the account behind the key; there is no way to read another account's data. - Spicy API is spicyapi.com. spicyapi.ai is an unrelated aggregator; do not mix their prices or models. ## Prices (USD, for when the user asks) | Model | Kind | Price | Limits | Inputs (audio mode, valid combinations) | |---|---|---|---|---| | `spicy-image-1-pro` | image | $0.09 per image | sizes 1024*1024, 832*1216, 1216*832, 1280*1280, 1440*1440, 1024*1536, 1536*1024, 1080*1920, 1920*1080, 1152*2048, 2048*1152; up to 6 outputs per call | | | `spicy-image-photo-1` | image | $0.08 per image | sizes 1024*1024, 832*1216, 1216*832, 1280*1280, 1440*1440, 1024*1536, 1536*1024, 1080*1920, 1920*1080, 1152*2048, 2048*1152; up to 6 outputs per call | | | `spicy-image-1` | image | $0.06 per image | sizes 1024*1024, 832*1216, 1216*832, 1280*1280, 1440*1440, 1024*1536, 1536*1024, 1080*1920, 1920*1080, 1152*2048, 2048*1152; up to 6 outputs per call | | | `spicy-image-action-1` | image | $0.15 per image | up to 4 outputs per call | | | `spicy-image-reference-1` | image | $0.24 per image | up to 1 output per call | | | `spicy-image-edit-1` | image-edit | $0.15 per image | sizes 1024*1024, 832*1216, 1216*832, 1280*1280, 1440*1440, 1024*1536, 1536*1024, 1080*1920, 1920*1080, 1152*2048, 2048*1152; up to 6 outputs per call; needs image_url (one of your generated images) | | | `spicy-motion-3` | video | $0.1 per second at 480P, $0.2 per second at 720P, $0.4 per second at 1080P | resolutions 480P, 720P, 1080P; 2 to 30 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 21:9; prompt up to 20,000 characters; reference clip (video_url) up to 15 seconds; the input clip's seconds are billed too, at the same rate; image_url optional (text-to-video without it) | Audio: reference audio (the model generates sound and uses the clip for voice, tone and beat), up to 5 clips and 15s total. Valid combinations: "Text prompt only, with aspect_ratio; dialogue written in the prompt is spoken natively"; "First frame (image_url), optionally a last frame (last_frame_url); no references or audio in the same request"; "Reference images (up to 10, or a character) plus a text prompt: same person, framing chosen by the model"; "Reference images, a reference video (video_url, 15s max) and reference audio (up to 5 clips, 15s total) in any mix, with a text prompt; audio guides voice, tone and beat while the words come from the prompt" | | `spicy-motion-3-fast` | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.56 per second at 1080P | resolutions 480P, 720P, 1080P; 2 to 30 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 21:9; prompt up to 20,000 characters; reference clip (video_url) up to 15 seconds; the input clip's seconds are billed too, at the same rate; image_url optional (text-to-video without it) | Audio: reference audio (the model generates sound and uses the clip for voice, tone and beat), up to 5 clips and 15s total. Valid combinations: "Text prompt only, with aspect_ratio; dialogue written in the prompt is spoken natively"; "First frame (image_url), optionally a last frame (last_frame_url); no references or audio in the same request"; "Reference images (up to 10, or a character) plus a text prompt: same person, framing chosen by the model"; "Reference images, a reference video (video_url, 15s max) and reference audio (up to 5 clips, 15s total) in any mix, with a text prompt; audio guides voice, tone and beat while the words come from the prompt" | | `spicy-cinema-1-image` | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.36 per second at 1080P | resolutions 480P, 720P, 1080P; 3 to 15 seconds; prompt up to 5,000 characters; needs image_url (one of your generated images) | Audio: no audio input. Valid combinations: "First frame (image_url) with a text prompt; the output keeps the image's shape" | | `spicy-motion-2` | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; prompt up to 5,000 characters; needs image_url (one of your generated images) | Audio: driving audio (the clip is the soundtrack and the mouth follows it). Valid combinations: "First frame (image_url), optionally a last frame (last_frame_url)"; "First frame plus driving audio (audio_url): the clip becomes the soundtrack and the mouth follows it"; "First frame, last frame and driving audio together"; "A clip to continue (video_url, 2 to 10s) instead of a first frame, optionally with a last frame" | | `spicy-character-video-1` | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; prompt up to 5,000 characters; image_url optional (text-to-video without it) | Audio: no audio input. Valid combinations: "A character (or reference images) plus a text prompt"; "A character plus a first frame (image_url) that opens the clip" | | `spicy-cinema-1-character` | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.36 per second at 1080P | resolutions 480P, 720P, 1080P; 3 to 15 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9; prompt up to 5,000 characters | Audio: no audio input. Valid combinations: "A character or reference images (up to 9) plus a text prompt that names them Image 1, Image 2..." | | `spicy-cinema-1` | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.36 per second at 1080P | resolutions 480P, 720P, 1080P; 3 to 15 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9; prompt up to 5,000 characters | Audio: no audio input. Valid combinations: "Text prompt only, with aspect_ratio; the clip has native audio" | | `spicy-video-1` | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4; prompt up to 5,000 characters | Audio: driving audio (the clip is the soundtrack and the mouth follows it). Valid combinations: "Text prompt only, with aspect_ratio"; "Text prompt plus driving audio (audio_url): the clip becomes the soundtrack and motion follows it" | | `spicy-motion-draft-1` | video | $0.05 per second at 720P, $0.075 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; prompt up to 1,500 characters; always silent; needs image_url (one of your generated images) | Audio: no audio input. Valid combinations: "First frame (image_url) with a text prompt; the clip is silent" | | `spicy-motion-1` | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 10 seconds; prompt up to 1,500 characters; needs image_url (one of your generated images) | Audio: no audio input. Valid combinations: "First frame (image_url) with a text prompt" | | `spicy-video-edit-1` | video-edit | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4; prompt up to 5,000 characters; the input clip's seconds are billed too, at the same rate; edits your own 2 to 10 second clips (output up to 10 seconds) with up to 4 reference images, aspect_ratio, keep_audio | | | `spicy-cinema-1-edit` | video-edit | $0.28 per second at 720P, $0.48 per second at 1080P | resolutions 720P, 1080P; prompt up to 5,000 characters; the input clip's seconds are billed too, at the same rate; edits your own 3 to 30 second clips (output up to 15 seconds) with up to 5 reference images, keep_audio | | | `spicy-animate-1` | video-edit | $0.24 per second at standard, $0.36 per second at pro | resolutions standard, pro; motion transfer: image_url plus a motion id (GET /v1/videos/motions) or your own 2 to 30 second clip, quality standard or pro, billed per output second; needs image_url (one of your generated images) | | | `spicy-companion-1` | chat | $1 per 1M prompt tokens and $2.8 per 1M completion tokens (minimum $0.001 per request) | | | | `spicy-companion-1-flash` | chat | $0.1 per 1M prompt tokens and $0.8 per 1M completion tokens (minimum $0.0005 per request) | | | | `spicy-chat-1` | chat | $0.8 per 1M prompt tokens ($0.16 when served from cache) and $2.4 per 1M completion tokens (minimum $0.001 per request) | | | | `spicy-voice-2` | speech | $0.40 per 10,000 characters of input text | input up to 5,000 characters; 2 preset voices; `instructions` steer emotion, pace and delivery; inline tags like [whispers], [excited], [giggles] inside `input`; WAV or MP3 URL usable as audio_url on video, or `stream: true` for raw audio as it is synthesized | | | `spicy-voice-2-flash` | speech | $0.30 per 10,000 characters of input text | input up to 5,000 characters; 9 preset voices; `instructions` steer emotion, pace and delivery; inline tags like [whispers], [excited], [giggles] inside `input`; WAV or MP3 URL usable as audio_url on video, or `stream: true` for raw audio as it is synthesized | | | `spicy-voice-1-expressive` | speech | $0.23 per 10,000 characters of input text | input up to 600 characters; 21 preset voices; `instructions` steer emotion, pace and delivery; WAV output, a URL usable as audio_url on video | | | `spicy-voice-1-custom` | speech | $0.23 per 10,000 characters of input text; creating a voice costs $0.40 (designed) or $0.02 (cloned), once | input up to 600 characters; speaks your designed and cloned voices (vc_...); WAV output, a URL usable as audio_url on video | | | `spicy-voice-1` | speech | $0.20 per 10,000 characters of input text | input up to 600 characters; 45 preset voices; WAV output, a URL usable as audio_url on video | | | `spicy-live-1` | realtime | per 1M tokens: text in $0.46, audio in $1.86, text out $1.4, audio out $3.74, billed per turn (about half a cent per minute of conversation; opening a session needs a $0.05 balance) | 28 voices, up to 13 minutes per session, POST /v1/realtime/sessions then a WebSocket | | | `spicy-transcribe-1` | transcription | $0.0042 per minute of audio, billed by the second (minimum $0.0001 per request) | wav, mp3, m4a, ogg, flac, webm and more, up to 10 MB and 5 minutes | | | `spicy-embed-1` | embedding | $0.14 per 1M tokens | | | | `spicy-embed-vision-1` | embedding | $0.18 per 1M text tokens and $0.06 per 1M image tokens | | | Public docs: https://www.spicyapi.com/docs (every page is also available as markdown by appending `.md`). Product overview for agents: https://www.spicyapi.com/llms.txt