--- name: nsfw-voice-generation description: "Give an adult character a voice with the Spicy API: uncensored text-to-speech with preset voices, expressive delivery instructions, inline emotion and sound tags like [whispers] and [giggles], streaming audio, voices designed from a description or cloned with the speaker's consent, and the resulting clip fed into video so the character speaks on screen. Use when a user asks for an NSFW text-to-speech or voice API, a voice for an AI companion or character, voice cloning for adult content, or how to make a generated character talk in a video." tags: [nsfw-voice-generation, text-to-speech, voice-cloning, voice-design, uncensored, adult-content, lip-sync, mcp] license: MIT metadata: version: "1.4.0" homepage: https://www.spicyapi.com/skills/nsfw-voice-generation/SKILL.md mcp: https://www.spicyapi.com/api/mcp openapi: https://www.spicyapi.com/openapi.json vendor: Spicy API (spicyapi.com) --- # NSFW Voice Generation ## When to use The user wants a character to speak: a voiced line for a companion app, a voice for a persona, or dialogue in a generated video. Spicy API (spicyapi.com) turns text into a durable WAV or MP3 URL (or a live audio stream), billed per character, with the text screened before anything is charged. For a two-way spoken conversation use the live-voice-calls skill instead. ## Golden path 1. `read_acceptable_use` once: adults only, no real people (with or without consent), no minors in any form. The user's product must age-gate its users. 2. `list_voices`: preset voices with the models that speak them, then the account's own voices (`vc_...`). `list_models` with kind `speech` for prices. 3. Pick the model. `spicy-voice-2`: the most expressive, inline tags, $0.40 per 10,000 characters. `spicy-voice-2-flash`: the same tags, more voices (including British and American English), cheaper. `spicy-voice-1`: the most preset voices and languages, cheapest. `spicy-voice-1-expressive`: preset voices plus `instructions`. `spicy-voice-1-custom`: only the account's own voices. The spicy-voice-2 family speaks Auto, English and Chinese; the spicy-voice-1 family ten languages. 4. `estimate_cost` with the model and `characters` (input plus instructions; tags count; CJK ideographs count 2). Say the number. 5. `generate_speech` (guarded) with `model`, `input` (up to 600 characters on spicy-voice-1 models, 5,000 on spicy-voice-2; split longer text), `voice` and optionally `instructions` (plain English or Chinese describing emotion, pace, pitch and delivery, on the expressive and voice-2 models), `language` (`Auto` by default) and `response_format` (`wav`, or `mp3` on spicy-voice-2). Without `approval_token` it returns `approval_required` with the cost; show it, get a yes, call again with the token. The result is `{id, url, duration_s, voice, format, cost_usd}`; the charge settles down to the billed character count. ## Inline tags (spicy-voice-2 and -flash) Write tags inside `input`. Control tags change the delivery until the next tag: `[sad]`, `[amazed]`, `[deep and loud shouting]`, `[trembling]`, `[angry]`, `[excited]`, `[sarcastic]`, `[curious]`, `[like dracula]`, `[bored]`, `[tired]`, `[scornful]`, `[shouting]`, `[asmr]`, `[panicked]`, `[mischievously]`, `[empathetic]`, `[whispers]`, `[reluctantly]`, `[crying]`, `[serious]`, `[very slowly]`, `[very fast]`. Sound tags insert a sound at that point: `[gasp]`, `[sighing]`, `[clears throat]`, `[giggles]`, `[laughing]`, `[cough]`, `[snorts]`. Example: `[excited]Hey you, I've been waiting all night.[giggles] Come here.[whispers] Closer.` Tags are billed as characters. `list_models` shows them under `speech.tags`. ## Streaming in an app For a companion app that should start talking at once, call `POST https://api.spicyapi.com/v1/audio/speech` with `stream: true` on a spicy-voice-2 model (the tool always returns a stored clip). The response body is raw audio as it is synthesized, first audio in about 1.5 seconds: `response_format` `pcm` (default, 24 kHz 16-bit mono little-endian), `wav` or `mp3`. Headers carry `x-spicyapi-audio-id` (the clip is stored afterwards under that id, `GET /v1/audio/{id}`), `x-spicyapi-sample-rate` and `x-spicyapi-cost-usd`. Stream from your backend and relay to the client; never put the key in the app. ## A character's own voice - Design (`create_voice` with `name` and `description`, $0.40 once): describe vocal qualities only, such as gender, age range, pitch, pace, tone and accent. A name or description that names, references or imitates a real person is refused (`real_person`); never try to recreate a celebrity or anyone the user knows. The result has a `preview_url` to play to the user. - Clone (`create_voice` with `name`, `audio_url` and `consent: true`, $0.02 once): the clip must be one of the account's own (`upload_audio`, then `list_audio`), ideally a clean 10 to 20 second sample. `consent: true` is the user's attestation that the voice is their own or that they hold the speaker's written consent; it is stored with the voice. Ask the user explicitly and never set it on their behalf. A missing flag is refused before any charge. - Speak with it on `spicy-voice-1-custom`, passing the `vc_...` id as `voice`. `delete_voice` (guarded) removes it. ## Make the character speak in a video The speech `url` is already one of the account's audio clips: pass it exactly as `audio_url` to `generate_video`. It must be 2 to 30 seconds long. - Driving audio (`spicy-motion-2`): with a first frame from `list_images`, the clip becomes the soundtrack and the mouth follows the words. The usual choice for a talking shot. - Reference audio (`spicy-motion-3`, `spicy-motion-3-fast`): with reference images or a `character` and the dialogue written in the prompt, never with a first frame; the model uses the clip for voice, tone and beat and generates the sound itself. ## Driving audio vs reference audio `audio_url` does two different things, and `list_models` says which under `inputs.audio.mode`. Driving audio (`spicy-motion-2`, `spicy-video-1`): the clip is the soundtrack and the mouth follows it, one clip, alongside a first frame where the model takes one. Reference audio (`spicy-motion-3`, `spicy-motion-3-fast`): the model generates its own sound and uses the clip for voice, tone and beat while the words come from the prompt; up to 5 clips (15 s total) in `audio_urls`, only with a text prompt or references. On `spicy-motion-3` a first frame (`image_url`) pairs only with `last_frame_url`: no references, clip or audio in the same request; for lip-sync from a first frame use driving audio on `spicy-motion-2`. `spicy-motion-1` and `spicy-character-video-1` take no audio. Pass `audio_mode` (`driving` or `reference`) to assert what you expect; a mismatch is rejected naming a model that offers it. Read `inputs.combinations` before calling; an invalid mix is rejected locally with those sentences quoted, before any spend. ## Speech models | Model | Kind | Price | Limits | Inputs (audio mode, valid combinations) | |---|---|---|---|---| | `spicy-voice-2` | speech | $0.40 per 10,000 characters of input text | input up to 5,000 characters; 2 preset voices; `instructions` steer emotion, pace and delivery; inline tags like [whispers], [excited], [giggles] inside `input`; WAV or MP3 URL usable as audio_url on video, or `stream: true` for raw audio as it is synthesized | | | `spicy-voice-2-flash` | speech | $0.30 per 10,000 characters of input text | input up to 5,000 characters; 9 preset voices; `instructions` steer emotion, pace and delivery; inline tags like [whispers], [excited], [giggles] inside `input`; WAV or MP3 URL usable as audio_url on video, or `stream: true` for raw audio as it is synthesized | | | `spicy-voice-1-expressive` | speech | $0.23 per 10,000 characters of input text | input up to 600 characters; 21 preset voices; `instructions` steer emotion, pace and delivery; WAV output, a URL usable as audio_url on video | | | `spicy-voice-1-custom` | speech | $0.23 per 10,000 characters of input text; creating a voice costs $0.40 (designed) or $0.02 (cloned), once | input up to 600 characters; speaks your designed and cloned voices (vc_...); WAV output, a URL usable as audio_url on video | | | `spicy-voice-1` | speech | $0.20 per 10,000 characters of input text | input up to 600 characters; 45 preset voices; WAV output, a URL usable as audio_url on video | | ## Tools used by this skill - `list_models`: Every model with its kind (image, image-edit, video, video-edit, chat, speech, realtime, transcription, embedding), USD price (per image, per second by resolution or quality, per 1M tokens, per 10,000 characters of speech input, or per minute of audio), limits (image sizes, outputs per call, resolutions, clip lengths, aspect ratios, prompt character caps), whether it needs or accepts an input image or references, and example clips. Video rows carry `inputs`: which of first frame, last frame, reference images, clip and audio the model takes, `inputs.audio.mode` (driving: the clip is the soundtrack and the mouth follows it; reference: the model generates sound and uses the clip for voice, tone and beat; none) and `inputs.combinations`, the valid input combinations in plain sentences. Read those before generate_video. `silent: true` marks a video model that renders without sound; `bills_input_video_seconds: true` one that also bills the seconds of an input clip. Video-edit rows carry `video_edit` (instruction or motion mode, clip lengths, reference images, whether input seconds are billed) for edit_video. Speech rows carry `speech`: preset voices, whether `instructions` are accepted, custom voices, languages, inline `tags` and voice creation fees. Realtime rows (spicy-live-1) are live voice calls over a WebSocket: a server or app opens a session with POST /v1/realtime/sessions and connects the returned ticket URL; they cannot be driven from a tool call. Prices here are exactly what the API bills. Works with a $0 balance. Arguments: `kind`?: image | image-edit | video | video-edit | chat | speech | realtime | transcription | embedding (Filter by kind.). - `estimate_cost`: USD cost of a request before making it, from the same price table the gateway bills: images by count, video by seconds and resolution (plus the input clip or action clip on models that bill it), video edits by clip length (input plus output seconds for instruction edits, output seconds for motion transfer), chat by tokens, speech by characters of input (plus instructions), transcription by seconds of audio, embeddings by tokens, live calls by minutes. Use it to quote the user and to check against get_account before a spend. Works with a $0 balance. Arguments: `model`: string (A model id from list_models, e.g. spicy-image-1, spicy-image-1-pro, spicy-image-photo-1.); `n`?: integer (Images per call (image models), clamped to the model's maximum.); `resolution`?: string (Video or video-edit resolution, e.g. 720P (default) or 1080P.); `quality`?: standard | pro (Motion transfer tier on spicy-animate-1, default standard.); `duration`?: integer (Video seconds, default 5; -1 for a model-chosen length where supported.); `input_seconds`?: number (Length of the input clip in seconds: the clip to edit or motion clip on video-edit models (default 5), or a video_url on video models that bill input seconds.); `motion`?: string (A library motion id (list_motions) on spicy-animate-1; its length sets the output length.); `action`?: string (An action id (list_actions) on spicy-motion-3 and -fast: its reference clip seconds (4 on most) are billed as input, and duration defaults to 8.); `fps`?: any (Video frame rate: 30 (default) or 60. 60 smooths motion with frame interpolation for 20% on top of the video price (input or action clip seconds included); audio is kept. If that step fails the 30 fps clip is delivered and the fee refunded.); `tokens`?: integer (Chat: total tokens, default 1000. Embeddings: text tokens, default 1000.); `image_tokens`?: integer (Embeddings on spicy-embed-vision-1: image tokens, default 0.); `characters`?: integer (Speech models: characters of input plus instructions (CJK ideographs and inline tags count too), default 100.); `audio_seconds`?: number (Transcription: seconds of audio, default 60.); `minutes`?: number (Live calls (spicy-live-1): minutes of conversation, default 1.); `count`?: integer (How many such requests, default 1.). - `list_images`: The account's generated images, newest first. These are the only URLs accepted as image inputs by edit_image, generate_video, edit_video and create_embeddings (no uploads, no third-party URLs). Use it to pick a first frame to animate or a performer for motion transfer. Arguments: `limit`?: integer (Default 20.). - `generate_video` (guarded): Image to video, text to video or references to video. Asynchronous: returns a task id at once with the price already debited; poll get_job every 10 to 15 seconds until succeeded (output.video_url, usage.output_seconds) or failed (refunded). Every input must be the account's own: images from list_images, clips from its own finished tasks (video_url), audio from upload_audio or generate_speech. Price = seconds x the per-second rate for the resolution; models with bills_input_video_seconds (spicy-motion-3, -fast) also bill the seconds of a video_url clip or of an action clip. duration -1 on spicy-motion-3 lets the model pick the length, debited at the maximum and refunded down to what is rendered. Text to video: spicy-motion-3, -fast, spicy-video-1, spicy-cinema-1. Image to video: spicy-motion-2, spicy-cinema-1-image, spicy-motion-3. Character or references to video: spicy-character-video-1, spicy-cinema-1-character (up to 9 references named Image 1, Image 2 in the prompt), spicy-motion-3. Cheap previews: spicy-motion-draft-1 is a silent draft tier from a first frame; render the keeper on spicy-motion-2 or 3. The spicy-cinema-1 family runs 3 to 15 seconds at 480P to 1080P with native audio always on, no negative_prompt, no enhance_prompt and no audio inputs. Which inputs may travel together differs per model: read list_models inputs.combinations first. Audio comes in two modes (inputs.audio.mode): driving (spicy-motion-2, spicy-video-1: the clip is the soundtrack, the mouth follows it, works with a first frame) and reference (spicy-motion-3, spicy-motion-3-fast: the model generates sound and uses the clip for voice, tone and beat; only with a text prompt or references, never with a first frame). An invalid combination is rejected here before any spend with the valid combinations spelled out. Actions (spicy-motion-3, -fast): pass an action id from list_actions and its reference clip of that sex act supplies the camera, pose and motion second for second while image_url supplies the woman, e.g. {model: "spicy-motion-3", action: "pov-missionary", image_url: , prompt: "hotel room at night, red lingerie"}. To change a finished clip or transfer motion onto a person, use edit_video. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token. Arguments: `model`: string (A video model id, e.g. spicy-motion-2, spicy-motion-3, spicy-cinema-1, spicy-motion-draft-1.); `prompt`?: string (What should happen in the clip; required unless image_url or action is sent. With action it is only the scene (setting, outfit, lighting, extra direction). Capped per model (limits.maxPromptChars: 1500 on spicy-motion-1 and spicy-motion-draft-1, 20000 on spicy-motion-3 and -fast, 5000 on the rest).); `image_url`?: string (First frame, from list_images or a generate_image result. list_models inputs.first_frame says whether it is required, optional or not taken. Not with reference_image_urls; on spicy-motion-3 a first frame pairs only with last_frame_url.); `last_frame_url`?: string (Closing frame (one of the account's images); the clip travels from image_url to it. Models with inputs.last_frame true; needs image_url (or a continuation video_url).); `reference_image_urls`?: array (Up to inputs.reference_images of the account's images guiding identity and look without fixing the first shot (a character counts as one). Exclusive with image_url. See inputs.combinations.); `video_url`?: string (output.video_url of one of the account's own finished tasks. inputs.continuation_clip (spicy-motion-2): continues the clip in place of image_url (2 to 10 s in; duration is the whole output incl. the clip, billed in full). inputs.reference_video (spicy-motion-3, -fast): reference clip to extend or edit (15 s max; clip plus duration within 30 s), never with image_url.); `audio_url`?: string (A clip from upload_audio or generate_speech (its url, exact; 2 to 30 seconds). What it does depends on inputs.audio.mode: driving (spicy-motion-2, spicy-video-1) makes it the soundtrack with the mouth following it, and works with a first frame; reference (spicy-motion-3, -fast) guides voice, tone and beat while the model generates the sound, and only goes with a text prompt or references, never with image_url. Rejected on models with mode none (spicy-motion-1, spicy-character-video-1).); `audio_urls`?: array (Several upload_audio clips, up to inputs.audio.max_clips (5 on spicy-motion-3 and -fast, 15 s total); other models take a single audio_url.); `audio_mode`?: driving | reference (Optional guard: what you expect the audio to do. Must equal the model's inputs.audio.mode or the call is rejected naming a model that offers it. Omit to use the model's mode.); `resolution`?: string (480P, 720P (default) or 1080P, within the model's limits.resolutions.); `duration`?: integer (Seconds, default 5 (8 with action), within the model's limits (2 to 30; 3 to 15 on the spicy-cinema-1 family). -1 where inputs.smart_duration is true (spicy-motion-3, -fast): the model chooses; debited at the maximum, settled to the rendered length.); `aspect_ratio`?: string (Text-to-video, references and action only (no first frame or video_url), where inputs.aspect_ratio is true; values in limits.aspectRatios. 16:9, 9:16, 1:1, 4:3, 3:4 on spicy-video-1 (default 16:9); plus 21:9 on spicy-motion-3 and -fast, where omitting it lets the model choose; plus 4:5, 5:4 and 9:21 on spicy-cinema-1 and spicy-cinema-1-character.); `negative_prompt`?: string (What to avoid. Ignored on spicy-motion-3, spicy-motion-3-fast and the spicy-cinema-1 family.); `enhance_prompt`?: boolean (Let the model expand the prompt with lighting, camera and detail cues. Default false. Not on the spicy-cinema-1 family.); `audio`?: boolean (Generate audio where the model supports it. spicy-motion-draft-1 is always silent; the spicy-cinema-1 family always has audio.); `seed`?: integer; `character`?: string (A saved character id (list_characters). Only on models that take references: spicy-character-video-1 (no first frame needed), spicy-cinema-1-character, spicy-motion-3, spicy-motion-3-fast.); `action`?: string (An action id from list_actions, e.g. pov-missionary. Runs on spicy-motion-3 and spicy-motion-3-fast: the action's reference clip sets camera, pose and motion; image_url becomes the woman (her identity, not a first frame) and prompt is optional, only the scene. Not with video_url or last_frame_url. duration defaults to 8 and aspect_ratio to the action's own (any supported ratio may be requested); clip plus duration within 30 s. The clip's seconds are billed as input seconds on top of the output.); `style`?: photorealistic | studio | anime | 3d | cartoon (Art style (list_styles): photorealistic, studio (glossy glamour), anime, 3d (animated-film CGI) or cartoon (2D adult cartoon). Prompt text only, same price.); `fps`?: any (Video frame rate: 30 (default) or 60. 60 smooths motion with frame interpolation for 20% on top of the video price (input or action clip seconds included); audio is kept. If that step fails the 30 fps clip is delivered and the fee refunded.); `approval_token`?: string (Approval token from a previous approval_required response, after the user said yes.). - `upload_audio` (guarded): Bring in a WAV or MP3 (2 to 30 seconds, 15 MB max) from a public https URL for lip-sync and audio-driven video. The clip is transcribed and the transcript screened like a prompt; flat $0.01 for that screening. A clip that fails the screen is discarded with the same 422 shape as a blocked prompt; a 503 means screening was unavailable and nothing was charged. The returned url is what audio_url and audio_urls on generate_video accept; external audio URLs are rejected there. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token. Arguments: `url`: string (Public https URL of the WAV or MP3 to fetch (private and plain-IP hosts are refused).); `approval_token`?: string (Approval token from a previous approval_required response, after the user said yes.). - `list_audio`: The account's audio clips, newest first: uploads and generated speech (source upload or speech), with id, url, duration_s, mime and transcript. These urls are the only values generate_video accepts as audio_url or in audio_urls. Arguments: `limit`?: integer (Default 20.). - `generate_speech` (guarded): Text to speech, synchronous. Returns {id, url, duration_s, voice, format, cost_usd}: a durable clip (WAV 24 kHz mono, or MP3 on the spicy-voice-2 family) that is also one of the account's audio clips, so its url goes straight into generate_video as audio_url (driving audio on spicy-motion-2: the mouth follows the words, with a first frame; reference audio on spicy-motion-3: voice, tone and beat, no first frame). Video audio must be 2 to 30 seconds. Models: spicy-voice-1 (preset voices), spicy-voice-1-expressive (preset voices plus instructions for emotion, pace and delivery), spicy-voice-1-custom (the account's own vc_... voices from create_voice), spicy-voice-2 and spicy-voice-2-flash (their own preset voices, instructions, languages Auto, English and Chinese, up to 5,000 characters, and inline tags inside input: control tags switch delivery until the next tag, [sad] [amazed] [deep and loud shouting] [trembling] [angry] [excited] [sarcastic] [curious] [like dracula] [bored] [tired] [scornful] [shouting] [asmr] [panicked] [mischievously] [empathetic] [whispers] [reluctantly] [crying] [serious] [very slowly] [very fast]; sound tags insert a sound, [gasp] [sighing] [clears throat] [giggles] [laughing] [cough] [snorts]; e.g. "[excited]Hey you.[giggles] Come here.[whispers] Closer."). Tags count as billed characters. Input caps: 600 characters on the spicy-voice-1 family, 5,000 on spicy-voice-2. Billed per 10,000 characters of input plus instructions (CJK ideographs count 2), settled down to the billed count. Streaming (stream: true) is a REST feature for apps (POST /v1/audio/speech); this tool always returns a stored clip. The text is screened like a prompt; blocked text costs nothing. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token. Arguments: `model`: string (A speech model id: spicy-voice-1, spicy-voice-1-expressive, spicy-voice-1-custom, spicy-voice-2 or spicy-voice-2-flash.); `input`: string (The text to speak: up to 600 characters on the spicy-voice-1 family, 5,000 on spicy-voice-2 and -flash, where inline tags like [whispers] or [giggles] are performed.); `voice`: string (A preset voice name the model speaks (list_voices), e.g. Cherry on spicy-voice-1, Lingxin on spicy-voice-2, Eva on spicy-voice-2-flash, or a vc_... id on spicy-voice-1-custom.); `instructions`?: string (spicy-voice-1-expressive and the spicy-voice-2 family: how to say it (emotion, pace, pitch, delivery) in plain English or Chinese.); `language`?: string (Auto (default), English, Chinese, German, Italian, Portuguese, Spanish, Japanese, Korean, French or Russian. The spicy-voice-2 family takes Auto, English or Chinese.); `response_format`?: wav | mp3 (wav (default) on every model; mp3 on the spicy-voice-2 family.); `user`?: string (Your own id for the end user making this request.); `approval_token`?: string (Approval token from a previous approval_required response, after the user said yes.). - `list_voices`: Voices for generate_speech: every preset voice with the models that speak it (spicy-voice-1 has them all, spicy-voice-1-expressive a subset), then the account's designed and cloned voices (vc_... ids, spoken by spicy-voice-1-custom), newest first. Works with a $0 balance. Arguments: no arguments. - `create_voice` (guarded): Create a custom voice for spicy-voice-1-custom, returned as a vc_... id. Design: name plus a description of vocal qualities (gender, age range, pitch, pace, tone, accent); a description that names, references or imitates a real person is refused. Clone: name plus audio_url (one of the account's clips from list_audio or upload_audio, 10 MB max) plus consent: true, the user's attestation that the voice is their own or that they hold the speaker's written consent; it is stored with the voice. Never set consent without the user confirming it. One-off fee per voice, refunded if creation fails. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token. Arguments: `name`: string (A label for the voice, e.g. Mara.); `description`?: string (Design: the voice's qualities. Omit when cloning.); `preview_text`?: string (Design: a line spoken in the preview clip.); `audio_url`?: string (Clone: a clip url from list_audio. Omit when designing.); `consent`?: boolean (Clone: must be true, the user's attestation that the voice is theirs or that the speaker gave written consent.); `language`?: string (English (default), Chinese, German, Italian, Portuguese, Spanish, Japanese, Korean, French or Russian.); `user`?: string (Your own id for the end user making this request.); `approval_token`?: string (Approval token from a previous approval_required response, after the user said yes.). - `delete_voice` (guarded): Delete a designed or cloned voice (vc_...) so it can no longer be spoken. Clips already generated with it are unaffected. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token. Arguments: `id`: string (A vc_... id from list_voices.); `approval_token`?: string (Approval token from a previous approval_required response, after the user said yes.). - `get_job`: Status of a video task from generate_video or edit_video (queued, processing, finalizing, succeeded with output.video_url and usage, or failed and refunded) or of a logged image, chat, speech, transcription or embedding request by id. Arguments: `id`: string (Task or request id, e.g. sj_...). - `get_topup_link`: Checkout URL to add funds by card (hosted checkout) or crypto for a preset amount ($50, $100, $250, $500, $1000; minimum $50). The USER opens it and pays in the browser; never enter payment details yourself. The balance is credited by the payment webhook; call get_account afterwards. If it returns commitment_required, the user must first sign the one-page Customer Commitment Letter on the Billing page; an agent cannot sign it. Works with a $0 balance. Arguments: `amount_usd`?: integer (One of 50, 100, 250, 500, 1000. Default 50.); `method`?: card | crypto (Default card.). - `read_acceptable_use`: The acceptable use policy as markdown: prohibited content (minors in any form, images or video of real people with or without consent, non-consensual scenarios and the rest), age-verification duties for the customer's product, enforcement. Read it before the first generation and tell the user what their product must do. Works with a $0 balance. Arguments: no arguments. ## Setup (once per user) 1. Key: the user signs in once at https://www.spicyapi.com/auth (Google or email). The account is live at once with a $0 balance and a Default key shown once at https://www.spicyapi.com/dashboard/api-keys. Keep the key in an environment variable (`SPICYAPI_KEY`), never in chat, never in client-side code. With one key you can mint more with `create_api_key`. 2. Connect. MCP (Streamable HTTP): `https://www.spicyapi.com/api/mcp` with header `Authorization: Bearer `. Claude Code: `claude mcp add --transport http spicyapi https://www.spicyapi.com/api/mcp --header "Authorization: Bearer sk-spicy-..."`. Cursor, Codex and any URL-plus-headers client: same URL and header. REST instead of MCP: `POST https://www.spicyapi.com/api/v1/tools/` with the same header and a JSON body of arguments; `GET https://www.spicyapi.com/api/v1/tools` lists them; OpenAPI 3.1 at https://www.spicyapi.com/openapi.json. OAuth 2.1 clients (Claude custom connectors, ChatGPT, directory scanners) need only the URL: the endpoint advertises its authorization server, the user signs in and consents in the browser, and the token it returns is an API key they can revoke in the dashboard. 3. Dry run first: a key created with Sandbox ticked (or `create_api_key` with `sandbox: true`) answers every endpoint and every tool from fixture output flagged `sandbox: true`, needs no balance and bills nothing. Build against it, then swap the key. Never present sandbox output to the user as a real generation. 4. Money: `get_account` shows the balance; `estimate_cost` prices a request from the same table the API bills; `get_topup_link` returns a card or crypto checkout URL (presets $50, $100, $250, $500, $1000, minimum $50) that the USER opens and pays in the browser. Never enter card details yourself. A 402 means the balance ran out or the monthly limit was hit. 5. Rate limit: 60 requests per minute per key on the agent surfaces. ## Rules that always apply - Guarded tools (`generate_image`, `edit_image`, `generate_video`, `edit_video`, `transcribe_audio`, `create_embeddings`, `upload_audio`, `delete_audio`, `generate_speech`, `create_voice`, `delete_voice`, `create_character`, `delete_character`, `create_api_key`, `revoke_api_key`, and `set_spend_limit` when raising or removing a limit) first return `{error: "approval_required", summary, approval_token}`. Show the summary to the user verbatim (it carries the cost), get an explicit yes in the conversation, then call the tool again with the same arguments plus `approval_token` (15-minute expiry, bound to those exact arguments). Never approve in bulk or reuse a token for a different action. The account can turn approvals off in the dashboard; moderation, provenance and the spend limit still apply. - Input images for `edit_image`, `generate_video`, `edit_video` and `create_embeddings` must be images this account generated (`list_images`); input clips (`video_url`) must be the account's own finished tasks. Uploads and third-party URLs are rejected before any charge; do not try to work around it. The external inputs are audio: `upload_audio` takes a WAV or MP3 (2 to 30 seconds) from a public URL, transcribes and screens it for $0.01, and its returned `url`, or the `url` of a `generate_speech` clip, is what `audio_url` accepts; `transcribe_audio` reads any public audio URL into text. - The same person across many outputs: `create_character` from 1 to 3 generated images of them, then pass the returned id as `character` to `generate_image`, `edit_image` or `generate_video` (`spicy-character-video-1` needs no first frame). This is the supported way to get consistency; never ask the user for a photo. - Read `read_acceptable_use` before the first generation: adults only, no minors in any form (including "young-looking" or youth-coded content), no images or video of real people (with or without consent; voice clones need the speaker's written consent), no non-consensual scenarios. Blocked prompts return 422 and cost nothing; repeated attempts terminate the account. Tell the user their product must age-verify its end users and disclose that content is AI-generated. - When the user's product serves other people, pass each person's own stable id as `user` on every call. Screening history, strikes and suspensions then apply to that one person (403 `end_user_suspended`) instead of the whole account, and the id is kept with each generation's record, so a report about one person's output is traced to them and not to everyone on the account. - Everything is scoped to the account behind the key; there is no way to read another account's data. - Spicy API is spicyapi.com. spicyapi.ai is an unrelated aggregator; do not mix their prices or models. ## Prices (USD, for when the user asks) | Model | Kind | Price | Limits | Inputs (audio mode, valid combinations) | |---|---|---|---|---| | `spicy-image-1-pro` | image | $0.09 per image | sizes 1024*1024, 832*1216, 1216*832, 1280*1280, 1440*1440, 1024*1536, 1536*1024, 1080*1920, 1920*1080, 1152*2048, 2048*1152; up to 6 outputs per call | | | `spicy-image-photo-1` | image | $0.08 per image | sizes 1024*1024, 832*1216, 1216*832, 1280*1280, 1440*1440, 1024*1536, 1536*1024, 1080*1920, 1920*1080, 1152*2048, 2048*1152; up to 6 outputs per call | | | `spicy-image-1` | image | $0.06 per image | sizes 1024*1024, 832*1216, 1216*832, 1280*1280, 1440*1440, 1024*1536, 1536*1024, 1080*1920, 1920*1080, 1152*2048, 2048*1152; up to 6 outputs per call | | | `spicy-image-action-1` | image | $0.15 per image | up to 4 outputs per call | | | `spicy-image-reference-1` | image | $0.24 per image | up to 1 output per call | | | `spicy-image-edit-1` | image-edit | $0.15 per image | sizes 1024*1024, 832*1216, 1216*832, 1280*1280, 1440*1440, 1024*1536, 1536*1024, 1080*1920, 1920*1080, 1152*2048, 2048*1152; up to 6 outputs per call; needs image_url (one of your generated images) | | | `spicy-motion-3` | video | $0.1 per second at 480P, $0.2 per second at 720P, $0.4 per second at 1080P | resolutions 480P, 720P, 1080P; 2 to 30 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 21:9; prompt up to 20,000 characters; reference clip (video_url) up to 15 seconds; the input clip's seconds are billed too, at the same rate; image_url optional (text-to-video without it) | Audio: reference audio (the model generates sound and uses the clip for voice, tone and beat), up to 5 clips and 15s total. Valid combinations: "Text prompt only, with aspect_ratio; dialogue written in the prompt is spoken natively"; "First frame (image_url), optionally a last frame (last_frame_url); no references or audio in the same request"; "Reference images (up to 10, or a character) plus a text prompt: same person, framing chosen by the model"; "Reference images, a reference video (video_url, 15s max) and reference audio (up to 5 clips, 15s total) in any mix, with a text prompt; audio guides voice, tone and beat while the words come from the prompt" | | `spicy-motion-3-fast` | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.56 per second at 1080P | resolutions 480P, 720P, 1080P; 2 to 30 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 21:9; prompt up to 20,000 characters; reference clip (video_url) up to 15 seconds; the input clip's seconds are billed too, at the same rate; image_url optional (text-to-video without it) | Audio: reference audio (the model generates sound and uses the clip for voice, tone and beat), up to 5 clips and 15s total. Valid combinations: "Text prompt only, with aspect_ratio; dialogue written in the prompt is spoken natively"; "First frame (image_url), optionally a last frame (last_frame_url); no references or audio in the same request"; "Reference images (up to 10, or a character) plus a text prompt: same person, framing chosen by the model"; "Reference images, a reference video (video_url, 15s max) and reference audio (up to 5 clips, 15s total) in any mix, with a text prompt; audio guides voice, tone and beat while the words come from the prompt" | | `spicy-cinema-1-image` | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.36 per second at 1080P | resolutions 480P, 720P, 1080P; 3 to 15 seconds; prompt up to 5,000 characters; needs image_url (one of your generated images) | Audio: no audio input. Valid combinations: "First frame (image_url) with a text prompt; the output keeps the image's shape" | | `spicy-motion-2` | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; prompt up to 5,000 characters; needs image_url (one of your generated images) | Audio: driving audio (the clip is the soundtrack and the mouth follows it). Valid combinations: "First frame (image_url), optionally a last frame (last_frame_url)"; "First frame plus driving audio (audio_url): the clip becomes the soundtrack and the mouth follows it"; "First frame, last frame and driving audio together"; "A clip to continue (video_url, 2 to 10s) instead of a first frame, optionally with a last frame" | | `spicy-character-video-1` | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; prompt up to 5,000 characters; image_url optional (text-to-video without it) | Audio: no audio input. Valid combinations: "A character (or reference images) plus a text prompt"; "A character plus a first frame (image_url) that opens the clip" | | `spicy-cinema-1-character` | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.36 per second at 1080P | resolutions 480P, 720P, 1080P; 3 to 15 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9; prompt up to 5,000 characters | Audio: no audio input. Valid combinations: "A character or reference images (up to 9) plus a text prompt that names them Image 1, Image 2..." | | `spicy-cinema-1` | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.36 per second at 1080P | resolutions 480P, 720P, 1080P; 3 to 15 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9; prompt up to 5,000 characters | Audio: no audio input. Valid combinations: "Text prompt only, with aspect_ratio; the clip has native audio" | | `spicy-video-1` | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4; prompt up to 5,000 characters | Audio: driving audio (the clip is the soundtrack and the mouth follows it). Valid combinations: "Text prompt only, with aspect_ratio"; "Text prompt plus driving audio (audio_url): the clip becomes the soundtrack and motion follows it" | | `spicy-motion-draft-1` | video | $0.05 per second at 720P, $0.075 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; prompt up to 1,500 characters; always silent; needs image_url (one of your generated images) | Audio: no audio input. Valid combinations: "First frame (image_url) with a text prompt; the clip is silent" | | `spicy-motion-1` | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 10 seconds; prompt up to 1,500 characters; needs image_url (one of your generated images) | Audio: no audio input. Valid combinations: "First frame (image_url) with a text prompt" | | `spicy-video-edit-1` | video-edit | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4; prompt up to 5,000 characters; the input clip's seconds are billed too, at the same rate; edits your own 2 to 10 second clips (output up to 10 seconds) with up to 4 reference images, aspect_ratio, keep_audio | | | `spicy-cinema-1-edit` | video-edit | $0.28 per second at 720P, $0.48 per second at 1080P | resolutions 720P, 1080P; prompt up to 5,000 characters; the input clip's seconds are billed too, at the same rate; edits your own 3 to 30 second clips (output up to 15 seconds) with up to 5 reference images, keep_audio | | | `spicy-animate-1` | video-edit | $0.24 per second at standard, $0.36 per second at pro | resolutions standard, pro; motion transfer: image_url plus a motion id (GET /v1/videos/motions) or your own 2 to 30 second clip, quality standard or pro, billed per output second; needs image_url (one of your generated images) | | | `spicy-companion-1` | chat | $1 per 1M prompt tokens and $2.8 per 1M completion tokens (minimum $0.001 per request) | | | | `spicy-companion-1-flash` | chat | $0.1 per 1M prompt tokens and $0.8 per 1M completion tokens (minimum $0.0005 per request) | | | | `spicy-chat-1` | chat | $0.8 per 1M prompt tokens ($0.16 when served from cache) and $2.4 per 1M completion tokens (minimum $0.001 per request) | | | | `spicy-voice-2` | speech | $0.40 per 10,000 characters of input text | input up to 5,000 characters; 2 preset voices; `instructions` steer emotion, pace and delivery; inline tags like [whispers], [excited], [giggles] inside `input`; WAV or MP3 URL usable as audio_url on video, or `stream: true` for raw audio as it is synthesized | | | `spicy-voice-2-flash` | speech | $0.30 per 10,000 characters of input text | input up to 5,000 characters; 9 preset voices; `instructions` steer emotion, pace and delivery; inline tags like [whispers], [excited], [giggles] inside `input`; WAV or MP3 URL usable as audio_url on video, or `stream: true` for raw audio as it is synthesized | | | `spicy-voice-1-expressive` | speech | $0.23 per 10,000 characters of input text | input up to 600 characters; 21 preset voices; `instructions` steer emotion, pace and delivery; WAV output, a URL usable as audio_url on video | | | `spicy-voice-1-custom` | speech | $0.23 per 10,000 characters of input text; creating a voice costs $0.40 (designed) or $0.02 (cloned), once | input up to 600 characters; speaks your designed and cloned voices (vc_...); WAV output, a URL usable as audio_url on video | | | `spicy-voice-1` | speech | $0.20 per 10,000 characters of input text | input up to 600 characters; 45 preset voices; WAV output, a URL usable as audio_url on video | | | `spicy-live-1` | realtime | per 1M tokens: text in $0.46, audio in $1.86, text out $1.4, audio out $3.74, billed per turn (about half a cent per minute of conversation; opening a session needs a $0.05 balance) | 28 voices, up to 13 minutes per session, POST /v1/realtime/sessions then a WebSocket | | | `spicy-transcribe-1` | transcription | $0.0042 per minute of audio, billed by the second (minimum $0.0001 per request) | wav, mp3, m4a, ogg, flac, webm and more, up to 10 MB and 5 minutes | | | `spicy-embed-1` | embedding | $0.14 per 1M tokens | | | | `spicy-embed-vision-1` | embedding | $0.18 per 1M text tokens and $0.06 per 1M image tokens | | | Public docs: https://www.spicyapi.com/docs (every page is also available as markdown by appending `.md`). Product overview for agents: https://www.spicyapi.com/llms.txt