---
name: oceantoken-media
description: Create images, video, voice-overs and subtitles with OceanToken models and save them into the project. Use when the user wants a generated or edited image (hero, illustration, icon, product shot, background, a change to an existing picture), a video clip (from text, or animating an image), narration / text-to-speech, or a transcript or subtitles, and the OceanToken tools are connected. Also for Chinese requests such as 生成图片、做图、改图、海报、生成视频、图生视频、配音、语音合成、朗读、字幕、转写.
metadata:
openclaw:
homepage: https://oceantoken.ai
requires:
bins:
- curl
---
# Images, video and audio with OceanToken
Every generation spends the user's OceanToken balance. Work like a producer on a
budget: settle the brief, price it, generate once well, save the result, report.
If the host also has a built-in image tool, use OceanToken when the user asks
for OceanToken or a specific model, wants several options or a batch, needs
video, audio or subtitles, or wants the cost up front.
## Connect first
These steps use the OceanToken MCP tools (`search_models`, `estimate_cost`,
`generate_image`, ...). If they are not available, the server is not connected
yet. It is `https://mcp.oceantoken.ai/mcp` (streamable HTTP) with an OAuth
sign-in. In OpenClaw:
```bash
openclaw mcp set oceantoken '{"url":"https://mcp.oceantoken.ai/mcp","transport":"streamable-http","auth":"oauth"}'
openclaw mcp login oceantoken
```
For other hosts see https://github.com/NextFormAI/oceantoken-plugins#install.
## 1. Settle the brief
Know before generating: what the asset is for, where the file goes, its shape
(aspect ratio, pixel size, duration), and the look. Read the project first
(existing assets, design tokens, the `
`/CSS that will hold it) so you ask
the user only what the code cannot tell you.
## 2. Pick the model
- `search_models` with `type` = `image`, `video`, `speech` or `transcription`.
`featured` models are OceanToken's directly served defaults; `sort="price"`
finds cheap ones. Never invent a model id.
- Call `get_model` before using non-default options: it lists the sizes, aspect
ratios, resolutions, durations, audio and first/last-frame support the model
really accepts. Sending an option the model lacks fails the request.
- Rough guide (confirm with search): text-heavy or precise edits favour
GPT Image; illustration and vector styles favour Recraft; photoreal and
product shots favour Seedream or Gemini image models; video: Seedance is the
value pick, Veo and Sora for top quality.
## 3. Price it, and say the price
Call `estimate_cost` for any video, any batch (`n > 1`), 4K images and long
speech, and tell the user the figure and its `confidence`. Ask before spending
more than about $2 unless the user already agreed to a budget. When they gave a
budget, pass it as `max_cost_usd`: the tool then refuses anything over it before
spending. Never loop regenerations without telling them: each attempt costs money.
## 4. Generate
- **Image**: `generate_image(model, prompt, size, quality, n)`. A good prompt
names subject, setting, composition, lighting, style and mood. Few models
render text reliably: leave text out of the image and overlay it in HTML/CSS
unless the user wants it baked in (then quote the exact words).
- **Edit / restyle / keep a character consistent**: pass the source images in
`reference_image_urls`. Local files go through `upload_file` first (small
files as `content_base64`, big ones with the returned `curl` command).
- **Video**: `create_video` returns a `job_id`; call `get_job(job_id,
wait_seconds=50)` until `status` is `completed` (usually 1-5 minutes). To
animate a still, generate or upload the image and pass it as
`first_frame_url`; add `last_frame_url` to control the ending. Keep prompts
about motion and camera ("slow dolly-in, waves roll left to right").
- **Voice-over**: `generate_speech(model, text, voice)`; `instructions` steers
tone on models that support it.
- **Subtitles / transcript**: `transcribe_audio(model, audio_url,
response_format="srt")` (or `"vtt"`, or `"verbose_json"` for timestamps).
If `generate_image` answers `in_progress` with a `job_id`, the model is slow:
poll `get_job` the same way.
## 5. Save, check, report
- Download at once; links expire in about 24 hours:
`curl -L -o public/hero.png ""`. Use the path and name the code expects,
and ask before overwriting an existing file.
- Look at the preview (image tools return one) before wiring the file in. If it
is off, change the prompt specifically ("camera lower, warmer light") rather
than re-rolling blindly.
- Wire it in with the right dimensions and alt text.
- Tell the user: what was made, which model, what it cost (`cost_usd`), where
it was saved.
## Errors
- Balance too low: the tool says so; `get_account` shows the balance. The user
adds credit in the OceanToken console. Offer a cheaper model as the alternative.
- Parameter rejected: re-read `get_model` and use a listed value.
- Failed video jobs are not charged; the hold is released automatically.
- Reconnect prompt: the user's OceanToken sign-in or key was revoked; they need
to reconnect the plugin.