# ๐ŸŽง dsh-audiogen **AI audio generation for DeepSeek Harness (DSH)** โ€” turn your DSH web GUI into an audio studio: text-to-speech, music, sound effects and voice design, from the sidebar panel or straight from the Agent. [English](README.md) | [็ฎ€ไฝ“ไธญๆ–‡](README.zh-CN.md) ![npm version](https://img.shields.io/npm/v/dsh-audiogen?style=flat-square&color=8B5CF6) ![license](https://img.shields.io/npm/l/dsh-audiogen?style=flat-square) ![DSH plugin](https://img.shields.io/badge/DSH-plugin-brightgreen?style=flat-square) ![Node](https://img.shields.io/badge/node-%3E%3D20-blue?style=flat-square) ![Main panel](docs/images/hero.png) ## โœจ Features - **Four generation modes**: text-to-speech, music, sound effects, and voice design - **Multi-vendor channels in one place**: OpenAI-compatible TTS, MiniMax, ElevenLabs, Stability AI, or any custom OpenAI-compatible / generic POST endpoint - **Per-channel model & voice catalogs** with one-click discovery, display aliases, capability categories, and per-model advanced fields (duration, seed, steps, cfg_scale, loop, prompt influence, and ElevenLabs SFX output format split into format/sample rate/bitrate, combined host-side into the single `output_format`, โ€ฆ) - **Model comparison**: run the same prompt across 2โ€“4 models at once with per-model parameter overrides โ€” results are grouped side by side - **Prompt enhancement**: rewrite a rough idea into a ready-to-generate description with an LLM (pick any model from *Settings โ†’ Models*; falls back to the agent default model) - **History with one-click restore**: prompt, config, model set *and the original audio* come back into the panel โ€” no regeneration, no extra cost - **Resource library**: auto-save generated audio (or opt in per run), organized by type โ€” voices / music / SFX / TTS โ€” with search, tags, rename, category moves, and full provenance (channel, model, voice id, prompt, params snapshot). Reuse a voice or music bed instead of regenerating - **Agent tools**: `generate_audio` and `search_audio_library`, `manage_audio_voices` (vendor voice browsing/deletion + prompt-based voice recommendation + **role voice casting**), plus bundled session skills โ€” the Agent can generate and find audio on demand - **Role voice casting**: assign a primary voice (+ backups) to each character of a novel/game โ€” `manage_audio_voices` `action=cast` takes character profiles (JSON array/object or a text description structured first) and applies deterministic hard filters (gender/age/use_case strict; accent is a preference relaxed only when the strict pool is empty) per character; the Agent picks voices globally (no primary reuse across lead/major roles) and `action=save_cast` validates membership, auto-fills backups, flags reuse and persists the plan to `~/.dsh/dsh-audiogen/cast-selections.json`; then TTS with the chosen `voice_id` (or design a custom voice first via `generate_audio(mode=voice_design)`) - **Panel voice management**: a ยซ้Ÿณ่‰ฒยป entry in the studio's left mode row (next to TTS/music/SFX/voice-design) โ€” browse/filter vendor voices (language/keyword/source + official ElevenLabs shared-voice filters), ask the agent default model to recommend voices for a natural-language requirement (e.g. ยซๆธ…ไบฎ็”œ็พŽ็š„ๅฐ‘ๅฅณ้Ÿณยป), preview, delete account-owned voices (confirmed) and backfill the chosen `voice_id` into the TTS form; **every AI recommendation is recorded automatically** (last 50, shared by panel and Agent) so you can revisit requirements/channels/reasons and reuse a voice later - **Keys stay local**: API keys live in the local DSH settings document and generation is proxied by the local host; the browser and the Agent never touch plaintext credentials ## ๐Ÿ“ธ Screenshots | Generation panel | Resource library | | --- | --- | | ![Generation](docs/images/hero.png) | ![Library](docs/images/library.png) | | Library โ€” full provenance drawer | Channels settings | | --- | --- | | ![Library detail](docs/images/library-detail.png) | ![Settings](docs/images/settings-channels.png) | | Channel editor (model catalog & auto capabilities) | LLM models (Settings โ†’ Models) | | --- | --- | | ![Channel editor](docs/images/settings-stability.png) | ![Models page](docs/images/models-page.png) | ## ๐Ÿ“ฆ Installation The plugin is published on npm. DSH host (Node โ‰ฅ 20) required. ```bash dsh plugin --profile web add dsh-audiogen ``` Local development install: ```bash dsh plugin --profile web add /path/to/dsh-audiogen ``` Restart `dsh web` after install โ€” the sidebar will show the **AI Audio** entry. ## ๐Ÿš€ Quick start 1. Open **Settings โ†’ Plugins โ†’ AI Audio** 2. Add a channel: pick a preset provider (+ Add provider) or a custom endpoint (+ Add custom provider) 3. Fill in the API URL, API key, and the model/voice catalog (use *Fetch available models* to import them) 4. Save, then open the **AI Audio** sidebar panel: - choose a mode (Speech / Music / Sound effects / Voice design) - type your text or prompt (optional: โœจ Enhance prompt) - pick a model โ€” or tick **Model comparison** for 2โ€“4 models at once - press **Start generation** and play the results, download them, or add them to the resource library ## ๐ŸŽ› Modes supported by each vendor | Mode | MiniMax | ElevenLabs | Stability AI | OpenAI-compatible / custom | | --- | --- | --- | --- | --- | | TTS | โœ… (8 voices) | โœ… (voices + streams) | โ€” | โœ… | | Music | โœ… (`music-3.0` / `music-2.6` / `music-cover`) | โœ… (`music_v2`) | โœ… (`stable-audio-*`) | โœ… (generic POST) | | Sound effects | โ€” | โœ… (`eleven_text_to_sound_v2`, loop / prompt influence / output format as codec+sample rate+bitrate โ†’ `output_format`) | โœ… (`stable-audio-*` โ€” same text-to-audio protocol; auto-detected in both Music and SFX) | โœ… (generic POST) | | Voice design | โœ… (`/v1/voice_design`) | โœ… (`/v1/text-to-voice/design`) | โ€” | โ€” | ## ๐Ÿค– Agent usage | Tool | Purpose | | --- | --- | | `generate_audio` | Submit a TTS / music / SFX / voice-design task; waits for completion and returns same-origin audio URLs. Optional `enhance_prompt`, `save_to_library`, per-vendor params. | | `manage_audio_voices` | Browse/filter the vendor voice libraries (MiniMax, ElevenLabs) with language/keyword/source filters; recommend top-k voices for a natural-language requirement (`action=recommend`, uses the agent default model, ids validated against the pool); **role casting** (`action=cast` prepares per-character filtered candidate pools from character profiles; `action=save_cast` validates + persists the plan); delete account-owned voices (official/shared/system voices are read-only and refused). Then use the returned `voice_id` with `generate_audio` (mode=tts). | | `search_audio_library` | Search the local resource library (type / category / keyword) and reuse an existing voice, music bed or effect. | Typical session commands (skills bundled with the plugin): ```text /audio:tts Read this sentence with a warm voice /audio:music Generate a 30-second lo-fi background track /audio:sfx Create a sci-fi UI cue /audio:design Craft a warm retro synth voice ``` ## ๐Ÿ” Security & data notes - API keys are stored in the local DSH settings document; requests are proxied by the local host (`/api/dsh-audiogen/*`, loopback-only routes) - Generation consumes your upstream provider quota; audio content is produced by the upstream model - History & library persist under `~/.dsh/dsh-audiogen/` - Prompt enhancement calls the LLM model you choose (default: agent default model) โ€” no extra API key ## ๐Ÿ›  Development ```bash pnpm install pnpm run typecheck pnpm run build # outputs lib/ (host + client bundles) ``` ## ๐Ÿ“„ License [Apache-2.0](LICENSE)