💻 Try our model on the HuggingFace Space · ModelScope Space!
## News - **[2026/09/16]** 💻 Reduced encoder memory by **~7.5 GiB**, making local inference more accessible on **consumer GPUs**. See [PR #19](https://github.com/Tencent-Hunyuan/AuK/pull/19). - **[2026/09/13]** 🍎🖥️ **AuK** now officially supports **MLX inference on Apple Silicon** (available on the [feat/mlx-apple-silicon](https://github.com/Tencent-Hunyuan/AuK/tree/feat/mlx-apple-silicon) branch) and **CPU offload for CUDA inference**. - **[2026/09/13]** 🏆 **AuK** serves as the end-to-end baseline for the **Single Model Track** of the [ICASSP 2027 Audio Editing Challenge](https://audio-editing-challenge.github.io/). - **[2026/09/09]** 🙌 Thanks to **SGLang-Omni** for Day 0 support for **AuK** and **AuK-Flash**! Check out the [SGLang-Omni cookbook](https://sgl-project.github.io/sglang-omni/cookbook/auk.html) to get started. - **[2026/09/09]** 🎉 We open-source **AuK**. Code and model weights are publicly available. Try it on the [🤗 Demo Space](https://huggingface.co/spaces/tencent/AuK) or the [🤖 ModelScope Space](https://modelscope.cn/studios/Tencent-Hunyuan/AuK)! ### Demo| Category | Task | Description | Cookbook |
|---|---|---|---|
| Speech Generation | Zero-shot TTS | Speak the target text in the voice of the reference audio. | Zero-shot TTS |
| Instruct TTS | Generate speech from a voice description alone — no reference audio. | Instruct TTS | |
| Content Editing | Speech Content Editing | Rewrite what is said — replace, insert, or remove text. | Speech Content Editing |
| Lyric Editing | Rewrite lyrics in a singing recording while preserving the melody and voice. | Lyric Editing | |
| Acoustic Editing | Pitch Editing | Raise or lower the pitch by semitones. | Pitch Editing |
| Speed Editing | Adjust the speaking rate; output length scales with the speed factor. | Speed Editing | |
| Volume Editing | Raise or lower the volume by decibels. | Volume Editing | |
| Paralinguistic Editing | Emotion | Change the emotion while preserving content and voice. | Emotion |
| Timbre | Change the timbre to a description while keeping the content unchanged. | Timbre | |
| De-accent | Remove a regional accent while preserving the speaker's voice and content. | De-accent | |
| Nonverbal Editing | Remove or add nonverbal sounds such as breaths, laughs, or coughs. | Nonverbal Editing | |
| Whisper Conversion | Convert between normal speech and whisper while preserving speaker and content. | Whisper Conversion | |
| Enhancement & Separation | Speech Enhancement | Denoise, dereverberate, or restore natural, clear speech. | Speech Enhancement |
| Speech Separation | Keep one speaker by talking order and remove the others. | Speech Separation | |
| Music Separation | Extract the singing voice from a mix, or keep all human voices. | Music Separation | |
| Target Speaker Extraction | Keep the target speaker identified by what they say. | Target Speaker Extraction |