AI product Open source

AuK

AuK is an open-source 1.5-billion-parameter foundation model for speech generation and editing, developed by Tencent Hunyuan. It exposes zero-shot and instruction-based text-to-speech, content and lyric rewriting, pitch, speed, volume, emotion, timbre, accent, whisper and nonverbal-sound editing, speech enhancement, speaker separation, music separation, and target-speaker extraction through a unified natural-language instruction interface. Inputs can include text instructions and optional source or reference audio; outputs are generated or edited audio. The repository provides an AuK base model for configurable high-quality generation and AuK-Flash, a distilled variant using four-step inference, along with command-line, Python, Gradio, ComfyUI, and fine-tuning interfaces. The code is released under the MIT License, and model weights are distributed through Hugging Face and ModelScope.

View repository Visit site Mentioned in 2 videos ↓

What AuK is used for

2 uses taken from transcripts — each links to the moment in the video.

  • A 1.5-billion-parameter speech model that can synthesize and edit speech from plain-language instructions. It supports reference-voice synthesis, rewriting, emotion and pacing changes, noise removal, and speaker separation.

  • An open-source 1.5-billion-parameter speech model from Tencent for generating and editing speech through natural-language instructions, including text-to-speech, speech editing, enhancement, and speaker separation.

Videos mentioning AuK

2 in the library.