AI product Open source
AuK is an open-source 1.5-billion-parameter foundation model for speech generation and editing, developed by Tencent Hunyuan. It exposes zero-shot and instruction-based text-to-speech, content and lyric rewriting, pitch, speed, volume, emotion, timbre, accent, whisper and nonverbal-sound editing, speech enhancement, speaker separation, music separation, and target-speaker extraction through a unified natural-language instruction interface. Inputs can include text instructions and optional source or reference audio; outputs are generated or edited audio. The repository provides an AuK base model for configurable high-quality generation and AuK-Flash, a distilled variant using four-step inference, along with command-line, Python, Gradio, ComfyUI, and fine-tuning interfaces. The code is released under the MIT License, and model weights are distributed through Hugging Face and ModelScope.
2 uses taken from transcripts — each links to the moment in the video.
A 1.5-billion-parameter speech model that can synthesize and edit speech from plain-language instructions. It supports reference-voice synthesis, rewriting, emotion and pacing changes, noise removal, and speaker separation.
An open-source 1.5-billion-parameter speech model from Tencent for generating and editing speech through natural-language instructions, including text-to-speech, speech editing, enhancement, and speaker separation.
2 in the library.