AI product Open source
Kandinsky 6.0 Video is a family of diffusion models from Kandinsky Lab for synchronized text-to-audio-video and image-to-audio-video generation. Its Lite and Pro models produce five-second video clips with synchronized 44 kHz audio, including lip synchronization. A separate super-resolution model can raise the output to Full HD. The project supports NVIDIA GPUs and provides integrations for Diffusers, ComfyUI, vLLM-Omni, SGLang, and FastVideo, with command-line and API-based generation workflows.
1 use taken from transcripts — each links to the moment in the video.
Generates picture and soundtrack together: given a text prompt or starting image it produces a 5-second video clip with synchronized audio including lip sync. Includes a light model, a larger pro model, and a separate upscaler for full HD output; runs on an Nvidia GPU or via Comfy UI and diffusers.
1 in the library.