AI product Open source

MAGI-2 Preview

MAGI-2 Preview is the inference implementation for SandAI’s unified audio-video generation model. The 114-billion-parameter architecture, built on MagiMoE, activates about 6 billion parameters per token and generates 10-second clips from text prompts (T2V) or a prompt plus a still image (I2V), with sound generated alongside the video and muxed into the output file.

View repository Visit site Mentioned in 2 videos ↓

Overview

Generation runs in two stages: `magi2_preview` denoises the clip at low resolution, after which `magi2_refiner` increases it to 1080p. The current base release is not step-distilled, so denoising steps account for most of the generation time; a distilled release is described as forthcoming. The repository contains inference code, while the model weights are downloaded separately from the `sand-ai/MAGI-2-preview` Hugging Face repository and occupy hundreds of gigabytes in total.

An optional prompt-enhancement stage sends the input to an OpenAI-compatible instruction-following LLM, which rewrites it as a structured JSON caption for the 10-second clip, renders that result as Markdown, and passes it to the generation pipeline. It can be disabled to use the raw prompt. The documented runtime requires eight NVIDIA Hopper GPUs, Python 3.12, a recent CUDA toolkit, and FFmpeg for audio-video muxing; Docker images and a source installation are provided.

What MAGI-2 Preview is used for

2 uses taken from transcripts — each links to the moment in the video.

  • An open-source 114-billion-parameter mixture-of-experts model that generates video with sound from text or text plus an image. It produces a 10-second clip through a low-resolution generation stage followed by a 1080p refinement stage.

  • Generates 10-second videos with synchronized sound from a prompt or prompt plus still image, using a low-resolution denoising pass followed by a 1080p refiner.

Videos mentioning MAGI-2 Preview

2 in the library.