AI product Open source
MiniMax Music 3 is a music-generation model from MiniMax AI that creates complete songs of up to five minutes from lyrics and a music description. Lyrics can use section tags such as verse, chorus, bridge, instrumental, solo, and outro; the music description specifies style, emotional progression, vocals, instrumentation, arrangement, and production. A structured-caption format can organize these instructions into global metadata, vocal details, and arrangement sections.
The model uses a hierarchical autoregressive hybrid architecture: an 8B Global LLM predicts the first residual-vector-quantization codebook frame by frame to model long-range musical structure, while a 0.6B Local LLM predicts the remaining acoustic codebooks within each frame. Its continuous hidden-state synthesis module combines the two models' final hidden states with Flow Matching and Flow-VAE to produce 32 kHz, 16-bit stereo WAV audio, preserving information for vocal articulation, instrumental texture, and temporal continuity.
1 use taken from transcripts — each links to the moment in the video.
Generates full five-minute songs from lyrics and a music description rather than short clips. An 8B global model tracks song structure while a 0.6B local model supplies frame-level acoustics.
1 in the library.