What Is MiniMax Music 3?
MiniMax Music 3 is an open-weight text-to-music model developed by MiniMax. It can generate songs up to five minutes long while maintaining relatively consistent vocals, rhythm, musical themes, and song structure.
The model accepts two main inputs:
Lyrics, including tags such as [Verse], [Chorus], [Bridge], and [Outro]
Music captions describing genre, mood, vocals, instruments, tempo, arrangement, and production style
MiniMax Music 3 produces 32 kHz, 16-bit stereo audio. It is designed for complete song generation rather than short sound effects or isolated music loops.
How MiniMax Music 3 Works
MiniMax Music 3 combines several components to handle both composition and audio quality.
An 8B Global LLM manages the long-range structure of the song, while a 0.6B Local LLM adds frame-level acoustic detail. Flow Matching and Flow-VAE components then convert the model’s hidden representations into the final audio.
This architecture helps MiniMax Music 3 maintain recurring choruses, vocal identity, instrumentation, and emotional progression across longer tracks.
Using MiniMax Music 3 in ComfyUI
ComfyUI provides an official MiniMax Music 3 workflow. After updating ComfyUI, open the template browser and search for MiniMax Music 3.
The workflow requires three model components:
ComfyUI/models/
├── diffusion_models/
│ └── minimax_music3_dit_fp16.safetensors
├── text_encoders/
│ └── minimax_music3_text_encoder_pruned_int8_convrot.safetensors
└── vae/
└── minimax_music3_dav.safetensorsA low-VRAM INT8 diffusion model is also available. The required files can be downloaded from the ComfyUI MiniMax Music 3 repository.
Once the models are loaded, enter a music caption, add structured lyrics, choose the maximum duration and seed, and queue the workflow.
For better results, describe more than the genre. Include the vocal style, primary instruments, emotional progression, and how the arrangement should change between sections.
Current Limitations
MiniMax Music 3 currently requires CUDA in the official implementation and only supports non-streaming generation. The released workflow focuses on text and lyrics to music; it does not yet provide reference-audio conditioning, cover generation, or audio-to-audio editing.
Musical instructions are also generative rather than exact. The output may not always follow the requested BPM, key, instruments, lyrics, or section structure perfectly.
MiniMax Music 3 is best described as an open-weight model, since its custom community license includes attribution, commercial, and acceptable-use requirements.
Final Thoughts
MiniMax Music 3 brings one of the most capable open-weight music-generation models into ComfyUI. Its support for structured lyrics, expressive vocals, detailed prompts, and songs up to five minutes makes it useful for music demos, background tracks, AI videos, and experimental creative workflows.
For ComfyUI users, MiniMax Music 3 also makes music generation easier to combine with image, video, and audio-processing nodes—all within one local workflow.






