Documentation Navigation
Models & pricing / Models / MiniMax-Music3
MiniMax-Music3 Overview
This model's data has not been fully verified and benchmarks may be missing.
MiniMax-Music3 is a audio-speech AI model created by MiniMaxAI, featuring undisclosed parameters and a context window of unknown.
SGLang-Omni is a multi-stage serving runtime for omni, speech, and TTS models, designed for multi-stage decoding and generation. It owns the pipeline topology, stage lifecycle, inter-stage transport, model-family integration layer, and OpenAI-compatible serving surface.
Model Lineage & Specification
minimax-music3 is MiniMaxAI's primary release in the current family.
minimax-music3Comparable models
| Feature | MiniMax-Music3 | Audio8-TTS-Preview-0.6b | VibeVoice-ASR-BitNet | Stable Audio 3 |
|---|---|---|---|---|
| Description | SGLang-Omni is a multi-stage serving runtime for omni, speech, and TTS models, d... | Audio8-TTS-Preview-0.6b is a 601M-parameter text to speech model developed by Au... | VibeVoice-ASR-BitNet is a 323M-parameter automatic speech recognition model deve... | Stable Audio 3 is a model developed by Stability AI, released on 2026-07-15. It ... |
| API Identifier | minimax-music3 | audio8-tts-preview-06b | vibevoice-asr-bitnet | stability-ai-stable-audio-3 |
| Parameters | — | 601M | 323M | — |
| Context Window | unknown | unknown | unknown | 128K tokens |
| License / Type | open weights | open weights | open weights | research preview |
MiniMax-Music3
SGLang-Omni is a multi-stage serving runtime for omni, speech, and TTS models, d...
minimax-music3Audio8-TTS-Preview-0.6b
Audio8-TTS-Preview-0.6b is a 601M-parameter text to speech model developed by Au...
audio8-tts-preview-06bVibeVoice-ASR-BitNet
VibeVoice-ASR-BitNet is a 323M-parameter automatic speech recognition model deve...
vibevoice-asr-bitnetStable Audio 3
Stable Audio 3 is a model developed by Stability AI, released on 2026-07-15. It ...
stability-ai-stable-audio-3SGLang-Omni is a multi-stage serving runtime for omni, speech, and TTS models, designed for multi-stage decoding and generation. It owns the pipeline topology, stage lifecycle, inter-stage transport, model-family integration layer, and OpenAI-compatible serving surface.