MusicGen
MusicGen
Model Overview
MusicGen is Meta's single-stage auto-regressive transformer model designed for high-quality music generation. Released on June 8, 2023, under a Creative Commons Attribution-NonCommercial (CC-BY-NC 4.0) license, MusicGen operates on compressed discrete audio tokens generated by Meta’s EnCodec audio tokenizer. It is available in multiple parameter sizes, most notably a 3.3 billion parameter version.
Capabilities
- Text-to-Music: Generates high-quality stereo music directly from text prompts describing genre, mood, instruments, or style.
- Melody Conditioning: Supports conditioning generation on reference melodies (audio files), guiding the melodic structure while the text prompt sets the stylistic context.
- Single-Stage Architecture: Utilizes an efficient token interleaving pattern to predict audio codebooks, avoiding the need for cascaded multiple stages or hierarchical upsampling.
- Extensive Training: Trained on over 20,000 hours of licensed music to ensure diverse and high-fidelity outputs.
Example Use Cases
- Background Tracks: Quickly generating royalty-free background music for videos, podcasts, and games.
- Musical Ideation: Assisting composers and musicians in exploring new melodies, chord progressions, and stylistic fusions.
- Audio Prototyping: Providing sound designers with rapid, controllable music generation for various media projects.
Performance & Benchmarks
MusicGen, particularly the Large (3.3B) variant, serves as a foundational benchmark in the field of open-source music generation.
- FAD (Fréchet Audio Distance): Achieved a verified score of 3.4, indicating high statistical similarity between the generated audio and real music.
- Human Preference Studies: Consistently ranks highly in pairwise audio comparisons for perceived quality and prompt alignment.
Intended Use & Limitations
- Intended Use: Self-hostable, open-weights model intended for non-commercial research, creative exploration, and audio tool development.
- Limitations: The model has a limited context window of 30 seconds for direct generation. Generated audio may occasionally drift from the text prompt or exhibit artifacts in highly complex or unstructured genres.
About Meta
Meta is a global technology company focused on connecting people through social media platforms, virtual reality, and advanced artificial intelligence. Through initiatives like AudioCraft, Meta is a leading contributor to the open-source AI community, advancing research in audio and multimodal generation.
Key Features
Generates stereo music from text prompts
Supports conditioning generation on reference melodies
Trained on 20,000 hours of licensed music
Open weight files under CC-BY-NC 4.0
You might also want to compare
Verified Sources
Tags
Model Specs
Parameters
3.3B
Context Window
30s
License
Other/Custom
Deployment
Resources & Links
Lineage
Model Family
Part of the AudioCraft family
Only release in this line currently tracked.
Curator Notes
Weights released under Creative Commons Attribution-NonCommercial 4.0 International license.
Compare Specs
Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.
Compare Model