Back to Muse
Meta /

Muse Video
Muse Video

Research PreviewVideo GenvideoaudioUpdated July 1, 2026
Some details or benchmark scores on this page are self-reported by developers and unconfirmed.

Muse Video

There is currently no API availability or technical specifics disclosed.

Model Overview

Muse Video is a state-of-the-art video generation model developed by Meta Superintelligence Labs. Announced in July 2026 alongside the Muse Image model, it is designed to deliver high visual fidelity with native audio support. The model leverages the same robust pretraining base as the rest of the Muse media generation family, showcasing complex physical comprehension and seamlessly integrated audio rendering.

Capabilities

As an advanced multimodal model, Muse Video is designed for:

  • Native Audio Integration: Audio generation is natively baked into the video rendering process, rather than being stitched on as a post-processing step.
  • Physical Plausibility: The model exhibits high physical plausibility for complex, continuous movements (e.g., juggling), maintaining consistency across frames.
  • Multimodal Generation: Supports generating rich video and audio outputs simultaneously.

Example Use Cases

  • Content Creation: Assisting creators in producing high-fidelity video clips with perfectly synchronized audio.
  • Entertainment & Media: Generating conceptual videos or storyboards with realistic physics and soundscapes.
  • Meta AI Ecosystem: Future integration into Meta's platforms to empower users with advanced multimedia tools.

Performance & Benchmarks

However, early demonstrations show significant improvements in visual consistency and native audio synchronization compared to previous generations of video models.

Intended Use & Limitations

  • Intended Use: Designed as an API-only service (once released) for creators, researchers, and developers seeking high-fidelity video generation. As with most generative video models, it may occasionally struggle with prolonged narrative consistency or artifacting in highly complex scenes.

About Meta

Meta is a global technology company focused on connecting people through social media platforms, virtual reality, and advanced artificial intelligence. The Meta Superintelligence Labs division drives frontier research in multimodal generative models, open-source AI, and immersive technologies.

Key Features

Audio generation is natively baked into the video rendering process

Feature 01

High physical plausibility for complex movements like juggling

Feature 02

You might also want to compare

Verified Sources

Model Specs

research-preview

Parameters

Undisclosed

Context Window

undisclosed

License

Proprietary

Deployment

api-only

Resources & Links

Lineage

Model Family

Part of the Muse family

Only release in this line currently tracked.

Curator Notes

Currently in preview mode with no API availability or technical specifics disclosed in the transcript.

Compare Specs

Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.

Compare Model