Back to Newsroom

Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action Prediction

By MarkTechPost / Modelverse Editorial·July 26, 2026·3 min read
Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action Prediction

NewsHub NewsHub Premium Content Read our exclusive articles FacebookInstagramX Home Open Source/Weights AI Agents Tutorials Voice AI Robotics Newsletter → Partner with Us NewsHub Search Home Open Source/Weights AI Agents Tutorials Voice AI Robotics Newsletter → Partner with Us Home Technology AI Shorts Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image,... TechnologyAI ShortsArtificial IntelligenceApplicationsComputer VisionEditors PickLanguage ModelLarge Language ModelNew ReleasesPhysical AIStaffTech News Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action Prediction By Michal Sutter - July 26, 2026 Black Forest Labs (BFL) has released FLUX 3, a multimodal foundation model that learns from images, videos and audio inside a single architecture. It is also the first FLUX model to ship video, audio and action prediction from one set of weights.

The Black Forest Labs (BFL) research team argues that no single modality gives a complete description of the world. Images capture spatial structure at one instant. Video restores time and exposes physical dynamics. Audio reveals causal relationships between mechanical events and sound. Each is treated as a lossy projection of the same underlying reality.

Training on all of them at once means the modalities constrain each other. The sound has to match the impact. The motion has to obey the mass. The research team calls FLUX 3 its first model built entirely on that principle.

FLUX 3 builds on Self-Flow, BFL's method for aligning multimodal generation and understanding in one architecture. Self-Flow combines the flow matching objective with a self-supervised feature reconstruction objective. The reference implementation on GitHub is Apache-2.0 and uses SiT-XL/2 with per-token timestep conditioning. It trains with a 25% per-token mask ratio and self-distillation from an EMA teacher at layer 20 to a student at layer 8.

That released checkpoint is an ImageNet 256×256 research model, not FLUX 3. BFL states that it 'significantly scaled up compute and data resources' on the same approach to train FLUX 3 across video, images and audio simultaneously. Self-Flow itself was introduced in March 2026, so it is not new to this launch. What is new is the scale.

FLUX 3 Video generates clips up to 20 seconds long in a single generation, with native audio. The supported modes cover text-to-video, image-to-video, video-to-video from a reference clip, keyframe-to-video for controlled transitions, and generative video-audio continuation from input video and audio.

BFL also lists multilingual dialogue, agentic chaining of clips into multi-shot sequences, and strong typography generation with animated designs. The BFL team reports particular strength in human facial expressions and in associating sounds with physical events.

BFL team published preliminary human preference results. The setup was 10-second text-to-video clips at 720p with audio. FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77%. Against Grok Imagine Video the figure is up to 69%, then Kling v3 Pro at 60%, Happy Horse v1 at 59% and Happy Horse 1.1 at 57%. Against Seedance 2.0 and Gemini Omni Flash the result is 52%, close to a coin flip.

Official Announcement

Read the full update directly from the official source at MarkTechPost News.

Stay tuned to Modelverse for real-time model analysis and benchmark coverage.

ai-newsbreakingmarktechpost

Related content

Are brain waves the next unlock for physical AI?

Forget YouTube videos—frontier physical AI models need multiple camera angles, dense annotation, and soon, brain wave readings.

Read article

Ben Bernanke

Official announcement from Anthropic: Ben Bernanke

Read article

Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security

Open source software is a critical pillar of the global economy. It underpins cloud computing, financial services, manufacturing, telecommunications, government and internet servic...

Read article
© 2026 Modelverse®. All rights reserved.Modelverse Newsroom