Unified multimodal model for images, video with synchronized audio, and robotic action. FLUX.3 Video makes 20-second clips; Image does synthesis and editing; Action predicts robotic manipulation. All use flow matching.
ScalingStage 5 of 5
Released Jul 2026. Video/action in early access; image and open-weight to follow.
Updated 23 Jul 2026·Checked 10 Oct·0 updates this week
Milestones
No announced next step
Black Forest Labs founded by Stable Diffusion researchers2024Complete.
FLUX.1 released2024Complete.
Meta partnership announced, $140M contractSep 2025Complete.
FLUX.2 released with three editions25 Nov 2025Complete.
Series B funding at $3.25B valuationDec 2025Complete.
FLUX.3 multimodal model released23 Jul 2026Complete.
Multi-shot coherence and continuityMaintaining character and scene consistency across long videos and multi-shot sequences.
Cross-modal alignmentEnsuring tight synchronization between generated audio and video actions, and aligning action predictions with visual dynamics.
Physics limits
Video generation durationCurrent generation is limited to 20-second clips; extending to longer-form coherent narratives requires solving memory and consistency challenges.
Real-time action prediction for roboticsWhile action prediction is supported, real-time robotic feedback and online learning during deployment remain open research questions.
How it works
3 parts
Unified architecture
Images, video, audio and action
Learning from video and audio builds models of physical behavior. Supports images, video with synced audio, and robotic action in one system.
Flow matching
Efficient generative foundation
Replaces diffusion with learnable paths, reducing computation and inference time across all modalities.
Video control
Keyframes and composition
Keyframe-controlled transitions, image-to-video with character consistency, video continuation, and dialogue control.
FLUX, image generation, video generation, flow matching
German AI company by latent diffusion researchers. FLUX multimodal models for images, video, audio and robotics using flow matching. Partnerships exceed $400M.
Raised $300M Series B at $3.25B in Dec 2025, led by Salesforce Ventures and AMP with a16z, NVIDIA, Temasek.
Meta signed $140M multi-year contract Sep 2025; total partnerships exceed $400M with Adobe, Canva, Snap, Deutsche Telekom.
FLUX.3 released Jul 2026: multimodal model for images, videos up to 20 seconds with audio, and robotic action prediction.