StorySoftware
Black Forest Labs releases FLUX.3 multimodal model with video, audio and robotics
Black Forest Labs unveiled FLUX.3, a multimodal frontier model that jointly learns image, video, audio and action prediction in a unified architecture. FLUX.3 Video generates 20-second clips with native synchronized multilingual dialogue; FLUX.3 Action enables robotic learning with partner mimic robotics.
- Single unified model spans image generation, video synthesis, audio, and action prediction
- FLUX.3 Video produces 20-second clips with native synchronized audio
- Text-to-video, image-to-video, video continuation and keyframe-controlled transitions supported
- FLUX.3 Action enables robotic manipulation through video-action learning with mimic robotics