StorySoftware

Black Forest Labs releases FLUX.3 multimodal model with video, audio and robotics

Black Forest LabsFLUX.3

Black Forest Labs unveiled FLUX.3, a multimodal frontier model that jointly learns image, video, audio and action prediction in a unified architecture. FLUX.3 Video generates 20-second clips with native synchronized multilingual dialogue; FLUX.3 Action enables robotic learning with partner mimic robotics.

  • Single unified model spans image generation, video synthesis, audio, and action prediction
  • FLUX.3 Video produces 20-second clips with native synchronized audio
  • Text-to-video, image-to-video, video continuation and keyframe-controlled transitions supported
  • FLUX.3 Action enables robotic manipulation through video-action learning with mimic robotics
Read the original · Press

More on FLUX.3

Primer