Black Forest Labs launches Flux 3, generating video and audio simultaneously
Black Forest Labs has released Flux 3, a multimodal AI model capable of generating video clips up to 20 seconds long with natively synchronized audio in a single pass. Unlike previous AI video tools that added sound as a separate post-production step, Flux 3 uses a unified transformer architecture — dubbed Self-Flow — that learns image, video, and audio together. This approach eliminates common sync issues such as mismatched lip movements or delayed sound effects. In internal benchmarks on 720p, 10-second clips, Flux 3 outperformed Runway Gen-4.5 in 77% of comparisons and Luma Ray 3.2 in 93%, though it only matched Seedance 2.0 and Gemini Omni Flash at around 52%. The model also supports text-to-video, image-to-video, multilingual dialogue, and multi-shot sequence composition.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in