Back to Home

Black Forest Labs Drops FLUX 3: AI Video, Audio & Robotics

Black Forest Labs is no longer just an image-generation company. This week, the German AI startup unveiled FLUX 3, a multimodal AI model that does what few have attempted: generate images, video with synchronized audio, and power robotic systems — all within a single architecture.

Video Generation Finally Has Native Audio

FLUX 3 can generate video clips up to 20 seconds long with audio baked in — no separate sound generation pipeline needed. The model supports text-to-video, image-to-video animation, video editing via reference clips, keyframe-controlled transitions, and multilingual dialogue. It can even chain shorter clips into longer sequences while keeping characters consistent.

Black Forest Labs ran early internal comparisons using 10-second, 720p text-to-video clips. The results are eye-catching: FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93%. It also outperformed Grok Imagine Video and Kling v3 Pro. That said, these are self-reported benchmarks — independent validation is still pending.

From Content Creation to Physical AI

The bigger story here isn't just better videos. Black Forest Labs is betting that understanding how objects move on screen translates directly into teaching robots how to act in the physical world.

They partnered with robotics company mimic to build FLUX-mimic, a video-action model built on the FLUX 3 backbone that predicts robot movements. It is already being tested in manufacturing environments — including Audi production lines — for tasks like assembling components, handling flexible materials, and placing parts into fixtures.

Audi Production Lab's Christoph Schneider noted that these robots have solved complex soft-body manipulation work that would have been impossible with conventional robotics approaches.

Why This Architecture Matters

Black Forest Labs argues that video inherently encodes how the physical world works — movement, timing, cause and effect. According to the company, video prediction accounts for more than 95% of FLUX 3 training compute, while audio is a much smaller slice. The thesis: once a model understands visual change over time, adding sound and physical action becomes far easier.

For enterprises, this could mean consolidating what currently requires multiple specialized AI systems — marketing content, product visualization, simulation, and robotics — into a single backbone. But the approach remains unproven at production scale, and businesses will want independent benchmarks and safety validation before deploying it in environments where robots interact with people.

Open Weights Are Coming

FLUX 3 enters early access now with video and action capabilities rolling out first. Image generation features are expected to follow. An open-weight version called FLUX 3 Dev is also planned for later release, continuing the strategy that helped earlier FLUX models gain wide adoption among developers and researchers.

Open access has been central to Black Forest Labs' growth, and FLUX 3 Dev will bring a multimodal backbone to the open-source community — supporting image, video, audio, and action prediction in one package.

FLUX 3 is early, and enterprises will likely wait for pricing, independent benchmarks, and broader availability before committing. But by pushing beyond image generation into video, audio, and robotics, Black Forest Labs is positioning itself as a contender in the race to build foundation models that span both digital content and the physical world.

  • FLUX 3 generates video up to 20 seconds with native synchronized audio
  • Internal benchmarks show 77% preference over Runway Gen-4.5 and 93% over Luma Ray 3.2
  • FLUX-mimic, a robotics variant, is already being tested on Audi production lines
  • Open-weight FLUX 3 Dev version planned for later release

Comments

No comments yet. Be the first to share your thoughts!