AI & Creative Tools

Black Forest Labs' FLUX 3 Puts Image, Video, Audio, and Robot Actions in One Model

FLUX 3 Video entered gated early access on July 23 as the lab's first model trained jointly across modalities — a single network that generates 20-second video with native audio, still images, and even action predictions for robots.

Black Forest Labs, the German lab behind the FLUX image-generation family, opened gated early access to FLUX 3 on July 23 — its first attempt at a genuinely multimodal foundation model rather than a specialized one. Where FLUX 1 and FLUX 2 were image models with video and editing variants bolted on, FLUX 3 is trained jointly on images, video, and audio from the start, with a fourth capability — robot action-prediction — folded into the same architecture.

What’s actually new here

The headline feature is FLUX 3 Video: it generates clips up to 20 seconds long with native audio, synthesized in the same pass rather than added afterward by a separate audio model bolted onto silent video. The system supports generation from text, from a still image, or from existing footage, with continuation and keyframe-transition modes for stitching scenes together, multilingual dialogue, and clip chaining for longer sequences. Black Forest Labs’ own preference testing claims a 77% win rate against Runway Gen-4.5 and 93% against Luma Ray 3.2 — numbers worth treating as a lab’s self-reported comparison rather than an independent benchmark, but a signal of where they think they stand.

The robot action-prediction piece is the part that reframes what FLUX 3 is for. Training one model across pixels, sound, and physical-action sequences is a bet that generative video and embodied-AI world models are converging problems — that predicting “what does this scene look like next” and “what should this robot arm do next” benefit from shared representations rather than separate stacks. That’s a research direction more than a shipping product today, but it’s a notable architectural choice for a lab whose business has been creative image tools.

What’s not out yet

FLUX 3 Video is the only piece live, and it’s gated: access is via API and private weights, available to selected partners who apply through Black Forest Labs’ site. FLUX 3 Image, an open-weight FLUX 3 Dev variant, and a separate FLUX-mimic robotics model are described as following in later rollout phases, with no firm dates attached. That staged rollout — video first, open weights later — is a departure from FLUX 2’s pattern, where an open Dev variant shipped closer to the proprietary release.

For creative technologists, the practical read right now is: this is a lab to watch closely over the next couple of months rather than a tool to build with today. The interesting test will be whether FLUX 3 Dev, once it lands, keeps the audio-video joint generation or trims it down the way open releases sometimes do to fit consumer hardware.