Black Forest Labs, the German lab behind the FLUX image-generation family, opened gated early access to FLUX 3 on July 23 — its first attempt at a genuinely multimodal foundation model rather than a specialized one. Where FLUX 1 and FLUX 2 were image models with video and editing variants bolted on, FLUX 3 is trained jointly on images, video, and audio from the start, with a fourth capability — robot action-prediction — folded into the same architecture.
What’s actually new here
The headline feature is FLUX 3 Video: it generates clips up to 20 seconds long with native audio, synthesized in the same pass rather than added afterward by a separate audio model bolted onto silent video. The system supports generation from text, from a still image, or from existing footage, with continuation and keyframe-transition modes for stitching scenes together, multilingual dialogue, and clip chaining for longer sequences. Black Forest Labs’ own preference testing claims a 77% win rate against Runway Gen-4.5 and 93% against Luma Ray 3.2 — numbers worth treating as a lab’s self-reported comparison rather than an independent benchmark, but a signal of where they think they stand.
The robot action-prediction piece is the part that reframes what FLUX 3 is for. Training one model across pixels, sound, and physical-action sequences is a bet that generative video and embodied-AI world models are converging problems — that predicting “what does this scene look like next” and “what should this robot arm do next” benefit from shared representations rather than separate stacks. That’s a research direction more than a shipping product today, but it’s a notable architectural choice for a lab whose business has been creative image tools.
What’s not out yet
FLUX 3 Video is the only piece live, and it’s gated: access is via API and private weights, available to selected partners who apply through Black Forest Labs’ site. FLUX 3 Image, an open-weight FLUX 3 Dev variant, and a separate FLUX-mimic robotics model are described as following in later rollout phases, with no firm dates attached. That staged rollout — video first, open weights later — is a departure from FLUX 2’s pattern, where an open Dev variant shipped closer to the proprietary release.
For creative technologists, the practical read right now is: this is a lab to watch closely over the next couple of months rather than a tool to build with today. The interesting test will be whether FLUX 3 Dev, once it lands, keeps the audio-video joint generation or trims it down the way open releases sometimes do to fit consumer hardware.
Related Reading
- FLUX 3 Launches: Black Forest Labs Enters Video, Audio, and Physical AI in One Model — Tech Times
- Black Forest Labs FLUX 3 Multimodal Video Model — DataNorth
- Black Forest Labs Launches FLUX 3, but Its 20-Second Video Claim Still Needs a Real-World Test — Remio
- FLUX 3 — Black Forest Labs (official)
- FLUX 3 Released: Black Forest Labs Turns Its Image Model Into a Multimodal Video, Audio and Robotics Engine — Coursiv