FLUX 3 Launches Multimodal Video Generation With Native Audio
Black Forest Labs presents FLUX 3 as its first video model, trained jointly on images, video and audio to generate up to 20-second clips with dialogue, sound effects and ambience created alongside the frames. The source argues that its Self-Flow training approach improves semantic understanding without a separate vision model, underpinning claimed gains in motion, faces, hands and in-scene text. It positions video continuation and low-cost Draft mode as the most practical workflow features, while noting that the preview release lacks editing and reference support.

A video model built around a shared audio-visual representation
FLUX 3 is Black Forest Labs’ first video model, rather than another entry in its earlier text-to-image line. It generates clips of up to 20 seconds with audio produced alongside the frames: dialogue, sound effects, and ambience are generated with the video rather than added in post-production.
Black Forest Labs trained FLUX 3 jointly on images, video, and audio in one architecture. The more specific technical distinction is its training method, published in an arXiv paper as “Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis.”
Conventional denoising training adds noise to an input and asks a model to reconstruct it. The presenter’s argument is that uniformly cleaning up noise does not necessarily require understanding what an image, frame, or scene depicts; many systems address that by coupling the generator to a second model that already understands images.
Black Forest Labs instead applies different amounts of noise to different regions of the same input. The model must infer heavily obscured regions from the portions it can still see, which Black Forest Labs says teaches semantic understanding without a separate understanding model.
The on-screen research comparison makes the intended outcome concrete. Given a prompt for a neon sign reading “Flux is multimodal,” the vanilla model produces “FLUX is mutilmolal,” while the newer model renders “FLUX is multimodal.” Black Forest Labs’ own comparisons also report gains in faces, hands, motion, and text rendering.
The useful controls are coherence, typography, and range
FLUX 3 supports text-to-video and image-to-video generation, including a start image and an end image that bracket a generated transition. That endpoint control is familiar in video generation. The more consequential claim is that a single render can contain multiple scenes and camera angles while retaining coherence, so a cut between shots need not require a separate generation.
Typography is intended to appear inside the scene rather than as an overlay pasted on afterward. FLUX 3 also claims multilingual dialogue and lip sync.
| Capability | What FLUX 3 is presented as doing |
|---|---|
| Text-to-video | Generate a clip from a prompt, with native dialogue, effects, and ambience. |
| Image-to-video | Generate from an image, including transitions between supplied start and end frames. |
| Multi-shot generation | Hold multiple scenes and camera angles in one coherent sequence. |
| Typography | Render text as part of the scene. |
| Dialogue and lip sync | Support English dialects and listed languages including Chinese, Spanish, French, Japanese, Turkish, German, Russian, Italian, Indonesian, Hindi, and Punjabi. |
Its style range is presented as a differentiator. Outputs can be raw, handheld, animated, nostalgic, or deliberately strange rather than locked into a polished cinematic look. The presenter contrasts that with models that need substantial negative prompting to avoid turning every prompt into a film-trailer aesthetic.
Black Forest Labs also says FLUX 3 combines world knowledge from pre-training with real-time grounding, allowing synchronized multi-camera generations.
Continuation preserves motion and sound, not just a final frame
The clearest workflow implication is video continuation. FLUX 3 can accept up to four seconds of an existing clip plus an instruction for what happens next, then continue from that moving and sounding context.
That differs from extending footage from a final still. The source illustrates the conventional approach as retaining one frame and generating anew from it. Everything before that frame is discarded: camera movement can stop, a voice can cut off mid-word, and the join can become conspicuous.
FLUX 3 instead receives four seconds of preceding motion and sound. The intended result is continuity in camera behavior, subject motion, and the voice and flow of ongoing dialogue. Chaining continuations could make the nominal 20-second generation limit less restrictive, since a creator can build a longer sequence in connected stages rather than request it all at once.
The constraint is equally clear: four seconds is a small context window. The presenter had not yet tested how far repeated continuations can go before drift or hallucinations appear, describing the feature as promising rather than proven.
Draft mode turns final rendering into a later decision
FLUX 3 offers Draft mode for faster, cheaper prompt previews. When a draft has the desired subject, composition, and motion, FLUX 3 can render the same video at full quality.
The displayed comparison of a sunset-over-ocean clip labels one version “DRAFT” and the other “FULL RENDER.” Black Forest Labs describes the latter as a full render rather than an upscale. Draft is also a separate setting from output resolution, rather than simply a smaller result enlarged afterward.
Black Forest Labs has not explained what makes the draft process cheaper or what changes internally between draft and final rendering. The presenter compares the workflow to settling a prompt with a lighter model before moving to a more capable model for the final result. In practical terms, Draft mode is a cost and iteration feature, not a separate creative capability.
Treat it as a strong option, with important workflow gaps
The practical decision is to test FLUX 3 against the prompts and workflows already in use. Black Forest Labs’ own numbers place it at or near the top of the field on its first video-model attempt, while the presenter describes it as a solid competitor for unusual concepts and vague prompts.
The shown comparisons place FLUX 3 beside Seedance 2.0 for an iceberg-calving scene and Gemini Omni Flash for a cheetah chasing a gazelle. The presenter particularly emphasizes the distinctiveness of FLUX 3’s outputs.
There are material limits. Black Forest Labs labels FLUX 3 a preview model. Full HD output is described as upscaled rather than generated natively at that resolution. Video editing is not available, and neither is reference support for images and videos, though both are listed as coming. Continuation remains capped at four seconds of input context.
Twenty seconds is longer than the roughly 10-second generation length described as common, but it is not the field’s ceiling: Seedance 2.5 is shown at 30 seconds. Longer clips are becoming part of the competitive baseline.
FLUX 3 Image has been announced but not released. Black Forest Labs says preliminary results indicate stronger handling of complex prompts and text generation than earlier FLUX image models, including high-accuracy multilingual text. No release date is given beyond its roadmap status.