What is FLUX 3 Video?
FLUX 3 Video is Black Forest Labs' first video model, released August 4, 2026 through the lab's own API and select partners. It is not yet callable on Unifically. We will open it here once the integration is live, and every spec on this page comes from the release and will be re-verified against the callable API on day one.
The model generates 5 to 20 second videos with native audio in one pass: dialogue, sound effects, and ambience render together with the picture instead of being added afterward. It comes out of FLUX 3, a single multimodal model trained jointly on images, video, and audio, which is why one generation can switch scenes and camera angles, render readable typography inside the scene, and keep sound tied to the physical event that makes it.
Output is 720p natively at 24 fps, with 1080p available through an upscaling pass.
Key features of FLUX 3 Video
- 5 to 20 seconds in one generation. Any whole duration in that range, or auto to let the model fit the length. Most competing models stop at 8 to 15 seconds.
- Native audio by default. Speech, effects, and ambience generate in the same pass. Dialogue supports English, Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi, Punjabi, and more, with lip-sync.
- Up to 10 keyframes. Start from one image, pin a first and last frame, or lay out a storyboard of up to ten frames, including frames tied to exact timestamps.
- Video continuation. Feed in up to 4 seconds of existing video and audio and describe what happens next. Motion, camera behavior, dialogue, and sound carry across the seam.
- Multiple shots in one video. Scene changes and camera cuts inside a single generation, with the sequence staying coherent.
- Draft mode. A fast, cheap preview first; a draft-enhance pass then re-renders the chosen draft at full quality with the same seed, composition, and motion.
Best for
Dialogue-led scenes
Lip-synced speech in 13+ languages renders with the picture, so a talking scene comes back finished.
Storyboard-driven shots
Pin up to 10 keyframes, with timestamps, and the model generates one continuous video that hits each one.
Longer single-pass videos
A 20-second ceiling covers a full short-form story without joining separate generations.
Extending existing videos
Continuation picks up motion and audio from up to 4 seconds of source video and keeps going.
Cheap iteration
Draft previews let you test prompts at low cost, then re-render the winner at full quality unchanged.
Use cases
Build short ads that open, turn, and close inside one 20-second generation instead of stitching three videos. Script a product shot as keyframes, pin the frames that matter, and let the model create the motion between them. For social formats, generate a 15-second vertical video with spoken dialogue in the viewer's language and skip the dubbing pass. Use continuation to grow a strong 10-second result into a scene that runs past the single-pass limit. And when a brief is vague, run a batch of draft previews first, pick the best one, and promote only that draft to a full-quality render.
Limitations
Native output stops at 720p; 1080p comes from an upscaling pass rather than direct generation. The 20-second ceiling applies per generation, so longer pieces need continuation passes. Reference-based generation from image, video, and audio combinations is on the roadmap but not released. Quiet, low-motion scenes can come back with little audible sound unless the prompt names the audio it wants. Scenes with several coordinated actors, or full-body motion where hands, feet, and posture all matter at once, are less reliable than single-event shots. And on the Arena image-to-video board it ranks #5 as of August 16, 2026, behind MiniMax H3, Seedance 2.0, Gemini Omni Flash, and Grok Imagine 1.5, so image-driven work has stronger options today.
FLUX 3 Video vs SeeDance 2.0
On the Arena text-to-video board, FLUX 3 Video ranks #2 at 1496 Elo, one place above Seedance 2.0's 720p build at 1478, as of August 16, 2026. The order flips on image-to-video, where Seedance 2.0 holds #2 at 1479 and FLUX 3 Video sits at #5 at 1453. SeeDance 2.0 also reaches 1080p and 4K output on Unifically today, against FLUX 3's native 720p. Until FLUX 3 Video lands here, SeeDance 2.0 is the practical pick; once it does, FLUX 3's prompt-led generation, keyframe control, and continuation are the reasons to switch a text-driven pipeline over.
Availability
Black Forest Labs announced FLUX 3 on July 23, 2026 and released FLUX 3 Video through its own API and select partners on August 4, 2026. Reference-based generation, FLUX 3 Image, and an open-weight FLUX 3 Dev are on the lab's roadmap. Unifically pricing, variants, and exact callable parameters will appear on this page once the model is live here.
What FLUX 3 Video can do
#2 on the Arena text-to-video board
FLUX 3 Video holds 1496 Elo from 1,288 votes on the Arena text-to-video board as of August 16, 2026, behind only gemini-omni-flash and ahead of Seedance 2.0, Muse Video, and MiniMax H3. That rank came within two weeks of release.
20 seconds in a single generation
One 20-second pass with no cuts and no stitching: kids build a snowman, the light moves through sunset into night, and the season turns while the snowman and yard stay recognizable from first frame to last.
Keyframes pinned to exact timestamps
Four frames pinned at 0, 5, 10, and 15 seconds set this shot, from a rock headland above a harbour to a colossus standing above the clouds. The model returns one continuous video that hits every frame, which keeps long generations from drifting.
Continue an existing video across the seam
Continuation takes up to 4 seconds of source video with its audio and generates what happens next. Movement, camera behavior, dialogue, and sound carry over the join instead of resetting.
Sound generated on the frame it happens
Audio renders inside the same pass as the picture, so the roar of the flames and the clatter of the wok land exactly where the fire and metal move. No separate audio model, no syncing afterward.
Physics that runs in reverse
Every shard of a shattered bottle slides and leaps back into place until the bottle stands whole. Reverse motion is a hard physics test: the model has to keep mass, timing, and glass behavior believable while running causality backward.
FAQs
People also ask
FLUX 3 Video is the first video model from Black Forest Labs, the lab behind the FLUX image line. It generates 5 to 20 second videos with native audio from text, images, keyframes, or an existing video. It is part of FLUX 3, a single model trained jointly on images, video, and audio.
Not yet. Black Forest Labs released it through its own API and select partners on August 4, 2026. We will open it on Unifically once the integration is live, and this page will get the playground, callable parameters, and pricing at that point.
5 to 20 seconds per generation. You can also leave duration on auto and let the model fit the length to the content. Video continuation runs 5 to 15 seconds per pass, so longer sequences chain from there.
720p natively, with 1080p available through an upscaling pass. Output runs at 24 fps, with aspect ratios from 21:9 widescreen through 16:9, 4:3, 1:1, 3:4, and 9:16 vertical.
Yes, by default. Dialogue, sound effects, and ambient sound render in the same pass as the picture, so a sound lands on the frame where its event happens. Spoken dialogue supports English, Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi, Punjabi, and more, with lip-sync.
A text prompt on its own, 1 to 10 images, a first and last frame pair, up to 10 keyframes (including keyframes pinned to exact timestamps), or up to 4 seconds of existing video and audio to continue from.
A fast, cheaper preview of a prompt. When a draft looks right, a draft-enhance pass re-renders that exact generation at full quality with the same seed, subjects, composition, and motion, so the final video matches the preview you approved.
As of August 16, 2026 it holds
