
FLUX 3 Video Goes GA With 20-Second Clips and Audio
Black Forest Labs opened FLUX 3 Video to every developer on August 4: 20-second clips with natively generated audio, starting at $0.06 per second.
Black Forest Labs moved FLUX 3 Video out of early access and into general availability on August 4, 2026, and the shift matters more than a typical status change. The model now generates clips up to twenty seconds long with natively produced audio — dialogue and sound effects included — from either a text prompt or a still image, all through one API surface. For teams that have been stitching together a separate video model, a separate voice model, and a separate sound-effects pass, that consolidation is the headline.
- General availability began August 4, 2026, after an early-access period that opened July 23
- Clips run up to 20 seconds with natively generated dialogue and sound effects
- Six published per-second price tiers, ranging from $0.06 to $0.54 per second
- Text-to-video, image-to-video, and video continuation all run through a single API
What Makes Native Audio Different From a Dubbing Pass
Most video generation pipelines in 2026 still treat sound as post-production. You render the visual, then you hand the result to a text-to-speech system and a foley model, then you spend real engineering effort on alignment so that lips, footsteps, and door slams land on the right frames. Every handoff is a place where synchronization drifts.
Generating audio inside the same model removes those handoffs. The system that decides a character turns their head is the same system deciding when the accompanying sound arrives. It is a smaller architectural change than it sounds, and a much larger practical one — the failure mode it eliminates is the one that made earlier AI video feel uncanny even when the pixels looked right.
How Much Does FLUX 3 Video Cost to Run?
Black Forest Labs published six per-second tiers spanning $0.06 to $0.54, which lets teams pick a resolution and quality point rather than accepting a single blended rate. A ten-second clip at the entry tier lands around sixty cents; the same clip at the top tier is closer to five dollars and forty cents. That spread is useful for the realistic workflow, where a team iterates cheaply on composition and framing, then re-renders only the approved shot at the higher tier.
The video continuation capability deserves attention here too. Because the model can extend an existing clip, a twenty-second ceiling is a per-call ceiling rather than a hard limit on finished output. Chaining continuations is not the same as generating a coherent two-minute sequence in one shot, but it does open the door to longer-form work built from consistent segments.
Why an API-First Launch Is the Right Read on Adoption
There is no consumer app attached to this release, and that is deliberate. FLUX 3 Video is arriving as infrastructure, aimed at the product teams building storyboard tools, ad-variant generators, localization pipelines, and previsualization software. Those buyers care about per-second pricing, latency, and a stable API contract far more than they care about a polished editing interface.
That pattern is consistent with what we have seen across the broader tooling layer this year — capability shipping as a primitive that other people compose. It mirrors the direction of the agent frameworks and runtime platforms covered in our recent AI coverage, where the winning move has repeatedly been to expose a clean interface and let the ecosystem build the surface. The same logic applied to Cloudflare's open agent workspace platform and to NVIDIA's single-class agent framework: the primitive ships, the products follow.
What to Watch Next
The interesting question for the rest of 2026 is whether twenty seconds with synchronized audio is enough for production work or merely enough for prototyping. The honest answer is that it depends entirely on the format. For social-length content, product demos, and animated explainer segments, it is already sufficient. For narrative work, the continuation feature will carry the weight, and how gracefully it preserves character and lighting consistency across chained segments is the thing worth measuring.
Either way, the consolidation of picture and sound into a single generation step is the durable part of this release. Everything downstream of it gets simpler.
Sources: XenoSpectrum — August 2026; The Decoder — August 2026; VentureBeat — July 25, 2026.
More AI Stories
Suno Watermarks AI Songs to Make Origins Verifiable
Suno will embed an inaudible signature in every track it generates, giving streaming platforms a way to identify AI-made music automatically.
Firmus Raises $2B for Renewable-Powered AI Factories
Firmus closed a fully subscribed $2 billion round at a $10.5 billion valuation to build renewable-powered AI data centers across Asia-Pacific.
ChatGPT Reasoning Slider Puts Thinking Effort in Your Hands
ChatGPT's new reasoning slider spans five effort levels, and the updated GPT-5.6 Sol makes factual errors 68% less often than GPT-5.5 Instant.



