
MiniMax H3 Makes 2K Video With Native Stereo Audio
MiniMax H3 generates 15-second 2K clips with native stereo sound and tops the video editing leaderboard at 1,130 Elo, priced at 0.8 yuan per second.
Sound That Was Never a Separate Step
MiniMax released H3 on July 31, 2026, an omni-modal model that takes text, images, video, and audio in a unified context and produces video up to 15 seconds at 2K resolution with native stereo audio. It is live in the platform API under the model ID MiniMax-H3 and in the consumer Hailuo AI app, at 0.8 yuan per second of output — roughly $1.95 for a 15-second 2K clip.
- 15-second clips at 2K resolution with native stereo audio, generated in the same pass rather than dubbed afterwards
- Ranked first globally in video editing with 1,130 Elo across 5,043 blind preference samples, plus second in text-to-video and third in image-to-video as of July 31
- 0.8 yuan per second of output — around one twelfth of the rate MiniMax cites for a leading competing model
- Open weights announced as coming "within days, subject to applicable laws and regulations"; as of early August the API is the only available path
Why Generating Audio in the Same Pass Matters
Almost every video generation pipeline in production today treats sound as post-production. The model produces silent frames, and audio is added by a second system that has to infer from the finished pixels what things should sound like. The result is the uncanny quality of a lot of AI video — footsteps that land slightly off, a voice whose mouth shape nearly matches, ambience that is plausible for the scene but not synchronized to it.
Generating both in one pass from a shared representation means the model is not guessing at correspondence, because it never separated the two. That is a structural advantage rather than a quality tweak, and it shows up most in exactly the moments a viewer notices: impacts, speech, and anything where timing carries meaning.
It is the same argument for native multimodality we made about Qwen3.7 Flash's native vision-language architecture — models that learn modalities together behave differently from models that meet them at a seam.
What Does First Place in Video Editing Actually Measure?
Blind human preference, which is both the most honest signal available in generative video and the one most easily over-read.
The 1,130 Elo figure comes from 5,043 paired comparisons where evaluators picked between outputs without knowing which model produced them. That resists the gaming that afflicts automated metrics, and it captures the thing users care about — does this look right — better than any frame-level score. Text-to-video second place and image-to-video third place put H3 credibly in the top tier across the board rather than winning one narrow category.
The caveats are worth stating plainly. Elo positions in a fast-moving field describe a snapshot, in this case as of July 31, 2026. Preference evaluations reflect the prompt distribution used to gather them. And "first in video editing" is a specific task — modifying existing footage — not a claim about generation quality overall.
Is This Actually an Open-Weights Release?
Not yet, and the distinction deserves care. MiniMax has said weights will follow "within days, subject to applicable laws and regulations," but at the time of writing they have not shipped. Today the model is reachable through the platform API and the Hailuo app only.
That matters because the open-weight video category is genuinely thin. Open text models are abundant, and open image models are well established, but video generation has stayed largely behind APIs because the training cost is enormous and the weights are correspondingly valuable. A frontier-tier video model with published weights would be a meaningful addition to what teams can self-host — which is why it is worth waiting for the actual release rather than filing this under open source in advance.
Our AI coverage has tracked how quickly the open-weight frontier has moved this year, including Kimi K3 shipping 2.8T parameters with a 1M-token context. Video has been the conspicuous gap in that story.
Who Is the Pricing Aimed At?
Volume producers, obviously, but the more interesting effect is on iteration.
At roughly $1.95 for a 15-second 2K clip with audio, the cost of a discarded take approaches nothing. That changes creative process more than it changes budgets: when generating an alternative is cheap, you generate ten and pick, rather than crafting one prompt carefully and accepting the result. Anyone who has worked with expensive generation credits knows how much that constraint shapes what gets made.
MiniMax positions the rate at about one twelfth of a leading competing model's. Price comparisons across video models are slippery — resolution, duration, audio inclusion, and quality tier all vary — so the useful takeaway is directional rather than exact: this is aggressively priced for the capability class, and the pressure it puts on the category is real.
What to Watch Next
Two things. First, whether the weights ship as described and under what licence, because that determines whether this is a cheap API or a self-hostable capability. Second, whether the audio quality holds up outside the evaluation set — native stereo generation is the differentiating claim, and it will be judged on speech synchronization and impact timing in real use rather than on a leaderboard.
For teams building on video generation today, the practical move is the cheap one: run your own prompts through the API and compare against whatever you currently use. At this price, finding out costs less than reading about it.
Sources: MarkTechPost — August 1, 2026; MiniMax research blog — July 31, 2026; South China Morning Post — August 2026.
More AI Stories
DeepSeek V4-Flash 0731 Tops V4-Pro at a Third the Price
DeepSeek's retrained V4-Flash 0731 beats its own V4-Pro preview on every published agentic benchmark at $0.28 per million output tokens, MIT licensed.
Local LLM Servers Compared: Ollama vs vLLM vs llama.cpp
Ollama, llama.cpp, vLLM and LM Studio all serve local models. Here's which one fits your hardware, from a 16GB laptop to a four-GPU workstation.
Oracle Puts Gemini Models Inside Fusion AI Agents
Oracle is bringing Google's Gemini models into Fusion Applications AI Agent Studio, letting thousands of enterprise customers build agents on Gemini.



