MiniMax H3
MiniMax H3 is MiniMax's general-purpose multimodal generation model, launched July 31, 2026. The research post says H3 jointly understands text, images, video, and audio, then writes video with native stereo sound at up to 2K and 15 seconds. API docs name the model MiniMax-H3 and list text-to-video, first/last-frame image-to-video, and reference generation from images, video clips, and audio. Consumer playground traffic goes through Hailuo. MiniMax says it plans to open the weights, subject to law. Confirm the current weight drop before you assume a local checkpoint.
Key Features
- Unified multimodal context: Official example: take camera movement from Video 1, a character from Image 2, vocals from Audio 3, and describe the relationship in text.
- Output specs: 768P or 2K, duration 4-15 seconds (integers only), 24 fps on the models table, common or adaptive aspect ratios.
- Reference limits: Up to 9 images; up to 3 video clips, each 2-15s, total video input at most 15s; width/height in 256-5760.
- Production claims: MiniMax highlights instruction following, text and brand rendering, and video-to-video motion transfer for ads, e-commerce, UI, and games.
- Open-weight intent: The July 31 post said weights would open in the coming days. Recheck the blog and GitHub before you plan an offline deploy.
Use Cases
- Brand and product video that must keep a logo or pack shot readable.
- Reference-driven shots when you already have a still, a camera move, and a voice bed.
- Draft at 768P, finish at 2K with the regeneration endpoint.
Pricing
H3 is not in MiniMax video point packages. Official pay-as-you-go (accessed 2026-08-17):
| Output | List price |
|---|---|
| 2K | $0.13 / second |
| 768P | $0.08 / second |
Input audio is free. The first 5 images are free, then $0.04 per extra image. Input video is billed at the same per-second rates as output at that resolution. Regenerating a 768P result to 2K is $0.05 / second of the new output, and original input materials are billed again. MiniMax's launch post said 2K per-second price is less than one-third of mainstream models; treat that as vendor copy and use the table above for quotes.
Getting Started
- Read the H3 post and try Hailuo if you only need a playground.
- For API work, use the pay-as-you-go video guide and the v2 create endpoint, not a Hailuo 2.3 package.
- Start at 768P for prompt search, then regenerate to 2K.
- If you need weights, verify they actually shipped. The July post was a plan, not a download link.
Frequently Asked Questions
Is this Hailuo 02 under a new name?
No. Hailuo 2.3 remains on the legacy video table (1080p/768p, 6-10s). H3 is the next-gen multimodal model with 2K and 4-15s.
Can I buy H3 with video points?
Not yet. The packages page says H3 is unsupported there. Use pay-as-you-go or a custom sales plan.
How does it compare to FLUX 3?
FLUX 3 is Black Forest Labs' multimodal generator with up to 20-second clips. H3 is MiniMax's 2K stereo path with a promised open-weight track.
Alternatives
- FLUX 3: Draft-then-render video and audio from BFL.
- Veo 3: Google video if you already live in Gemini.
- MiniMax: Company hub for Code, Speech, and Music next to H3.
Tips
- Write the relationship between references in the prompt. The model is built to bind those inputs, not to guess them.
- Stay inside the 15-second input-video budget. Extra clips will not fit the published limits.
- Recheck pay-as-you-go before a campaign. Package SKUs still describe Hailuo, not H3.
Conclusion
H3 is the MiniMax video model to evaluate in late 2026 if you need multimodal references and 2K stereo shorts. Start on Hailuo or the pay-as-you-go API, then confirm whether open weights have actually landed. First-party starting point: the H3 research post.
Comments
No comments yet. Be the first to comment!
Related Tools
MiniMax M3
www.minimax.io/blog/minimax-m3
MiniMax open-weight frontier model (June 1, 2026): 1M-token MSA context, native image/video input, desktop computer use, and discounted API pricing.
FLUX 3
bfl.ai/models/flux-3
Black Forest Labs multimodal model for video, audio, upcoming images, and action-prediction, with up to 20-second clips, native audio, and draft mode.
Seedance 2.5
seed.bytedance.com/en/seedance2_5
ByteDance Seed's audio-video model for 30-second storytelling, precise reference control, extend-twice generation, and pro tools like green-screen editing.
Related Insights
What Locks Codex Is the Picker, Not the Models
You already pay for OpenCode Go, Grok, and Z.ai, but the Codex picker still shows mostly GPT. The community project codex-router does not teach another install ritual. It puts subscriptions you already bought back into the selector, keeps keys on the machine, and leaves native GPT alone.
Seven AI Coding CLIs, Six Months: No Matter How Strong the Model, Work Needs Supervision
Claude Code, Codex, opencode, pi, omp and DeepSeek Harness all have personalities. After six months of deep use I run a division of labor: pi for the fastest cheapest reviews, omp for complex PRs, DeepSeek Harness on V4 Flash for high-frequency low-cost review, and Claude Code, Qoder and Cursor for writing. No matter how strong the model, work needs supervision — ideally from an independent third party.
After I Connected Obsidian to OpenClaw, It Started Helping Me Make Decisions
Once Obsidian stopped being just a place to store notes and started working with OpenClaw, it began helping me organize context, connect information, and improve real decisions.