Back to blog

Hailuo 3 vs Veo 3: Which AI Video Model Wins in 2026?

August 7, 2026By Bilal Azhar

MiniMax H3 and Google Veo 3.1 compared on duration, 2K vs 4K, native audio, reference control, aspect ratios, and what a finished clip really costs.

Pick MiniMax H3, the model consumers know as Hailuo 3, when the shot needs to run longer than eight seconds, needs a non-standard frame like 21:9 or 4:3, or needs a specific face and outfit carried across takes from your own reference images. Pick Google Veo 3.1 when the shot needs spoken dialogue baked in, needs a 4K master, or needs photoreal live-action that survives close inspection. H3 launched on July 31, 2026 and is roughly a week old at the time of writing, so the honest framing is a mature model against a very new one with a strong spec sheet.

Scorecard

CategoryWinner
Max single-generation lengthHailuo 3 (15s vs 8s)
Resolution ceilingVeo 3.1 (4K vs 2K)
Native dialogue and soundVeo 3.1
Reference images per generationHailuo 3 (up to 9 vs 3)
Aspect ratio rangeHailuo 3 (6 ratios vs 2)
Photoreal physicsVeo 3.1
Stylized and anime motionHailuo 3
Cost per secondHailuo 3
Open weights availableHailuo 3 (base model only)
Enterprise availabilityVeo 3.1

A naming note before the specs, because it trips people up. Veo 3.1 is the current shipping model, in three tiers: Veo 3.1, Veo 3.1 Fast, and Veo 3.1 Lite. When people search for Veo 3 they almost always mean the Veo 3 family, and Veo 3.1 is what they will actually get. On the other side, MiniMax H3 is the official name and Hailuo 3.0 is the consumer-facing one. Same model.

What Are the Hard Specs on Each Side?

Two numbers separate these models more than anything else: single-generation length and resolution ceiling. H3 produces 4 to 15 seconds in one pass at up to 2K. Veo 3.1 accepts only 4, 6, or 8 seconds per generation and reaches 4K, with 1080p and 4K restricted to the 8-second option according to Google's Gemini API video documentation.

SpecMiniMax H3 (Hailuo 3)Google Veo 3.1
Single-generation duration4-15s upstream, 5-15s on Morphed4s, 6s, or 8s
Longer runtimesNot needed under 15sScene extension, chained past 60s
Max resolution2K, 24fps4K (1080p/4K need the 8s option)
Native audioStereo audio upstreamYes, dialogue and effects, 48kHz
Input modesText, image (first and last frame), reference setText, image, up to 3 reference "ingredients"
Reference controlUp to 9 images, 3 videos, 3 audio clips, 12 files maxUp to 3 reference images
Prompt lengthUp to 7,000 charactersStandard prompt field
Aspect ratios21:9, 16:9, 4:3, 1:1, 3:4, 9:1616:9 and 9:16 in the API
Open weightsH3-Base (33B) only, released Aug 3, 2026None
Reported price per second~$0.13/s at 2K, ~$0.08/s at 768p$0.40/s standard with audio, $0.10/s Fast with audio at 720p, Lite from ~$0.03/s without audio
WatermarkingNo mandatory watermark disclosedSynthID on every output
Where it runsMiniMax API, resellers, MorphedGemini app, Flow, Vertex AI, Gemini API, Morphed, third-party APIs

The aspect ratio row is the one most roundups skip. If you cut 21:9 for a brand film or 4:3 for an editorial piece, H3 handles it natively and Veo 3.1 through the API does not. You would be letterboxing or cropping down from 16:9, which costs you vertical resolution on a model you are paying a premium for precisely because of resolution.

Which Model Handles Physics and Realism Better?

Veo 3.1 for photoreal live-action, H3 for stylized and full-body human motion. Independent comparisons through 2026 have consistently put Veo 3.1 and Sora 2 Pro at the top for realism, physics coherence, and cinematic staging. The Hailuo line built its reputation on a different axis: the viral noodle-eating demo that other models could not match, and per-frame limb stability through choreography.

That split holds up under the kind of shot you actually commission. Ask for a hand pouring espresso into a glass cup in raking window light and Veo 3.1 is the safer first call. Ask for a full-body dance take, an ink-wash sequence, or an anime fight beat and H3 is the safer first call. Reporting on Hailuo 2.3 already flagged it as the strongest stylized-motion model in its price bracket, ahead of Veo 3.1 Lite on character body fluidity, and H3 inherits that lineage.

Veo's known failure mode is multi-character interaction: two or more people touching, handing things over, or colliding is where its physics gets unconvincing despite otherwise excellent realism. H3's known failure modes are inherited from the line: legible in-frame text is unreliable, and complex finger work like guitar fretting or typing still degrades, less often than in Hailuo 02 but often enough to plan around.

How Do You Lock a Character's Face Across Shots?

This is H3's clearest structural advantage and the reason to test it first on any campaign with a recurring character. H3 accepts up to nine reference images, three reference videos of 2 to 15 seconds each, and three reference audio clips in a single request, capped at twelve files total. Veo 3.1's Ingredients to Video accepts up to three reference images.

Nine versus three is not a bragging-rights number. It changes what you can specify. With three ingredients you assign one to the face, one to the product or prop, and one to the location or style, and you are out of slots. With nine you can give the model four angles of the same face, two of the garment, one of the prop, and two style plates, which is much closer to how a real reference board is assembled. The video and audio reference slots go further still: a reference video can carry the camera move, and a reference audio clip can carry the rhythm.

Veo 3.1's counterargument is continuity across shots rather than within one. Scene extension takes a finished clip and continues it forward while preserving the subject, which is how Veo strings 8-second generations into sequences that run past a minute without the face resetting. If your problem is "the same person across a 60-second narrative," Veo's extension path is well trodden. If your problem is "this exact person, this exact jacket, this exact product, in one 12-second shot," H3's reference stack gives you more handles. On Morphed the reference workflow runs as a dedicated model at MiniMax H3 Reference.

Does Either Model Give You Usable Sound?

Veo 3.1, without qualification. It generates synchronized dialogue, ambient sound, and effects natively at 48kHz, with room tone that shifts between indoor and outdoor reverb. That is a genuine production capability, not a novelty, and it is the single strongest reason to pay Veo prices.

H3 ships native stereo audio upstream as part of its omni-modal output. The caveat matters more than the headline: Morphed's H3 integration does not advertise audio output, so do not plan a Morphed H3 generation around getting spoken dialogue back. If you need sound on a clip generated on Morphed, the right model is FLUX 3 Video, which is built for native audio work, and the right workflow for H3 is to treat it as picture and lay sound underneath in your editor.

The contrast in practical terms: an 8-second Veo 3.1 generation with audio arrives as a finished shot with a voice in it. A 12-second H3 generation on Morphed arrives as picture you still need to score, and you got four extra seconds and a wider frame for it.

What Does a 10-Second Clip Actually Cost?

A 10-second finished shot is one generation on H3 and at least two on Veo 3.1, because Veo caps a single generation at 8 seconds. That structural difference compounds the per-second gap rather than sitting beside it.

PathGenerations neededReported rateApproximate cost
MiniMax H3, 2K, upstream1 x 10s~$0.13/s~$1.30
MiniMax H3 via a higher-priced reseller1 x 10s~$0.26/s~$2.60
MiniMax H3 on Morphed1 x 10s40 credits/s400 credits
Veo 3.1 on Morphed, audio on6s + 4s40 credits/s400 credits
Veo 3.1 standard with audio8s + extension$0.40/s~$4.00
Veo 3.1 Fast, 720p with audio8s + extension$0.10/s~$1.00
Veo 3.1 Lite, no audio8s + extensionfrom ~$0.03/s~$0.30

Three things to hold onto here. First, at the top tiers the gap is roughly threefold in H3's favor, which matches the reporting around H3's launch positioning as a price undercut. Second, the Veo 3.1 Fast and Lite tiers erase most of that gap, so "Veo is expensive" is only true of the standard tier with audio on. Third, H3's upstream pricing is genuinely unsettled: MiniMax's own documentation listed H3 as unsupported in its standard video packages while resellers were already charging for it, and quoted rates spread from about $0.13 to about $0.26 per second depending on who you buy from. Check the rate on the day you buy.

On Morphed the H3 number is fixed and public: 40 credits per second across text-to-video, image-to-video, and reference-to-video, so a 15-second clip is 600 credits and a 10-second clip is 400. Veo 3.1 runs on Morphed too, at 20 credits per second for silent picture plus 20 for audio, which means Veo with audio and H3 land on the same 40 credits per second. At identical pricing, the decision stops being about money and becomes purely about which model's takes survive review for your kind of shot, and you can run that test side by side in one account. Morphed is a paid platform with no free generation tier, so credits come from a plan.

Why Per-Second Price Understates Your Real Bill

Here is the observation that changed how we brief these two models internally: the unit that actually matters is cost per usable clip, not cost per second, and the two models fail in ways that cost different amounts to fix.

The method is simple enough to replicate. For every delivered shot, log the number of generations it took to get there and the reason each rejected take was rejected. Do not log seconds. After a dozen shots the pattern separates cleanly by model, and it is not the pattern the pricing pages suggest.

Veo re-rolls cluster around audio and instruction adherence. The picture is usually fine and the take dies because the line reading is wrong, the lip sync drifts, or the model added a sound you did not ask for. Those re-rolls are expensive twice over: you are paying the audio-on rate, and a Veo re-roll buys you 8 seconds of replacement footage at most.

H3 re-rolls cluster around hands, in-frame text, and identity drift on the harder reference sets. Those are cheaper re-rolls, both because the per-second rate is lower and because a single 12-second regeneration replaces the whole shot rather than one segment of it. There is no published re-roll rate for H3 yet, since the model is a week old. The nearest published figure is for Hailuo 2.3, where reviewers reported roughly one generation in three diverging noticeably from the written prompt. Treat that as lineage, not as an H3 measurement.

The practical consequence: a model that is three times cheaper per second but needs twice as many takes is only 1.5 times cheaper in reality, and a model whose re-rolls only replace 8 seconds at a time compounds badly on longer shots. Run ten shots through both before you standardize on either. For a broader field, see our best AI video generators roundup and the text-to-video model comparison.

How Hard Is Each One to Get?

Veo 3.1 has the wider distribution by a large margin: the Gemini app for casual use, Flow for creators on Google AI Pro and Ultra plans, Vertex AI for enterprise, the Gemini API for developers, Google Ads for advertisers, and a long list of third-party API resellers. If your company already has a Google Cloud relationship, Veo is a procurement formality.

H3 is newer and thinner on distribution. The MiniMax API is the primary path, resellers picked it up within days of launch, and multi-model platforms added it quickly. Morphed runs it as MiniMax H3 for text-to-video and image-to-video and as a separate reference model, which is the shortest route to testing it without provisioning MiniMax API credentials yourself. Morphed also carries Veo 3.1, so the head-to-head this article describes is something you can actually run in one place instead of juggling a Google Cloud project and a MiniMax account. Sign-up is at morphed.app/register.

The open-weights question deserves precision because a lot of coverage flattened it. MiniMax released H3-Base, a 33B-parameter model, on Hugging Face on August 3, 2026. It did not release the in-context 2K regeneration stage or the orchestration layer that handles multi-reference prompting. You can run H3-Base on your own hardware. You cannot reproduce the hosted 2K output with it. Calling Hailuo 3 open source without that qualifier misleads anyone planning a self-hosted pipeline.

Can You Use the Output Commercially?

Both yes, with different strings attached. Veo output is commercially usable through paid access such as Vertex AI or Gemini Enterprise, subject to Google's content policies, and every frame carries a SynthID invisible watermark by design. SynthID is built to survive cropping, color grading, and recompression. On Flow, lower subscription tiers also get a visible corner mark, which Ultra removes.

For paid media that matters more than it sounds. If your ad platform or client contract requires AI disclosure, SynthID is an asset. If your brand guidelines forbid any embedded identifier in delivered masters, it is a hard blocker, and it is not one you can negotiate away. H3 has no equivalent mandatory watermark disclosed, which cuts both ways: less friction, less provenance.

Pick Hailuo 3 If, Pick Veo 3 If

Pick Hailuo 3 (MiniMax H3) if:

  • The shot is longer than 8 seconds and needs to be one continuous take
  • You are delivering 21:9, 4:3, 1:1, or 3:4 and do not want to crop from 16:9
  • You have a real reference board: multiple angles of a face, a garment, a product
  • The work is anime, ink-wash, game CG, or any stylized register
  • The shot is full-body human motion, dance, or choreography
  • Volume matters and you are generating dozens of variants per concept

Pick Veo 3.1 if:

  • The clip needs spoken dialogue generated with the picture
  • You are delivering a 4K master and upscaling is not acceptable
  • The shot is photoreal live-action that will be scrutinized frame by frame
  • You need a narrative that runs past a minute with a consistent subject
  • You are inside a Google Cloud or Google Ads workflow already
  • Provenance watermarking is a requirement rather than an obstacle

When Neither One Is the Right Call

Three scenarios where both models are the wrong tool, stated plainly because most comparison pages will not.

A 20-second single-take talking ad with sound. Veo cannot do it in one generation at 8 seconds a shot, and H3 tops out at 15 seconds with no advertised audio output on Morphed. This is a FLUX 3 Video job on Morphed, which is built for exactly this shape of ad.

Legible on-screen text. Neither model renders in-frame typography reliably. Generate the text as a still in an image model or lay it in your editor, then animate around it. Burning a promo code into a generated frame will cost you more re-rolls than the shot is worth.

Anything requiring exact product geometry. If a label has to be readable and a silhouette has to match a spec sheet, a reference-driven image-to-video pass from a locked product still beats text-to-video from either model. Our Kling vs Runway comparison covers the models built around that constraint.

And the meta-answer nobody selling a single model will give you: on any given shot, the model that wins is frequently not the model that won last week. H3 is one week old. Veo 3.1 has been iterating since October 2025 and picked up 4K in January 2026 and a Lite tier on March 31, 2026. Both are moving. Any comparison page including this one is a snapshot, and the workflow that survives version churn is testing both on the same prompt rather than committing to a brand.

Frequently Asked Questions

Does Hailuo 3 support first and last frame control?

Yes. H3 accepts image-to-video with both first-frame and last-frame conditioning, which is the cleanest way to control where a shot starts and ends without over-describing the middle in the prompt.

How long is the prompt limit on each?

H3 accepts up to roughly 7,000 characters per request, which is long enough for a full shot list with per-reference role assignments. Veo 3.1 uses a standard prompt field and rewards structured cinematic description over length.

Can Veo 3.1 really do 4K?

It outputs at 3840x2160, with reporting describing native generation at 1080p and a strong upscale to 4K rather than native 4K sampling. For most delivery that distinction is academic. For heavy punch-ins and reframing it is not, so test on your own footage before promising a 4K master.

Which is faster to generate?

Veo 3.1 Fast and Lite are built for turnaround and are the quickest path on the Google side. H3 generation time scales with duration and resolution, so a 15-second 2K clip is meaningfully slower than a 5-second one. Draft short, finish long.

Is there a Veo 4?

Not in general availability as of August 2026. Veo 3.1 across its three tiers is the current family.


Test MiniMax H3 on your own shot list, in your own aspect ratio, before you standardize on anything. Start on Morphed

Related: MiniMax H3 | MiniMax H3 Reference | FLUX 3 Video | Best AI Video Generators | Best Text-to-Video AI Generators | Kling vs Runway