Seedance 2.5 and MiniMax H3 launched on the same day at the end of July. Both generate audio in the same pass as picture, both take large reference packages, and both arrived with a wall of demo reels that prove very little. So we did the boring thing: read both spec sheets properly, listed everything each one actually ships, then wrote a prompt designed to break things and ran it through both.
Seedance 2.5MiniMax H3Feature comparisonPricingHead-to-head~11 min read
The rig — same prompt, same references, one take each
Seedance 2.5ByteDance Seed
MiniMax H3MiniMax · Hailuo 3
Shared timecode — jump both clips to a beat00:00 / 00:15
Each marker is a graded beat. The two hardest are 05–08, a held hand pose in macro followed by twin web strands with real line tension, and 13.5, where the same face has to come back after eleven seconds under a helmet.
The two models
Same launch day. Different companies, different bets.
Before any judgment about quality: what each one is, who builds it, when it shipped, everything it publicly claims to do, and where you can actually get at it.
ByteDance Seed · launched 31 July 2026
Seedance 2.5
Built by ByteDance Seed, ByteDance's research division. Previewed at the Volcano Engine FORCE conference in late June and shipped at the end of July. ByteDance skipped 2.1 through 2.4 entirely, framing this as a generational jump rather than a point release. The pitch is long-form: a whole scene, with its own internal shot changes, out of one generation.
DurationUp to 30 seconds in a single pass, with multi-round extension beyond that. No stitching and no extension pass needed to reach half a minute.
Multi-shotOrganises several logically connected shots inside one generation, with improved transitions and scene changes, so a story unfolds without a cut you have to assemble.
ReferencesUp to 30 images, 10 video clips and 10 audio clips in a single pass — 50 reference inputs total.
Reference typesCharacter, motion, clay-render and creative references, plus R2V support with green-screen and white-model references for blocking.
EditingTimestamp-level control for targeted edits to audio and video without regenerating the whole shot.
AudioNative joint audio-video generation, inherited from the Seedance 2.0 architecture.
CameraProfessional camera movement and performance blocking, including 3D camera blockout controls.
WeightsClosed. Platform and API access only.
ResolutionAnnounced with native 4K. Verify inside your own workspace before planning around it — see the note below.
Built by MiniMax, the lab behind Hailuo 02 and Hailuo 2.3. Officially named MiniMax H3; Hailuo 3, Hailuo 3.0 and Hailuo 03 all point at the same model. The pitch is that it stops behaving like a video model — one transformer reading text, images, video and audio as a single context.
Duration4 to 15 seconds per generation at 24fps, with platform tooling to extend further.
Resolution768p and 2K, native — not a separate upscale pass. The 1440p tier circulating in the community is not an official option.
Audio32 kHz native stereo generated in the same pass as picture: dialogue, effects and room tone, timed to on-screen action. No separate audio stage.
ReferencesUp to 9 images, 3 video clips and 3 audio clips — 12 files total — usable for identity, motion, framing, atmosphere or voice without forcing them as keyframes.
ModesText-to-video, first- and last-frame image-to-video, and reference-to-video with mixed inputs.
EditingInstruction-based: request a change in a sentence instead of re-rolling the shot. Also does video-to-video motion transfer and voice transfer.
Prompt limit7,000 characters.
WeightsOpen. A 33B dense single-stream Omni-Transformer, published on Hugging Face in early August 2026 under the MiniMax H3 Community License.
BenchmarksRanked #1 in Video Editing, #2 in Text-to-Video and #3 in Image-to-Video on the Artificial Analysis leaderboard at launch.
Both launches are recent enough that the public record still contradicts itself. Seedance 2.5's announced native 4K has been reported as not making it into the shipped model, with some routes topping out well below it. MiniMax marketed H3 as open-weight before any weights existed, so comparison pieces published in that window state flatly that neither model is open — those sections are now stale.
Access is layered too. On Seedance, "available" means one thing inside Dreamina or Jimeng and something else on BytePlus ModelArk and Volcano Engine, where the developer contract landed after the consumer product. Verify the model label inside your own signed-in workspace before you plan a pipeline around a number you read in a blog post. Including this one.
Side by side
The scene model and the shot model.
Read down this table and the strategic difference is obvious. Seedance is betting the unit of work is a scene. MiniMax is betting it is a shot. That decides where each one sits in a pipeline more than any quality score will.
Published specifications — August 2026
Capability
Seedance 2.5
MiniMax H3
Company
ByteDance Seed
MiniMax
Launched
31 July 2026
31 July 2026
Predecessor
Seedance 2.0
Hailuo 2.3
Max single generation
30 seconds
15 seconds
Extension
Multi-round extension
Platform extend tooling
Frame rate
Not officially published
24 fps
Native audio
Yes
Yes — 32 kHz stereo
Image references
Up to 30
Up to 9
Video references
Up to 10
Up to 3
Audio references
Up to 10
Up to 3
Multi-shot in one take
Yes, with scene changes
Yes, within 15s
Editing model
Timestamp-level
Instruction-based
Prompt limit
Not published
7,000 characters
Weights
Closed
Open — 33B, community license
Billing unit
Tokens, via provider
Per generated second
50Seedance reference inputs, max
12H3 reference files, max
33BH3 open-weight parameters
0Days between the two launches
Cost
You cannot divide one of these prices by the other.
This is the part most comparisons get wrong. The two models do not bill in the same unit, so any figure produced by comparing them directly is meaningless.
MiniMax publishes a straightforward per-generated-second rate for H3 — roughly $0.13 per second at 2K and $0.09 per second at 768p, with extra charges for excess image references and for input video duration. A 15-second reference clip is billed as fifteen seconds of input on top of your output.
Seedance 2.5 bills through token-based provider rates that change depending on whether video input is present, and the numbers differ across Volcano Engine, BytePlus ModelArk and aggregators. Video-input duration feeds into token consumption. There is no single published per-second figure to hold against H3's.
Under that framing the answer flips depending on the work. A short, bounded 2K shot is cheap to review and cheap to throw away, so H3 wins on iteration economics. One accepted 30-second scene can remove four generations, three stitch points and a continuity pass, so Seedance wins when the retry rate is low and the scene is long.
If you are testing prompts, H3 is cheaper to fail at. If you are delivering scenes, Seedance is cheaper to succeed at.
Method
We tested at fifteen seconds. Neither model stops there.
Seedance 2.5 will run to thirty seconds in one pass. H3 will run to fifteen. Testing at fifteen is not a statement about what either model can do — it is the only window where the two are doing the same job, and therefore the only window where a result means anything.
Running Seedance at thirty and declaring it the winner measures a spec sheet, not a model. Long-form generation, multi-shot continuity across half a minute, and Seedance's timestamp-level editing all deserve their own test, and they will get one. This one is about what happens inside the overlap.
Everything else is held constant: 16:9, the same two reference images attached to both runs, one generation each, no rerolls. Both outputs at each platform's highest available native resolution, downscaled to a shared 1080p canvas only at the comparison stage, with both original files retained.
Why this prompt
Detail alone stopped discriminating between frontier models a while ago — both of these will render a beautiful static portrait. To learn anything you have to attack the parts that still break, and you have to be able to see the break while scrolling, not in a lab. Every beat below is independently gradeable. Use the ruler at the top of this page to jump both clips to it.
00:00 – 00:03
Bare face on the ledge
Close on his unmasked face, rain on the skin, breath visible, neon across one cheek. The camera eases back to reveal the helmet held in both gloved hands, weight pulling the wrists down.
Watch for: whether the face matches the portrait reference at all, and whether the helmet has believable mass in the hands or floats like a decal.
00:03 – 00:05 · key
The helmet goes on
He raises it, turns it once so the visor throws the skyline back in curved miniature, then lowers it over his head — rim past the ears, hair pressing flat, shell seating at the collar, a roll of the neck to settle it.
Watch for: hands manipulating a rigid object with occlusion. Fingers should grip a hard shell, not squash it. No clipping through the skull, no skin visible once seated.
00:05 – 00:06.5 · key
The hand sign, in macro
Camera drops to the right hand. Index and little fingers snap straight up and spread; middle, ring and thumb fold flat to the palm. Glove leather creases where the knuckles fold. The wrist cocks back and holds.
Watch for: a specific held finger configuration at close range. The hardest thing in the shot. Count the digits, check the joint bends, check whether the pose survives the beat or quietly morphs.
00:06.5 – 00:08 · key
Twin strands, braiding
Macro on the two raised fingertips, rain beading on the glove. The wrist snaps and two separate web strands fire in parallel, braiding into one taut line a short distance out. Recoil travels back through fingers, wrist, forearm.
Watch for: two strands actually converging. Models cheat two ways — collapsing to one strand immediately, or letting the pair drift and never merge. The line should have thickness and sag, not glow.
00:08 – 00:11
Swing, car roof, wall run
Wide on a brownstone street with a neon MARA'S DELI sign. He swings low, boots skimming a parked car roof with real compression, redirects off it, then lands feet-first on a wall and runs four steps along it before pushing off and firing again.
Watch for: the sign staying spelled correctly, and whether the car roof and the brick push back. Floating contact is the most common failure at this speed.
00:11 – 00:13.5
Street, taxi, launch
Low over a street of yellow taxis, headlights smearing on wet asphalt. The web releases, boots hit pavement for two hard strides, a horn sounds, he plants a boot on a taxi and pushes off the roof to rocket up. Then a one-second hold at the apex.
Watch for: reflections on wet asphalt moving coherently with the camera rather than sitting painted in place. Also whether the splash lands on the same frame as the footfall.
00:13.5 – 00:15 · key
Landing, unmask, the smile
He lands in a low crouch, one gloved hand on wet concrete, knees absorbing. He rises, grips the helmet with both hands and lifts it clear. Hair pushes free, damp. He looks out over the city and breaks into a real smile.
Watch for: the same face. Produced at second one, hidden for eleven seconds under a rigid shell, reproduced at second fourteen. Also the landing — models float the touchdown instead of absorbing it.
Which to use
They are not competing for the same job.
Neither model has been independently benchmarked long enough for a verdict to survive the next release cycle. What follows is a fit decision, not a ranking.
Reach for MiniMax H3
The deliverable fits inside fifteen seconds
You want 2K without an upscale pass
On-screen type is part of the brief
You are iterating heavily and need cheap failures
You want instruction-based edits instead of re-rolls
Compliance needs weights outside a vendor API
Reach for Seedance 2.5
The scene needs more than fifteen seconds to breathe
You want multiple shots inside one generation
You are stacking large reference packages for character consistency
Motion, clay-render or white-model references matter to your blocking
Timestamp-level editing removes a real revision bottleneck
You already work inside the Dreamina, Jimeng or CapCut stack
And the answer most production teams will actually land on is both. Seedance for the establishing scene, H3 for the inserts and coverage. They are not competitors in a pipeline. They are different tools that happened to launch on the same day.
Appendix
The full prompt, verbatim.
Copy it into both models with the same two reference images attached. If your results differ from ours, send them — we will add them to this page with credit.
elaris-stress-test-01.txt · 6,988 characters
ELARIS STRESS TEST 01 — "THWIP"
CHARACTER REFERENCE — MANDATORY: Two images: a pose sheet (proportions, suit construction, helmet design) and a portrait (face, hair, beard, skin tone). Face, suit, helmet and gloves match them EXACTLY — no restyling, no color changes, no added or removed panels. The references are the sole source of truth for appearance; everything below governs motion, camera, light and sound only.
ANATOMY — NON-NEGOTIABLE: South Asian man, late twenties, black wavy hair, full dark beard, lean and broad-shouldered, approx. 178cm / 75kg. NEVER ragdoll, limp or floppy. Core engaged, every movement has muscle intention. Moves like a trained traceur — precise, casual, effortless. Minimal spinning. The bare face at 0–3s and at 13.5–15s is the SAME face: same bone structure, beard and hairline. No drift.
THE SUIT — NON-NEGOTIABLE: Fitted armored bodysuit. Deep navy base, crimson armored panels across chest, shoulders, forearms and outer thighs. Segmented plating, visible seams, matte-satin sheen — real material, not spandex. Angular dark-blue emblem on the chest. Wide black belt, armored gloves, dark boots. Panel layout, emblem and color split stay IDENTICAL every frame — no morphing, no color shift.
THE HELMET — NON-NEGOTIABLE: Full-face rigid helmet, glossy crimson shell, dark smoked faceplate, fully enclosing. Hard shell, NOT fabric — no stretching or clinging. ON at 5s, OFF at 13.5s; in between it never slips and NO face, skin or hair is visible. Shell shape, visor outline and panel lines identical frame to frame.
HANDS — NON-NEGOTIABLE: Correct five-finger anatomy every frame — gripping the rigid helmet on and off, and in macro on the web-sign. WEB-SIGN, right hand, gloved: INDEX and LITTLE fingers extend straight up and spread apart; middle finger, ring finger AND thumb fold in and press flat to the palm. Exactly this, clean and readable. No merging fingers, no extra digits, no warping, no helmet clipping through the hands.
THE WEB — NON-NEGOTIABLE: Fires from the FINGERTIPS of the extended index and little fingers — TWO separate strands launching in parallel, braiding into one taut line a short distance out. Not the palm, not the wrist, not thin air. Real thickness and tension: sags at rest, snaps straight under load, anchors to a real surface. Never a glowing beam.
PHYSICAL CONTACT — NON-NEGOTIABLE: He does NOT float. He interacts with the city physically — feet push off walls, hands grab ledges, knees absorb impact, surfaces react.
ON-SCREEN TEXT — NON-NEGOTIABLE: A neon bodega sign reading "MARA'S DELI" in clean block lettering — sharp, correctly spelled, legible in every frame it appears. No warping or garbled letters.
CAMERA — NON-NEGOTIABLE: ONE continuous take. ZERO cuts. ZERO teleporting. It moves physically between close and wide — pushing, whipping, banking, rolling.
REALISM — NON-NEGOTIABLE: Real Hollywood film. RED Monstro 8K, anamorphic. Real motion blur, real armor physics, wet-street reflections, night haze.
15-second continuous single-take shot. Night. Rain-wet city, neon reflections in every puddle. Starts from black.
0–3s: Black. Then CLOSE — his bare face on a rooftop ledge, exact match to the portrait reference. Rain on his skin, neon across one cheek, breath visible. He looks out over the city, eyes moving, weighing it. Jaw sets. Camera eases back to reveal the helmet held in both gloved hands at chest height, real weight pulling his wrists down.
3–5s: He raises the helmet until it is level with his face and turns it once — the dark visor catches the skyline and throws it back in curved miniature. Then he lowers it over his head in one committed motion: fingers guiding the rim past his ears, hair pressing flat, the shell seating at the collar with a solid mechanical click. He rolls his neck to settle it. Rain runs down the gloss. Head lifts. Camera holds tight on the finished helmet one full beat.
5–6.5s: Camera drops and pushes in on his right hand — CLOSE. The gloved fingers set into the web-sign: index and little fingers snapping straight up and spreading wide, middle, ring and thumb folding in and pressing flat to the palm. Glove leather creases hard where the knuckles fold. The wrist cocks back. Held one clean beat, perfectly readable.
6.5–8s: MACRO — the two raised fingertips fill the frame, rain beading on the glove. The wrist snaps forward. TWIN web strands fire from the fingertips, launching in parallel and braiding into one taut line a short distance out. The recoil travels visibly back through fingers, wrist and forearm. The line goes tight. Camera whips out with it as he steps off the ledge and drops.
8–11s: WIDE — residential street below. Brownstones, fire escapes, the "MARA'S DELI" neon sign above a bodega door. He swings LOW, three meters up — arms extended gripping the web, legs together trailing, core tight, back straight. At the bottom of the arc his boots skim a parked car roof: real contact, real compression, feet bounce off and he redirects. He releases at the peak, lands feet-first on a brownstone wall and RUNS along it — boots striking brick in a real running gait, body angled 45°, arms pumping, four quick steps — then pushes off hard with both feet, knees fully extending, and fires a new web mid-air. Camera banks 40°, then rolls 60°.
11–12.5s: He comes in LOW over a narrow street of yellow taxis, headlights smearing on wet asphalt. The web releases. Boots hit wet pavement for two hard strides, real splash on each. A taxi honks. He laughs inside the helmet, plants a boot on a taxi, pushes off the roof in one explosive step, fires a web upward and rockets HIGH. "WOOOO!"
12.5–13.5s: WIDE — rooftop level. He hangs at the peak one breathless moment, arms slightly open, body arched, a small figure dead center against the burning skyline. Then he drops.
13.5–15s: He lands on a rooftop in a low crouch, one gloved hand on the wet concrete, knees absorbing the impact. He rises. Both hands grip the helmet and lift it clear — hair pushing free, damp, real. The SAME face. He looks out over the glowing city, chest still rising, and breaks into a wide genuine smile. City lights behind him.
[END]
EMOTION — NON-NEGOTIABLE: Fully alive in every frame, never blank. Opening: quiet, focused, a decision being made. Once masked the BODY carries it — shoulders, chin, spine, hands. Car skim: instinctive flinch, then recovery. The horn: helmet tips back, audible laugh. Final smile: unguarded, warm, genuine — NOT a posed grin.
AUDIO — SFX ONLY, NO MUSIC: Rain on concrete. One breath before the helmet. The solid click of it seating. Glove leather creaking as the fingers set. The sharp snap of the web firing, line twang under load. Wind rushing. Boots on car metal, then wet pavement. Taxi horn. His laugh, muffled by the helmet. "WOOOO!" between buildings. Boots on the rooftop, the helmet lifting free, one long breath out.
NO cuts. NO slow motion. NO spinning. NO floating. NO ragdoll. Photorealistic.
ElarisLabs is the AI-native creative operating system for brands. We publish the tests we run internally — including the ones where the model we expected to win didn't — because a comparison you can reproduce is worth more than a reel you can only admire.
One prompt→two models→eight axes→published either way