MiniMax H3 (community name: Hailuo 3.0) is the multimodal video model MiniMax released on 31 July 2026: 2K, 24fps, 4-15 second clips with natively synced audio and dialogue, supporting text-to-video, first/last-frame control, reference images and videos, motion transfer and generative video editing — up to 9 reference images, 3 video clips and 3 audio clips per request, priced from about $0.13 per second on the official API. This page collects copy-ready H3 prompts, each with a sample clip and credit to its author, free to use.
Source: MiniMax (the model vendor’s official site)
15 seconds, 16:9 dark-pop music video performance. Use @Image 1 as the strict identity reference for the three women — faces, hair and wardrobe unchanged in every shot. Use @Image 2 as the reference for the two titles only: ultra-condensed heavy caps, bone white, photocopied print surface.
The three women perform and sing to camera throughout. Follow the beats below shot by shot with continuous natural camera movement and real physical motion inside every frame — this is live performance footage, never a slideshow, never a sequence of held poses.
VOCAL: a single female lead voice, gritty and low, close-mic'd with tape saturation. Three short lyric lines are sung at 2.7s, 6.0s and 10.0s. Lip-sync locks precisely to those lines on whichever woman is singing. Between the lines all three mouth along, breathe hard and move.
[0.0–2.7s] The three stand in near-blackness, sand-beige studio crushed dark. Handheld camera drifting slowly forward, breathing. The platinum-crop woman rolls her shoulders and steps toward lens. Her hair lifts. The auburn woman in mirrored shades tilts her head up. Real movement, low energy, building. First title card over the darkness.
[2.7–6.0s] The auburn woman in the mirrored shades sings the first line straight into the lens, medium-close, mouth clearly forming the words. Camera whip-pans off her to the platinum-crop woman mid-stride, then crash-zooms back. Hard cuts every 0.3 seconds. Bodies in constant motion — hair whipping, chains swinging, hips driving, hands reaching for camera. Single-frame white flashes between cuts.
[6.0–10.0s] All three together, performing hard. The chestnut-haired woman takes the second line, singing directly to lens while the other two move behind her, bodies loose and driving. Camera orbits the trio fast, then whip-pans, then handheld snap-zooms. Cuts every 0.2 seconds. Hair, fur and fishnet all moving. Faces alive, mouths working, sweat catching the light.
[10.0–12.7s] Slow down hard, two long shots. The chestnut-haired woman alone, medium-close, singing the third line slowly and holding eye contact with the lens all the way through. Camera almost still, only breathing. Her wet hair moves across her face. She finishes the line and her mouth stays slightly open.
[12.7–15.0s] Everything detonates. Flash cutting, single-frame plates, all three thrashing and singing at once, faces close to lens, black frames and inverted negatives between cuts. Camera whipping violently. Final title card hammers in. Cut hard to black with the music.
MUSIC: driving industrial dark-pop, distorted 808 kick landing at 2.7 seconds, full arrangement through 6.0 to 10.0 seconds with clipped brass stabs and heavy tape saturation, dropping to near silence and one breath at 12.0 seconds for a full second, then the final two seconds at full weight with one metallic crash on the closing title. The lead vocal sits forward in the mix. Do not imitate an existing melody.
Warm sand-beige studio sweep, hard raking light from camera-left, Kodak Portra 400 grain, tape dropout, scanline tear, chromatic fringing, light leaks. Hard cuts only, full speed throughout, every shot carrying real motion. Titles spelled exactly as written, English only.
Reference prompts (2)
This clip was made by generating reference images with the prompts below, then driving the video with the prompt above.
参考图 1 · 三人组角色定妆
Render a 16:9 three-woman group production reference sheet.
SHEET — layout: single group photograph, three women standing together on one continuous seamless ground, full body head to sole, arranged left to right as a trio; the frame reads as one editorial group shot rather than a panelled sheet. Aspect 16:9 horizontal, 4K.
MEDIUM — 35mm film photograph on Kodak Portra 400, analog film response with gentle highlight halation, fine natural grain carried evenly across the whole frame. One photographed frame of three living actors on set, one stock, one grain structure.
SUBJECT 1 — VESPER: late-twenties woman of Eastern European build, tall and lean with wide shoulders. Bleached platinum crop cut short and shaggy, dark roots grown out, buzzed close at the nape. Pale grey-green eyes. Markers: a fine vertical scar splitting the left eyebrow, a dense freckle scatter across the nose bridge, a healed empty piercing hole in the right nostril. Wardrobe: cropped shaggy sheepskin coat in bone and toffee worn open off one shoulder, silver chainmail mesh halter beneath, low-slung baggy carpenter denim in washed sand with a studded belt hanging loose, fingerless black leather gloves, platform moto boots in scuffed black leather. Mood: combative stillness, chin lifted, soft closed-lip expression. Standing at full height on the left third, frontal, weight dropped onto the right hip.
SUBJECT 2 — MARLOW: early-thirties woman of Mediterranean colouring, slender with long limbs. Dark auburn hair past the shoulder blades, half-wet and raked straight back. Hazel eyes behind narrow wraparound mirrored shield sunglasses. Markers: a faint pale scar along the right jawline, a small dark mole beside the left corner of the mouth. Wardrobe: oversized black leather bomber with a heavy grey wolf-fur collar pushed off both shoulders, white ribbed cotton tank cropped raw above the navel, layered steel chain harness, low-rise bootcut denim in near-black wash with laced open panels down the outer thigh, chunky black buckled boots. Mood: cold unbothered composure. Crouched low in the centre foreground, three-quarter to camera-right, forearms across the knees, one hand raised to the sunglasses temple.
SUBJECT 3 — SABLE: mid-twenties woman of Irish colouring, compact and athletic. Long dark chestnut hair to the waistline, heavily tousled and swept forward across the left cheek. Dark brown eyes with a heavy upper lid, very fair skin with a live flush. Markers: a visible gap between the two upper front teeth, a pale scar high on the right temple, a cluster of small dark moles down the left side of the neck. Wardrobe: rust-orange satin corset top boned and front-laced, shaggy Mongolian fur stole in cream falling off both shoulders, asymmetric brown leather wrap mini skirt, torn black fishnet tights, cream faux-fur leg warmers, spiked platform heels in oxblood leather. Mood: coiled predatory readiness, lips just parted. Leaning in from the right third, bent forward at the waist toward frame centre, three-quarter to camera-left.
SCENE — the trio holds one group pose as a unit: Vesper upright and open on the left, Marlow low and central pulling the eye down, Sable folding in from the right to close the triangle; the three silhouettes overlap slightly at the edges so the group reads as one mass. Location: an empty seamless studio cyclorama, warm sand-beige sweep running floor into wall.
LIGHT — key: 5400K single very large studio softbox raking from camera-left at 40 degrees, soft-edged, carving all three sets of cheekbones with one shadow direction. Fill: broad white bounce card camera-right at 3:1, cooler by 400K. Rim: narrow strip light behind camera-right separating platinum crop, wolf-fur collar and chestnut hair from the sweep. Material response: sheepskin and Mongolian fur breaking into bright curling fringes, chainmail throwing hard specular pinpoints, satin sweeping a long soft highlight along each crease, mirrored lenses returning a flat hard reflection of the softbox, fishnet casting a fine shadow lattice on fair skin. Firm contact shadows where every boot meets the floor, twin catchlights in every visible eye.
LOOK — warm analog grade led by sand-beige, bone and rust against the cool blacks of leather and denim, contrast held gentle, shadows lifted to a milky charcoal, highlights rolling off softly. Hasselblad H6D-100c, 80mm, f/8 with all three figures crisp from crown to sole. 35mm film still, Kodak Portra 400, fine natural film grain, matte fibre texture, late-1990s fashion editorial.
CONSISTENCY LOCK — three distinct women photographed in one frame, each keeping her own facial geometry, eye colour, hair colour and length, body proportion and complete wardrobe; all three under one continuous light recipe on one continuous ground with one film stock and one grain structure.
参考图 2 · 标题字体系统板
Render a 16:9 graphic design system board — a single flat printed reference sheet photographed head-on, filling the entire frame edge to edge.
LAYOUT — ten numbered panels in an asymmetric grid on a deep black ground, thin hairline rules dividing them, each panel carrying a small sans-serif index number in its upper-left corner with a short descriptive caption beside it.
Panel 01, upper-left, wide and tall — an ultra-condensed heavy all-caps title in bone white, letters stretched tall and pressed narrow, filling the panel in two stacked lines. A hairline vertical rule runs down the far left edge carrying tiny rotated type.
Panel 02, upper-centre-left, narrow vertical strip — two heavy condensed words set sideways, reading bottom to top, bone white on black, filling the full column height, with small type at the base.
Panel 03, upper-centre, wide — two heavy condensed words arcing outward in perspective, letters fanning from a vanishing point below the baseline so the type rushes toward the viewer, with a thin measurement scale and micro-type beneath.
Panel 04, upper-right, wide — a bone-white bordered box on black holding a single heavy condensed word, with a technical strip beneath carrying a small wireframe globe glyph, a row of checker squares, a numeral and a small burst mark, plus a line of micro-type below.
Panel 05, far-right narrow vertical strip — a tall machine-readable barcode in bone white running the full column height, with rotated type alongside.
Panel 06, lower-left, wide — two enormous heavy condensed words acting as a cut-out window, a grainy high-contrast black-and-white landscape photograph visible inside the letterforms, with micro-type beneath.
Panel 07, lower-centre — a solid oxblood-red block filling the panel carrying two bone-white heavy condensed words stacked on two lines, with a thin audio waveform strip and micro-type beneath.
Panel 08, lower-right, wide — a heavy condensed word standing in perspective on a receding wireframe grid, a second smaller ghosted instance behind it in dark grey, small labels pointing to each depth layer.
Panel 09, bottom-left, wide and short — a block of small sans-serif body copy in bone white on black arranged in three columns under short headings, with six small bordered chips beneath one heading and a boxed line at the left.
Panel 10, bottom-right, wide and short — six square texture swatches in a row showing coarse film grain, halftone dot pattern, scratched emulsion, dust and hair speckle, torn paper edge and scan misregistration streaking, with micro-labels beneath and a small oxblood-red block at the far corner.
TYPOGRAPHY: every title set in one family of ultra-condensed heavy all-caps sans-serif with tight tracking, bone white on black or bone white on oxblood red. Captions, index numbers and micro-type in a plain small sans-serif, uppercase, widely tracked. All text English, correctly spelled, each glyph holding its exact designed shape.
PALETTE: deep black ground, bone white type, one aggressive oxblood red used only in Panel 07 and the corner block of Panel 10.
SURFACE AND TREATMENT: the whole board reads as a physical printed page that has been photocopied and scanned — coarse film grain across every panel, halftone dot structure inside the large letterforms, rough inked print edges, faint scan misregistration offsetting the red channel by a hair, light scratches and dust speckle, paper worn and slightly torn along the outer frame border. Flat even illumination with the faintest vignette in the corners. Shot square-on, every panel equally sharp, late-1990s underground music-poster and zine design language.
Super casual real smartphone home video footage, cozy winter cabin gathering with snow visible outside, natural mobile phone camera with slight authentic handheld shake, normal frame rate with smooth natural motion, rapidfire montage with quick jump cuts every 1-2 seconds like scrolling through phone memories, unpolished authentic phone recording of mixed ages sipping hot cocoa, chatting by the fireplace, pure raw home video feel, no cinematic polish.
Use the provided reference photo as the strict ONLY visual reference for the main woman. Maintain her exact appearance with zero deviation. Generate a mixed group of family and friends around her inside the cabin, fireplace and snow visible through the window.
0-2.5s: Shaky rapid cuts — main woman laughing near the fireplace, warm cocoa mug in hand, quick flashes of snow falling outside.
2.5-5s: Abrupt jump cuts — close-up of her smiling while sipping cocoa, then friends passing blankets and snacks around.
5-7.5s: Fast shaky — she chats animatedly by the fire, friends and family cozy on the couch nearby.
7.5-10s: Quick cut close-up — warm smile toward camera, then jump to her laughing with a grandparent by the fireplace.
10-12.5s: Abrupt edit — group gathered close, sharing snacks and stories, casual toast with mugs.
12.5-15s: Final rapid transition — main woman relaxed by the fire with family, soft smile, calm winter memory ending with gentle phone sway.
Natural smartphone video quality, slight real handheld shake, smooth normal frame rate motion, authentic casual physics, stable main character consistency, no pro stabilization or effects.
Use @Image1 as the cloaked avenger identity lock and wardrobe reference. Use @Image2 as the environment, lighting, and street layout reference. Use @Image3 as the cyborg executioner identity lock, armor design, and weapon reference.
Create a 15-second ultra-cinematic confrontation in the ruined streets of Sector 7 at night. Heavy ashfall, wet pavement, orange emergency lights, smoke columns, and reflective industrial architecture strictly follow @Image2. The cloaked avenger from @Image1 faces the red cyborg executioner from @Image3 in the center of the empty street. Maintain facial features, wardrobe, armor geometry, and environment consistency across all shots. No onscreen text, no floating letters, clean visuals.
0–4s: Wide establishing shot. The cloaked avenger from @Image1 walks slowly through the rain-dark street toward the cyborg from @Image3. Camera performs a slow push-in. Coat edges move in the wind, boots splash through shallow water, ash drifts through orange backlight.
4–8s: Medium close-up on @Image1. He stops, lifts his chin slightly, and says clearly in English: “Long time I was looking for you.” Deliver slowly, low, and intense, with precise lip-sync. Camera holds a subtle handheld drift. Wet hood, tired eyes, and grim determination remain consistent with @Image1.
8–11s: Medium close-up on @Image3. The cyborg tilts its head, red visor brightening, and replies in a cold metallic voice: “finally”. Deliver with a short pause before the word and precise mouth-panel synchronization. Servo motion, red eye glow, and armored posture strictly follow @Image3.
11–15s: Sudden action payoff. Whip pan into a dynamic wide shot as @Image3 ignites the red energy blade and lunges forward. @Image1 instantly draws his blade and rushes to meet the attack. They collide at center frame in a shower of sparks, rain spray, and reflected orange-red light. End on a dramatic frozen impact moment with crossed blades, smoke and ash swirling around them, camera orbiting slightly into a powerful hero frame. Gritty cinematic realism, strong contrast, volumetric haze, realistic body momentum, normal proportions without stretch.
Use [Ref_Image1] as the strict character identity, design, wardrobe, color and multi-angle reference. Its six views depict the same traveler. Reconstruct one coherent 3D character without averaging or redesigning his face. Preserve exactly the crescent-shaped face, amber eyes, midnight-blue hair, proportions, mustard-yellow suit, ivory shirt, teal scarf, burgundy shoes and floating black bowler hat shown in [Ref_Image1] throughout the video.
👤 CHARACTER
The original stylized cartoon traveler from [Ref_Image1] has a tall slender silhouette, expressive crescent-shaped face and large curious amber eyes. Keep his identity, proportions, colors and clothing perfectly consistent.
🌌 WORLD
A whimsical Dalí-inspired surreal desert at golden hour, rendered as a premium 3D cartoon. Vast peach-colored sand, an impossibly curved turquoise horizon, long-legged houses walking slowly in the distance, floating doorframes, liquid staircases, stretched shadows, cloud-shaped fish and soft clock-like flowers bending in the breeze. Original imagery, not a recreation of any existing artwork.
🎬 15-SECOND ACTION
0–2s: A shadow peels itself away from the traveler’s feet, bows to him and becomes a moving staircase. He reacts with delighted surprise.
2–5s: The traveler steps onto the staircase as it glides across the desert; floating doors open onto tiny upside-down oceans.
5–9s: He returns to the sand and walks beneath extremely tall houses whose windows blink like sleepy eyes. His floating hat circles his head once.
9–12s: Gravity tilts; the horizon curves upward and he walks sideways while cloud-fish swim overhead.
12–15s: He opens a freestanding door and discovers the opening shot inside. He steps through as it closes, creating a seamless loop.
📷 CAMERA
Smooth side-tracking camera with subtle dolly movement, one continuous shot, no hard cuts, strong layered parallax, 35mm cinematic perspective and stable character scale.
✨ STYLE
Elegant surrealist cartoon, rounded expressive shapes, hand-painted textures, premium stylized 3D animation, playful impossible physics, warm peach, mustard, teal and burgundy palette, cinematic golden light, gentle film grain, 24fps.
🎵 MUSIC AND SOUND
Original instrumental LoFi: mellow vinyl crackle, dusty boom-bap drums, warm Rhodes chords, soft upright bass and dreamy reversed guitar at 78 BPM. Synchronize the shadow, doors, orbiting hat and final closure to the beat. No vocals or recognizable melody.
🔒 CONSISTENCY
Exactly one traveler. No identity or wardrobe changes, extra limbs, malformed hands, duplicate character, text, logos, watermark, flicker, abrupt cuts or photorealistic elements.
[FORMAT] Exactly 15 seconds, horizontal 16:9, photorealistic AAA fantasy MMORPG gameplay reveal with native synchronized game audio and music.
[OMNI REFERENCES]
[Image1] = Kael Ardyn, the exact playable arcane spellblade. Preserve his face, silver-black tied hair, pointed ears, cyan eyes, athletic proportions, navy-and-silver armor, gold filigree, teal cloth, chest focus and curved spellblade.
[Image2] = The exact Obsidian Tempest Dragon. Preserve its enormous four-legged anatomy, black volcanic-glass scales, crown horns, torn wings, cyan electrical fissures, storm-blue eyes and long spiked tail.
[Image3] = The strict Skybreak Citadel arena and spatial reference: circular black-stone platform, gold-and-cyan inlays, ruined towers, floating architecture, storm abyss and blue-hour lighting.
[Image4] = The strict visual reference for the 2026 gameplay interface: refined gold-and-cyan frames, boss bar, circular minimap, quest tracker, party frames, player status arcs, action bar, cooldown numbers and restrained typography.
Reconstruct each multi-view subject as one coherent 3D asset without averaging or redesigning it.
[DESIGN GOAL] Original next-generation fantasy MMORPG with cinematic AAA presentation and genuinely playable movement. The 2026 HUD uses elegant smoked glass, contextual information, subtle transitions, responsive cooldowns and excellent readability. It remains locked to the screen, never floating inside the world.
[15-SECOND SEQUENCE] 0–2s: Immediate hook with no HUD. Low camera behind Kael. The Obsidian Tempest Dragon crashes onto the platform ahead, wings filling frame, lightning splitting the sky and stone fragments rising from impact. Kael draws his spellblade as it roars. 2–4s: Camera settles into polished third-person gameplay. The interface assembles around the edges: top boss bar reading “OBSIDIAN TEMPEST,” player health, mana and stamina below, action bar, party frames left, minimap upper right and quest tracker reading “THE SKYBREAK SIEGE.” HUD stays sharp and screen-locked. 4–8s: Kael sprints forward and chains three abilities: cyan blade dash, circular gold runic shield and vertical storm strike into the dragon’s chest. Matching action icons illuminate once and begin radial cooldowns with changing numbers. Boss health falls from 100% to 78%. Accurate footwork, weapon contact and restrained camera shake. 8–11s: The dragon releases a ground-level electrical shockwave. A red telegraph ring appears on the arena. Kael performs a perfectly timed sideways dodge; a small centered “PERFECT EVADE” notification appears briefly and fades. Health stays stable while stamina dips and recovers. 11–13s: Kael targets the glowing chest fissure and unleashes a cyan-gold arcane slash. Brief hit-stop, one damage number, boss bar falls to 12%, then sparks and debris resume. 13–15s: Kael lands in a hero stance as the dragon recoils. HUD minimizes toward the edges. Centered title: “AETHERFALL ONLINE,” then “THE RAID BEGINS.” End on thunder and musical impact.
[CAMERA + VISUAL QUALITY] Seamless cinematic-to-gameplay transition, realistic third-person controls, premium real-time rendering, ray-traced reflections, volumetric storm light, grounded combat, coherent scale, controlled motion blur and cyan-gold contrast. Preserve the arena layout; avoid arbitrary cuts.
[UI BEHAVIOR] Anchor all UI to exact screen coordinates with consistent scale and typography. No flicker, duplication or drifting. Values change only in response to described actions. Never cover Kael or the dragon.
[AUDIO] Original hybrid fantasy score: war drums, low strings, metallic pulses and restrained electronic rhythm. Add dragon impact and roar, thunder, debris, footsteps, sword draw, arcane abilities, shield resonance, shockwave, dodge and final hit. Keep effects clear. No narrator, dialogue, vocals or recognizable melody.
[PRESERVE] Stable Kael and dragon identities, exact armor and creature design, one hero, one dragon, Skybreak Citadel geography, responsive 2026 HUD, readable specified labels and coherent combat causality.
[AVOID] Existing game logos or recognizable franchise symbols, copied interface layouts, extra heroes or dragons, turn-based combat, cinematic black bars, first-person view, random UI changes, unreadable text, floating HUD panels inside the world, excessive particles, cartoon rendering, anatomy drift, weapon changes, clipping, subtitles, additional text or watermark.
🖼️ OMNI REFERENCE — [Ref_Image1]
Use [Ref_Image1] as the strict character identity, wardrobe and multi-angle reference. Preserve Freya’s exact face, blue eyes, long platinum-blonde hair with darker roots, skin tone, age and body proportions from [Ref_Image1]. Keep the same buttoned white linen shirt, cream linen shorts, bare feet and gold bracelet shown in [Ref_Image1]. Its six views represent one person from different angles; reconstruct one coherent 3D character without averaging or redesigning her face.
🏝️ SCENE
Freya stands near the waterline on a quiet Caribbean beach in warm late-afternoon sunlight: pale sand, transparent turquoise sea, distant palms and blue sky. A gentle breeze moves her hair and linen shirt while small waves reach the shore.
🎥 15-SECOND 360° ORBIT
One uninterrupted shot with no cuts. A stabilized camera moves clockwise around stationary Freya in one complete, physically correct 360° orbit.
0–2s: Start medium-wide in a frontal three-quarter view; she makes eye contact, smiles and breathes naturally as the orbit begins immediately.
2–5s: Pass her left profile with strong shoreline parallax; she begins speaking: “If paradise had a heartbeat…”
5–8s: Reach the full back view; she turns only her head slightly toward the moving camera and continues: “…I think it would sound…”
8–12s: Pass her right profile; she looks briefly at the horizon, returns her gaze to the lens and finishes: “…exactly like this.” She smiles naturally.
12–15s: Complete the orbit and return precisely to the opening frontal composition for a seamless loop.
📷 CAMERA
Constant clockwise direction, orbital radius, horizon and subject scale. Smooth gimbal motion, 35mm cinematic lens, full-body framing from head to bare feet, realistic foreground/background parallax. The camera travels around Freya; she does not spin and the background does not rotate artificially.
✨ LOOK
Premium photorealistic travel-fashion film, natural skin and linen texture, warm sunlight, soft shadows.
1One action per prompt. H3 clips are 4-15 seconds. "She pushes the door open and glances back with a smile" beats "she opens the door, crosses the hall, sits down and starts typing" — the second gets compressed into mush.
2Spell out the camera move. Dolly in, pull back, pan, tracking shot, handheld drift, drone descent. Camera language is the biggest difference between video and image prompts; leave it out and you get a locked-off tripod.
3Write the sound — H3 generates it. H3 produces synced native audio. Name the ambience (rain, subway announcement), the action sound (heels on tile), even a line of dialogue, and the clip lands a whole tier more finished.
4Set light and palette first. "Cold blue night with neon spill" or "warm golden-hour backlight" — one lighting/color sentence decides the whole look, exactly as in image prompts.
5Give it speed and timing. Slow motion, real time, timelapse — and when the beat happens. Without a rhythm cue the model spreads the action evenly and it reads sluggish.
6References carry identity, the prompt carries action. H3 takes up to 9 reference images. Let them hold the face, wardrobe or product; keep the prompt focused on what happens and how the camera moves.
7State aspect ratio and the final beat. 16:9 for landscape, 9:16 for short-form. Adding where the shot ends ("hold on a close-up of the product") saves you an edit later.
Prompt formula
Subject + specific action + camera move + lighting & palette + environment detail + sound/dialogue + pacing + aspect ratio
Example: Outside an izakaya on a rainy night, a woman in a beige trench closes her umbrella and turns; the camera dollies slowly in from over her shoulder to a profile close-up. Cold blue streetlight mixed with warm orange lantern glow, rain beading on the umbrella. Rain ambience and a distant train rumble; she says quietly, "You are late." Real time, 16:9.
Frequently asked questions
What is MiniMax H3?
MiniMax H3 (commonly called Hailuo 3.0) is the multimodal video model MiniMax released on 31 July 2026. It outputs 2K video at 24fps, 4-15 seconds long, with natively generated synced audio and dialogue. It supports text-to-video, first/last-frame control, reference images and videos, motion transfer and generative video editing — up to 9 reference images, 3 video clips and 3 audio clips in one request.
Are these prompts free?
Yes. Every prompt here is free to copy and run anywhere you have access to H3 — no login, no payment.
Who wrote these prompts?
Each card credits the prompt author, and the credit links to the original post or profile. We only curate; copyright stays with the author. If you are an author and want a prompt removed or a credit corrected, email us.
Where can I run these prompts?
H3 is live on apimodels.app as the model minimax-h3 — native 2K with synchronized audio, any whole length from 5 to 15 seconds, and up to 9 reference images plus 3 reference videos and 3 audio clips in one request. It is billed per second at $0.145/s (5 seconds is $0.725), the first 5 reference images are free and each extra image is $0.03, and only successful requests are charged. The same key and the same endpoint also cover Seedance, Kling, VEO and Grok. You can also run H3 on the official MiniMax API or in the Hailuo product.
How do I write a good H3 prompt?
Write each prompt like a single storyboard card: what the subject does, how the camera moves, the lighting and palette, the environment details — then the sound you want, since H3 generates native audio and dialogue. Keep it to what fits in 4-15 seconds; one clear action beats three crammed beats.
Can I use H3 videos commercially?
That is governed by MiniMax's terms of use — check the official terms. The prompt text collected here is yours to edit and remix freely.
Run these prompts on apimodels.app
MiniMax H3 is live here alongside 85+ other image, video, audio and language models — one API, one key, pay as you go, $1 free on sign-up.
The prompts we reconstructed are published in full under MIT; author-written prompts are indexed with credit and a link to the original post. Open an issue to suggest more.