
minimax-h3MiniMax H3 is a general-purpose multimodal video model rather than a stack of single-purpose tools. It reads text, images, video and audio in the same pass, works out what the brief is actually asking for, and returns one coherent clip with picture, motion and sound already in agreement — no separate dubbing, lip-sync or motion-transfer step to stitch together afterwards. Practically, that means one endpoint covers the whole range: send a prompt on its own for text-to-video, or attach up to 9 reference images, 3 reference videos and 3 reference audio clips to pin down character, camera movement and voice at once. Every reference input is optional. Choose any whole-second length from 5 to 15 and an aspect ratio of adaptive, 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16. Billed per second at $0.145/s (5s = $0.725, 15s = $2.175); the first 5 reference images are free and each additional image is $0.03. Only successful generations are charged.
At $0.145/s it costs 29.5% of our own Seedance 2.0 1080p tier ($0.492/s) — and outputs native 2K rather than 1080p. Both figures are published on this site, so you can check the comparison yourself
Text, image, video and audio are understood together rather than by separate specialist tools — visuals, motion and sound arrive already in sync, in a single pass
Every generation renders at 2K with synchronized audio — no upscaling step and no lower-resolution tier to pick between
Prompt alone behaves as text-to-video; add images, videos or audio to steer character, camera motion and voice in the same request
Whole-second control rather than fixed buckets, billed per second at $0.145/s so you pay for exactly the length you asked for
Up to 9 reference images, 3 reference videos and 3 reference audio clips — the first 5 images are free, extras are $0.03 each
Leave the references below empty for pure text-to-video — every reference input is optional.
Steer character, style or composition. The first 5 are free; each image past the 5th adds $0.03.
MP4 / MOV, ≤50MB and 2-15s each. Use these to carry camera movement or motion style into the output.
MP3 / WAV, ≤15MB and 2-15s each. Used to match voice or drive lip movement.
Generated video will appear here
Provide URLs and click Generate
MiniMax Hailuo H3 is a Video Generation API provided by Minimax. MiniMax H3 is a general-purpose multimodal video model rather than a stack of single-purpose tools. It reads text, images, video and audio in the same pass, works out what the brief is actually asking for, and returns one coherent clip with picture, motion and sound already in agreement — no separate dubbing, lip-sync or motion-transfer step to stitch together afterwards. Practically, that means one endpoint covers the whole range: send a prompt on its own for text-to-video, or attach up to 9 reference images, 3 reference videos and 3 reference audio clips to pin down character, camera movement and voice at once. Every reference input is optional. Choose any whole-second length from 5 to 15 and an aspect ratio of adaptive, 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16. Billed per second at $0.145/s (5s = $0.725, 15s = $2.175); the first 5 reference images are free and each additional image is $0.03. Only successful generations are charged. Through APIMODELS platform, you can access this model via a unified API with transparent pay-as-you-go pricing. Current pricing: 2K · per second: $0.145, 2K · 5s: $0.725, 2K · 10s: $1.45, 2K · 15s: $2.175, Extra reference image (past 5): $0.03.
Quickly generate brand promotion videos for ad campaigns and social media marketing.
Create compelling short-form video content for platforms like TikTok, Instagram, and YouTube.
Generate product feature demonstrations and tutorials to improve user conversion.
Produce course explanations, knowledge explainers, and training videos at low cost.
MiniMax Hailuo H3 is available through APIMODELS at: 2K · per second: $0.145, 2K · 5s: $0.725, 2K · 10s: $1.45, 2K · 15s: $2.175, Extra reference image (past 5): $0.03. Billing is pay-as-you-go — you only pay for what you generate.
Sign up at APIMODELS, get your API key, and call our unified API endpoint. We provide detailed API documentation with code examples in cURL, Python, and Node.js.
APIMODELS offers the same MiniMax Hailuo H3 model through our aggregation platform. We provide a unified API interface so you do not need separate accounts for each provider - one API key to access all models.
MiniMax Hailuo H3 is MiniMax's multimodal video generation model. A single request produces native 2K footage together with synchronized audio, so there is no separate dubbing step. What sets it apart from most video models is that one endpoint covers every mode: reference images, reference videos and reference audio are all optional inputs. Send a prompt on its own for text-to-video, or attach media to steer character, camera motion or voice — you never switch endpoints.
On APIMODELS it is billed per second at $0.145/s: 5 seconds is $0.725, 10 seconds is $1.45 and 15 seconds is $2.175. The first 5 reference images are free and each image past that is $0.03 (9 max, so at most +$0.12); reference videos and audio cost nothing extra. There is no minimum top-up and no subscription, and only successful requests are charged — a failed generation costs you nothing.
Use the unified video endpoint POST /api/v1/video/generations with model set to minimax-h3 (the alias hailuo-h3 also works). prompt and duration (a whole number of seconds, 5-15) are required. Attach references as needed: images (up to 9), video_list (up to 3) and audio_list (up to 3), each accepting a public URL or base64. Pass ratio to pin the aspect ratio. The call returns a taskId — poll GET /api/v1/video/generations?task_id=xxx, or supply callback_url at creation and we will POST you the result when it finishes. One API key covers every model on the site, reachable from mainland China with no organization verification.
Resolution is fixed at native 2K — this channel offers no lower tier. Duration is any whole second from 5 to 15 rather than fixed buckets, and the aspect ratio can be adaptive, 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16. Reference limits: up to 9 images (≤30MB each), up to 3 videos (MP4/MOV, ≤50MB and 2-15s each) and up to 3 audio clips (MP3/WAV, ≤15MB and 2-15s each). Prompts can be up to 20,480 characters. We transfer the finished video to our own storage and the link stays live for 7 days, so download or re-host it before it expires.
It is the right pick when you need native 2K, synchronized audio and multimodal references at the same time: continuous shots that must keep a character consistent, ad footage that has to reproduce a specific camera move, talking-head clips with a chosen voice or lip-sync, and complex briefs that mix image, video and audio references. On value it holds up well: $0.145/s is 29.5% of our own Seedance 2.0 1080p tier ($0.492/s), and it returns native 2K rather than 1080p — both prices are published on this site, so the comparison is yours to check. If you mainly need fast vertical clips on a tight budget, LTX-2.3 (from $0.02/s at 480p) or Seedance 2.0 Fast (from $0.11/s) still go further per dollar; if you need 4K or video editing, look at Omni Flash. On APIMODELS all of these share one API key and the same endpoint, so switching models means changing a single field.
On APIMODELS, MiniMax Hailuo H3 runs alongside 60+ models on one API key and one balance, so choosing is about fit, not lock-in. It supports Native 2K、Multimodal Reference、5-15s、Up to 9 Ref Images、3 Ref Videos + 3 Audios、Synchronized Audio, and you can weigh it on price and capability against other Video Generation models, then switch by changing a single model-name string — no new account or integration. Browse every Video Generation option with live pricing at apimodels.app/models.
MiniMax Hailuo H3 supports: Native 2K、Multimodal Reference、5-15s、Up to 9 Ref Images、3 Ref Videos + 3 Audios、Synchronized Audio. See the APIMODELS docs for full parameters and call examples.
Yes. APIMODELS exposes MiniMax Hailuo H3 through a single unified API and one key — no separate provider accounts, and no need to handle each provider's regional network access yourself.
We support Stripe (Visa, Mastercard, and other international cards) and Alipay. Credits are available instantly after payment.
Prompts shared by their authors — copy and adapt them. Each one credits its author and links back to the original post.
Condor Heroes characters teach English word dream
神雕侠侣主角趣味讲单词 dream 教程
by @nicekate8888
Pizza night UGC Domino’s vlog
VIDEO PROMPT — "Pizza Night Vlog" (UGC iPhone Style) Duration: 15 seconds | Aspect Ratio: 16:9 | Style: Authentic UGC / iPhone selfie-vlog, handheld, natural light, slight motion blur, TikTok/Reels energy — NOT cinematic, NOT overly polished. Feels like a real creator filmed this on their phone. Product Reference: Use the uploaded Domino's Pepperoni Pizza image as the only product reference. Keep crust thickness, cheese texture, bake color, pepperoni placement, and proportions identical in every cut — no redesigning the pizza. Camera: iPhone 15 Pro front + back camera switching, handheld, natural wobble, autofocus hunting slightly (realistic), vertical-style framing cropped to 16:9, occasional finger near lens edge, natural room lighting + phone flash reflections on the pizza box. Character Description Name (for reference): Mia Awoman in her mid-20s, naturally attractive and beautiful with an approachable, girl-next-door charm — not overly done up. Wavy sandy-blonde hair pulled back loosely, light natural makeup, wearing a cozy oversized cream sweater. Warm, genuine smile, expressive eyes, casual energetic personality like a real lifestyle vlogger. Sits in a softly lit modern kitchen/living room. Shot Breakdown SHOT 1 (0–2s) — The Grab Selfie-angle, she's mid-laugh holding up the Domino's box to camera. Quick jump cut. Dialogue: "Okay so it's officially pizza night—" SHOT 2 (2–4s) — The Open Cut to overhead handheld shot, box flips open, steam rising off the pizza, slight camera shake as she leans in. SHOT 3 (4–6s) — The Zoom Quick zoom-punch into the pizza, phone camera autofocus adjusts naturally, cheese and pepperoni in focus, ambient kitchen sounds. SHOT 4 (6–8s) — The Pull Cut to her hands lifting a slice, natural cheese pull, filmed from a slightly low candid angle like a friend filming across the table. SHOT 5 (8–10s) — The Reaction Cut back to selfie-cam, she takes a bite, eyes widen, quick genuine reaction. Dialogue: "Oh my god, that's so good." SHOT 6 (10–12s) — The Candid Cutaway Jump cut to a close, slightly shaky shot of the pizza box on the counter, her hand grabbing another slice off-frame, casual b-roll energy. SHOT 7 (12–14s) — The Wrap-Up Back to selfie angle, she grins at camera, holding slice up like a toast. Dialogue: "Dominos, y'all know what to do." SHOT 8 (14–15s) — End Tag Quick freeze/cut to the box logo close-up, natural handheld wobble, soft text overlay in casual font: "pizza night = solved 🍕" — cut to black. Look & Feel Warm indoor lighting, slightly grainy natural phone sensor look, imperfect framing, real reactions, minimal dialogue (3 short lines total), authentic pacing with hard jump cuts instead of smooth transitions. Negative Prompt cinematic grade, overly smooth camera moves, studio lighting, professional voiceover, staged acting, CGI look, plastic cheese, distorted pepperoni, extra fingers, warped hands, text glitches, logo distortion, overly polished commercial feel.
by @ShamiWeb3
The World's Unluckiest Superhero
A documentary about a superhero who has extremely bad luck and ends up saving people by accident through the destruction caused by his own misfortune. Dialogue in English. Scene direction with unique composition. Every cut, every camera angle, and every movement is of exceptionally high quality; the composition is guided by an experienced film director. The comedy is genuinely interesting, and throughout the 15 seconds everything unfolds in a perfectly crafted way, with a comedic payoff that can make anyone laugh.
by @NACHOS2D_
We curate copy-ready prompt libraries — every entry shows its full text and a sample result, ready to adapt.
How to get access, regional availability, and how this model compares with its alternatives.