
minimax-h3-liteMiniMax H3 Lite is the budget tier of the MiniMax H3 multimodal video model: the same H3 weights as minimax-h3, on a cheaper channel, with a hard ceiling of 768P. It costs a fifth of the standard tier — 768P at $0.02 per second versus $0.097/s on minimax-h3, and 480P at $0.01/s — so a 768P 10-second clip is $0.20 instead of $0.97. What you give up is explicit: no native 2K and no 1080P (768P is the maximum), no reference video (motion transfer), and only two aspect ratios, 16:9 or 9:16. Everything else carries over: any whole length from 1 to 15 seconds, up to 9 reference images to lock character or product, up to 3 reference audio clips to drive voice and lip movement, or a first and last frame for an interpolated shot; references are free. It is not a fast model — plan on 5 to 10 minutes per clip for 10-15 second requests, and a 15-second clip can run longer when the queue is busy; clips of 5 seconds or less usually finish in 1-3 minutes. For comparison, minimax-h3 on our production traffic over the last 14 days had a median turnaround of 6.4 minutes and 90% finished within about 10 minutes, so the two tiers are similar in speed — the difference is price and ceiling, not latency. The limit that matters most is not on the spec sheet: motion amplitude. Lite holds together on animation and on shots where the subject moves a little — talking heads, slow pans, product turns, looping backgrounds, subtle expression changes. On live-action footage with large movement — running, fighting, dancing, fast camera moves — faces and limbs deform and characters lose their shape. That is a property of this tier, not a prompt problem, and no wording of the prompt fixes it. Treat Lite output as a draft: use it to lock composition, timing and framing cheaply, then re-shoot the approved cut on a higher tier if the delivery has to survive large movement. Pick Lite for animation, small-motion shots, drafts, social verticals, iteration and volume; pick minimax-h3 when the delivery needs 2K or has to reproduce the camera move of a reference video. Only successful generations are charged.
Leave the references below empty for pure text-to-video — every reference input is optional.
Steer character, style or composition. Reference images are free on Lite.
MP3 / WAV, ≤15MB and 2-15s each. Drives voice and lip movement; needs at least one reference image. Free.
Generated video will appear here
Provide URLs and click Generate
Animation, talking heads, slow pans, product turns, looping backgrounds, subtle expression changes — this is where Lite holds together. Motion amplitude, not resolution, is the real ceiling on this tier
On live-action with running, fighting, dancing or fast camera moves, faces and limbs deform and characters lose their shape. It is a property of this tier, not a prompt problem — no prompt wording fixes it. Use Lite to lock composition and timing cheaply, then re-shoot the approved cut on a higher tier
768P at $0.02/s versus $0.097/s on minimax-h3 — a 10-second 768P clip is $0.20 instead of $0.97. Character consistency and prompt-following come from the same model
No native 2K and no 1080P on this tier — 480P or 768P only, in 16:9 or 9:16. If the delivery needs 2K, use minimax-h3
Async by design: 10-15 second clips take 5-10 minutes, a 15-second clip can run longer under load, clips of 5 seconds or less usually 1-3 minutes. About the same as minimax-h3 (median 6.4 min on our traffic) — the saving is money, not time
Prompt alone is text-to-video; add reference images, reference audio or a first + last frame and the request is routed to the matching pipeline — same field names as minimax-h3, so switching tiers means changing model only
Up to 9 reference images and 3 reference audio clips at no extra charge; reference video (motion transfer) is not supported here — sending one returns a clear 400 rather than a silently ignored file
MiniMax H3 Lite은(는) MiniMax의 영상 생성 API입니다. MiniMax H3 Lite is the budget tier of the MiniMax H3 multimodal video model: the same H3 weights as minimax-h3, on a cheaper channel, with a hard ceiling of 768P. It costs a fifth of the standard tier — 768P at $0.02 per second versus $0.097/s on minimax-h3, and 480P at $0.01/s — so a 768P 10-second clip is $0.20 instead of $0.97. What you give up is explicit: no native 2K and no 1080P (768P is the maximum), no reference video (motion transfer), and only two aspect ratios, 16:9 or 9:16. Everything else carries over: any whole length from 1 to 15 seconds, up to 9 reference images to lock character or product, up to 3 reference audio clips to drive voice and lip movement, or a first and last frame for an interpolated shot; references are free. It is not a fast model — plan on 5 to 10 minutes per clip for 10-15 second requests, and a 15-second clip can run longer when the queue is busy; clips of 5 seconds or less usually finish in 1-3 minutes. For comparison, minimax-h3 on our production traffic over the last 14 days had a median turnaround of 6.4 minutes and 90% finished within about 10 minutes, so the two tiers are similar in speed — the difference is price and ceiling, not latency. The limit that matters most is not on the spec sheet: motion amplitude. Lite holds together on animation and on shots where the subject moves a little — talking heads, slow pans, product turns, looping backgrounds, subtle expression changes. On live-action footage with large movement — running, fighting, dancing, fast camera moves — faces and limbs deform and characters lose their shape. That is a property of this tier, not a prompt problem, and no wording of the prompt fixes it. Treat Lite output as a draft: use it to lock composition, timing and framing cheaply, then re-shoot the approved cut on a higher tier if the delivery has to survive large movement. Pick Lite for animation, small-motion shots, drafts, social verticals, iteration and volume; pick minimax-h3 when the delivery needs 2K or has to reproduce the camera move of a reference video. Only successful generations are charged. APIMODELS 플랫폼을 거치면 통합 API와 투명한 종량 과금으로 이 모델을 호출할 수 있습니다. 현재 가격: 480P · per second: $0.01, 480P · 5s: $0.05, 480P · 10s: $0.1, 480P · 15s: $0.15, 768P · per second: $0.02, 768P · 5s: $0.1, 768P · 10s: $0.2, 768P · 15s: $0.3, Reference images (up to 9): $0, Reference audio (up to 3): $0, For comparison: minimax-h3 768P · per second: $0.097, For comparison: minimax-h3 2K · per second: $0.145.
관련 모델: minimax-h3 — the standard tier — native 2K at $0.145/s, 768P at $0.097/s, reference video for motion transfer and seven aspect ratios; same speed, five times the price at 768P
관련 모델: grok-imagine-video-1.5 — the other budget video pick — 480p from $0.0294/s with native synced audio, any length 1-15s
관련 모델: minimax-h3-max-turbo — the fast H3 tier — clips in seconds instead of minutes at 768P $0.096/s, for text-to-video or a single first frame when you do not need reference images
관련 모델: wan-3.0-video — when the shot has to run up to 30 seconds, or start from a document or a web page instead of a prompt
Same H3 model on two channels. Lite is a fifth of the price at 768P but stops at 768P and takes no reference video; both tiers take minutes per clip, so choose on delivery spec, not on speed.
| Dimension | minimax-h3 | minimax-h3-lite · this page |
|---|---|---|
| 768P price | $0.097/s | $0.02/s (a fifth) |
| 480P price | — | $0.01/s |
| Native 2K | $0.145/s | Not available — 768P is the ceiling |
| 768P · 10-second clip | $0.97 | $0.20 |
| Clip length | 5-15s, any whole second | 1-15s, any whole second |
| Aspect ratios | adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 16:9 / 9:16 |
| Reference images | Up to 9 — first 5 free, then $0.054 each | Up to 9, all free |
| Reference video (motion transfer) | Up to 3, billed by their own length | Not supported (400 at create time) |
| Reference audio | Up to 3, free | Up to 3, free (needs ≥1 image) |
| Turnaround (measured) | Median 6.4 min, 90% within ~10 min (14-day production data) | 5-10 min per clip for 10-15s; 1-3 min for ≤5s; 15s can exceed 10 min under load |
| Best for | 2K delivery, motion transfer, wide / square ratios | Drafts, social verticals, iteration, volume |
→ See also minimax-h3 — the standard tier with native 2K and reference video.
All eight clips were generated on this exact model on 2026-08-28 (768p is its ceiling), with native audio and no watermark: seven at 768p ($0.02/s — 15 s $0.30, 12 s $0.24, 10 s $0.20) and one at the 480p tier ($0.01/s — 11 s $0.11). Three run on an audio track (a TTS line or music), four lock a character from reference stills, two are pure text-to-video. The prompt under each clip is exactly what produced it and the tier tag is the list price; expect 5–10 minutes per clip on this channel. For the same prompts at native 2K, use the standard MiniMax H3.
Picture 1 is the woman's face — lock her identity exactly: short black bob with a grey streak, the small scar above her left eyebrow. Picture 2 is her full look — olive military parka, red scarf, wet hair. Picture 3 is the location: the rain-soaked harbor pier at dusk with the single sodium lamp and fog. Picture 4 is the brass pocket watch with the cracked face. 12-second single continuous take, 16:9, photoreal, 35mm anamorphic, rain, native ambient audio (rain on concrete, distant foghorn, the watch's faint tick that stops). 0-3s: Low tracking shot gliding over wet concrete toward her boots; she crouches, picks the watch up from a puddle, water runs off it. 3-6s: Camera rises to a tight close-up as she turns it over in her hands; her face is guarded, jaw tight, breathing controlled. 6-9s: She hears the tick stop — a flicker of hope collapses; her lower lip trembles, eyes fill but she refuses to blink, a single tear breaks and runs down with the rain. 9-12s: She closes her fist around the watch, exhales, lifts her chin toward the fog with hardened resolve; slow push-in ends on her eyes catching the sodium light. Rack focus from watch to face at 6s. No dialogue. No text.
Picture 1 locks the woman's face (short black bob with grey streak, scar above left eyebrow). Picture 2 locks her wardrobe (olive military parka, red scarf) and the rain-soaked pier at dusk. Audio 1 is her voice — she speaks these exact words on camera, lip-synced, addressing someone just behind the lens. 15-second single take, 16:9, photoreal, handheld with slight breathing sway, rain and fog, sodium lamp key light, native audio = Audio 1 plus rain ambience only. Performance: she starts almost calm with a bitter half-smile; on "I waited on this pier every night" her voice cracks and her eyes glisten; on "and now it has stopped" she looks down at the pocket watch in her palm and swallows hard; on "Just like you did" she looks straight into the lens, tears finally falling, then she turns away into the fog. Medium close-up drifting to close-up. No subtitles.
Audio 1 was a minimax-speech-2.6-hd TTS clip of: "You said you would be back before the fog came in. I waited on this pier every night. I kept it ticking for you... and now it has stopped. Just like you did."
Picture 1 and Picture 2 are the same girl — white headphones, white slip dress — keep her face, hair and outfit consistent. Audio 1 is the music; the whole clip is cut to its beat. 15-second vertical 9:16 dance performance in a bright white cyclorama studio, photoreal, high-key light with a moving color wash (white → pink → blue), native audio = Audio 1. 0-4s: she stands still, eyes closed, then snaps them open on the first beat with a playful grin; 4-8s: full-body choreography — sharp arm isolations, a body roll, a spin that whips her hair, camera orbiting 90° around her; 8-12s: she drops into a low floor move and pops back up, laughing, breath visible in effort; 12-15s: she walks straight at the camera, ends on a confident wink and a freeze on the final beat. Fast whip-pans on beats, slow motion for the hair whip only. Big expressive movement, real muscle effort, sweat sheen.
Duration: 15 seconds | Aspect ratio: 16:9 | Style: 2D anime character composited into live-action photorealistic environments SCENE Anime girl with green hair, cat ears, blue and yellow dress, pink backpack, thigh-high boots, and a mechanical tail. Live-action settings include a sunlit bedroom, residential Japanese streets, concrete stairs, and a sandy community park. SHOT BREAKDOWN 0-4s — Character looks panicked beside a digital clock in a bedroom, framed in a medium shot. Cut to her leaping gracefully over a tall concrete wall into a sunny street. 4-8s — She sprints past a traditional wooden house featuring a hanging bronze bell. Cut to a low-angle tracking shot as she runs forcefully down the center of an empty paved road. 8-12s — She ducks and slides under a metal traffic barrier, framed from the side. Cut to her back as she charges up a wide flight of concrete stairs toward distant pedestrians. 12-15s — She flips completely upside-down high in the air. She lands dramatically on one knee in a park dirt field, framed against a line of children and adults standing near a white sign reading "夏休み ラジオ体操". CAMERA Fast-paced tracking shots, dynamic low angles during slides and leaps. Rapid pacing with sharp cuts matching the character's forward momentum. LIGHTING & PALETTE Bright morning sunlight, realistic harsh shadows grounding the 2D subject. Flat cel-shaded character colors featuring neon green, pastel blue, vibrant pink, and bright yellow against natural photorealistic street tones. AUDIO Upbeat energetic anime music, rapid footsteps on concrete, whooshing wind during jumps, heavy thud on the final dirt landing. AVOID 3D character models, missing drop shadows, blurry facial features, unreadable Japanese text, dark or overcast environments.
Duration: 15 seconds | Aspect ratio: 16:9 | Style: Handheld action sports parkour videography SCENE Athletic young man, short black hair, black t-shirt, olive green cargo pants, white sneakers. Urban environment transitioning from narrow concrete alleyways and tiled stairwells to high-rise commercial rooftops with metal railings and HVAC ventilation units. SHOT BREAKDOWN 0-4s — Man executes a parkour vault over a metal barricade fence in a narrow alley, followed by a dynamic leap over a concrete parapet against a clear blue sky. Framed from low tracking angles. 4-8s — Man jumps directly toward the camera down an outdoor flight of tiled stairs. Cuts to a high-angle view as he drops from a higher ledge toward a lower rooftop featuring industrial air conditioning units. 8-12s — Man vaults smoothly over a thick yellow metal railing. Cuts to an overhead perspective tracking him as he scales down a concrete wall to another roof tier. 12-15s — Man lands in a balanced crouch on the edge of a roof parapet. He executes a parkour safety roll onto the flat gray rooftop surface. The camera pulls back to a high aerial wide shot revealing a sprawling metropolitan skyline with skyscrapers. CAMERA Wide-angle action lens. Handheld, energetic tracking movements following the subject's momentum. Rapid cuts matching the physical impact of the jumps. Extreme low and high angles emphasize verticality and height. LIGHTING & PALETTE Bright midday sunlight, harsh direct lighting causing sharp, high-contrast shadows. Dominant colors include neutral concrete grays, muted olive greens, bright flat yellow, and clear sky blue. AUDIO Whooshing wind noise, heavy landing thuds, rapid footsteps, sneaker squeaks, clothing rustle, distant city traffic drone. Fast-paced electronic action instrumental. AVOID Slow motion, static tripod shots, nighttime scenes, indoor settings, formal clothing, stylized neon lighting.
Picture 1 的女生在宽敞的白色摄影棚里做大幅度的街舞动作:0–3 秒从静止突然起跳、双臂大开旋转一整圈落地;3–6 秒连续快速踏步、甩头、身体大幅前倾后仰,长发随动作飞扬;6–8 秒一个侧手翻接下腰,裙摆和头发大幅摆动;8–10 秒冲向镜头急停,做一个定格 pose。镜头低角度手持跟拍,随她的动作大幅推拉和环绕,动感强烈,动作干净利落、肢体完整不变形,保持与参考图完全一致的五官和服装。音乐用 Audio 1,动作卡在节拍上,鞋底摩擦地板和呼吸声清晰。
Picture 1 的女生戴着耳机随着 Audio 1 的音乐轻轻点头微笑,镜头缓慢推近,柔和的白色摄影棚光,自然的皮肤质感。
Picture 1 and Picture 2 are the same girl; she turns her head slowly and smiles
광고 캠페인과 SNS 마케팅에 쓸 브랜드 영상을 빠르게 만듭니다.
TikTok, Instagram, YouTube에 맞는 세로형 짧은 영상을 대량으로 뽑습니다.
기능 소개와 사용법 영상을 만들어 전환율을 끌어올립니다.
강의 해설과 지식 설명, 사내 교육 영상을 낮은 비용으로 꾸준히 제작합니다.
MiniMax H3 Lite은(는) APIMODELS를 통해 480P · per second: $0.01, 480P · 5s: $0.05, 480P · 10s: $0.1, 480P · 15s: $0.15, 768P · per second: $0.02, 768P · 5s: $0.1, 768P · 10s: $0.2, 768P · 15s: $0.3, Reference images (up to 9): $0, Reference audio (up to 3): $0, For comparison: minimax-h3 768P · per second: $0.097, For comparison: minimax-h3 2K · per second: $0.145에 이용할 수 있습니다. 과금은 종량제라서 생성한 만큼만 냅니다.
APIMODELS에 가입해 API 키를 받고 통합 엔드포인트를 호출하면 됩니다. cURL / Python / Node.js 예제를 담은 상세 문서를 제공합니다.
APIMODELS는 같은 MiniMax H3 Lite을(를) 집약 플랫폼을 통해 제공합니다. API 인터페이스가 통합되어 있어 공급자마다 계정을 만들 필요가 없고, 키 하나로 모든 모델에 닿습니다.
**Not recommended — use it with caution here.** This tier is reliable for ordinary scenes: animation, talking heads, slow pushes and pans, product turntables, looping backgrounds. But on live-action footage with fighting, running, dancing or fast camera moves, **faces break down, limbs deform and characters lose their shape**. That is a property of this tier, not a prompting problem — no wording fixes it, and re-rolling for a lucky take is not a strategy. The intended use is as a **draft tier**: lock composition, timing and shot order cheaply here (768P at $0.02/s, a fifth of the standard tier), then re-shoot the approved cut on minimax-h3 — same H3 weights, native 2K, far steadier under large movement. If you need the action shot to be the deliverable, go straight to minimax-h3; this is not the place to save that money.
Same MiniMax H3 model, different channel, ceiling and price. Lite tops out at 768P and bills per second: 480P at $0.01/s and 768P at $0.02/s; minimax-h3 is $0.097/s at 768P and $0.145/s at native 2K. So at 768P, Lite costs a fifth as much (a 10-second clip is $0.20 vs $0.97). What Lite leaves out: native 2K and 1080P, reference video (motion transfer), and any aspect ratio other than 16:9 and 9:16. Character consistency and prompt-following are the model itself and carry over. Speed is about the same on both tiers — both are minutes per clip — so Lite is not the faster option, only the cheaper one.
Plan on 5-10 minutes per clip. Measured: a 768P 10-second clip in about 6 minutes, 12 seconds in about 6.5 minutes, 15 seconds in 8-11 minutes (11 with six jobs submitted at once); clips of 5 seconds or less usually finish in 1-3 minutes. For comparison, minimax-h3 on our production traffic over the last 14 days had a median turnaround of 6.4 minutes with 90% done within about 10 minutes, so the tiers are effectively the same speed. Integrate it asynchronously: create returns a taskId, poll GET /api/v1/video/generations?task_id=xxx every 15-30 seconds, or pass callback_url and we POST the result when it finishes.
Per second, in two resolution tiers: 480P at $0.01/s (5s $0.05, 10s $0.10, 15s $0.15) and 768P at $0.02/s (5s $0.10, 10s $0.20, 15s $0.30). Any whole length from 1 to 15 seconds, billed for the seconds you ask for. Reference images (up to 9) and reference audio (up to 3 clips) cost nothing extra. The same 768P 10-second clip is $0.97 on minimax-h3 and $0.20 on Lite. No minimum top-up, no subscription, and only successful requests are charged — failures are refunded.
Use the unified video endpoint POST /api/v1/video/generations with model set to minimax-h3-lite. prompt is required; duration is any whole second from 1 to 15 (default 5), resolution is 480p or 768p (default 768p), ratio is 16:9 or 9:16 (default 16:9). Attach references as needed: images (up to 9), audio_list (up to 3, needs at least one image), or first_frame_url + last_frame_url for a first/last-frame shot — each accepts a public URL or base64. Field names are identical to minimax-h3, so switching tiers means changing model only.
Resolution tops out at 768P — no 1080P and no 2K. No reference video (sending one returns a clear 400 at create time rather than being dropped silently). Aspect ratio is landscape 16:9 or portrait 9:16 only. There is also a limit that is not on the spec sheet — motion amplitude: on live-action with large movement (running, fighting, dancing, fast camera moves) faces and limbs deform and characters lose their shape, and no prompt wording fixes it, so treat this tier as a draft tier; animation and small-motion shots are unaffected. First/last-frame mode needs both frames and cannot be mixed with reference images or audio; reference audio needs at least one reference image; prompts up to 20,000 characters. Expect 5-10 minutes per clip, longer for 15-second requests. Result files are kept for 7 days.
No — this is the one thing this tier does badly, and it is the thing to know before you pick it. On live-action footage with large movement (running, fighting, dancing, fast camera moves) faces and limbs deform and characters lose their shape. It is a property of the Lite tier, not a prompt problem: rewording, negative prompts and extra reference images do not fix it. What Lite does hold together on is animation and shots where the subject moves a little — talking heads, slow pans, product turns, looping backgrounds, subtle expression changes. So use it as a draft tier: lock composition, timing and framing at $0.02/s, then re-shoot the approved cut on a higher tier if the delivery has to survive large movement. Conversely, if your shot is animation or small motion to begin with, Lite output is deliverable as-is and you do not need to pay five times more.
Decide on the delivery spec, not on speed — both tiers take minutes. Use minimax-h3 when you need 2K or 1080P, when the shot has to reproduce the camera move or motion of a reference video, or when you need 21:9, 4:3 or 1:1. Use Lite for drafts, social verticals, iterating on the same prompt, batch output and tight budgets — five times the footage for the same money. A common workflow is to tune the prompt and references on Lite, then change model to minimax-h3 for the final 2K render. If you need even cheaper, Grok Imagine Video 1.5 (480p from $0.0294/s with native audio) and LTX-2.3 (480p from $0.02/s) go further per dollar; for 30-second shots or document-to-video, look at Wan 3.0.
All 46 of its own tasks in the last 30 days succeeded, though that sample is too small to stand on its own; platform-wide the figure is 95.1% across 950 production video tasks in the last 30 days, platform-wide, median 179 seconds. The definition is stated plainly: the infrastructure success rate counts only failures that are ours to fix (timeouts, upstream congestion, no clip returned) and excludes content-moderation rejections and malformed input, because those clear once you adjust the prompt or the source material, and neither is ever billed. Worth saying out loud: **almost no video-generation platform publishes a success rate at all.** Their API pages offer "stable" and "highly available" — adjectives you cannot check. We publish the real 30-day production numbers together with how they are computed, and you can ask any provider the same question.
No. Timeouts, upstream congestion and content-moderation rejections are all free — you pay only for clips actually delivered, and reserved credits are released automatically on failure. So beyond the success rate itself, your exposure is zero.
No subscription, no monthly fee, no minimum spend — billing is per second of output. The smallest top-up is $10 and one API key covers every model on the platform (video, image, audio and LLMs share a single USD balance). New accounts on consumer email domains get $0.10 in free credit. Paying by Stripe lets you enter a business tax ID at checkout and a formal invoice PDF is generated automatically; PayPal invoices are manual. The account is prepaid, so calls return 402 when the balance runs out rather than continuing to bill — the ceiling is hard.
Median 179s measured in production over the last 30 days. Video generation is heavy work and the industry sits in the minutes range; submit the job, then poll with the taskId or set a callback_url — no long-lived connection required.
Our rate is from $0.01/s at 768P, and failures are never billed. Judge platforms on four things rather than headline price: (1) **the real success rate and how it is computed** — ask whether content-moderation rejections sit in the denominator, and note that most platforms publish no number at all; (2) whether failures are billed; (3) whether input material (reference images and videos) is charged — some channels bill output seconds plus input video seconds; (4) how long result files are kept. All four are answered on this page.
Beyond price, the differences are mostly about **access and settlement**: · **Onboarding** — the official channel requires business verification and an account approval flow; here you sign up and go, with a $10 minimum top-up, no subscription and no minimum spend. · **Region** — the official channel is domestic; this one is reachable worldwide, so overseas teams do not need workarounds. · **Settlement** — everything is priced in USD with Stripe, PayPal and Alipay; paying by Stripe lets you enter a business tax ID at checkout and a formal invoice PDF is generated automatically. · **One key for the whole platform** — video, image, audio and LLMs share a single balance, so you are not opening an account and reconciling invoices with each vendor. · **Failures are never billed** — timeouts, upstream congestion and content-moderation rejections all cost nothing; you pay only for clips delivered. If you already hold an official account, use only this one model, and do not mind settling in CNY through a corporate account, going direct is reasonable. This solves a different problem.
APIMODELS에서는 MiniMax H3 Lite이(가) 60개가 넘는 모델과 같은 API 키, 같은 잔액 위에 나란히 놓입니다. 그래서 선택은 궁합의 문제이지 종속의 문제가 아닙니다. Best for Animation & Small Motion、Draft Tier — Not Delivery、768P Max (No 2K)、480P / 768P、1-15s、A Fifth of minimax-h3、Up to 9 Ref Images (Free)、Up to 3 Ref Audios、First + Last Frame、5-10 min / Clip을(를) 지원하며 다른 영상 생성 모델과 가격·성능을 나란히 놓고 따져볼 수 있습니다. 갈아타기는 모델 이름 문자열 하나만 바꾸면 되고 새 계정도 추가 작업도 필요 없습니다. 영상 생성 선택지와 실시간 가격은 apimodels.app/models에서 볼 수 있습니다.
MiniMax H3 Lite은(는) 다음을 지원합니다: Best for Animation & Small Motion、Draft Tier — Not Delivery、768P Max (No 2K)、480P / 768P、1-15s、A Fifth of minimax-h3、Up to 9 Ref Images (Free)、Up to 3 Ref Audios、First + Last Frame、5-10 min / Clip. 전체 파라미터와 호출 예제는 APIMODELS 문서를 참고하세요.
네. APIMODELS는 MiniMax H3 Lite을(를) 하나의 통합 API와 키 한 개로 제공합니다. 공급자별 계정도 필요 없고, 각 공급자의 지역별 네트워크 경로를 직접 챙길 필요도 없습니다.
Stripe(Visa, Mastercard 등 해외 카드)와 Alipay를 지원합니다. 결제 후 잔액은 즉시 반영됩니다.
Prompts shared by their authors — copy and adapt them. Each one credits its author and links back to the original post.
Seven-shot bedroom candid — cat, phone call, doorbell
Montage, multi-shot candid observational footage. Do not use a single camera angle or continuous take. Handheld documentary style with the feeling of accidental real-life capture. Slightly imperfect framing, subtle handheld shake, tiny reframing adjustments, gentle exposure breathing, and autofocus that settles half a beat late. Realistic skin texture, soft indoor natural light, film grain, shallow depth of field. Relaxed breathing, natural blinking, restrained and authentic performance. Total of 7 shots. The woman from @ Image 1 is wearing a matching cotton pajama set consisting of a scoop-neck sleeveless top and loose shorts made from the same fabric and design. Barefoot, she lies on her stomach across a lived-in, slightly messy bed in her bedroom, browsing a mobile shopping page on a black smartphone. A house cat jumps onto the bed seeking attention, and she affectionately plays with it. Moments later, a friend calls. She answers, chats, and eventually bursts into laughter. The doorbell rings, so she ends the call and gets up to collect her dinner. Shot 1 (0–2s): She is already moving. One thumb scrolls through a shopping page while she frowns slightly, thinking. She quietly mutters in Korean: "아, 뭐 살랬더라…" ("Ah... what was I going to buy again?"). Lips synchronize naturally. Her bare feet gently sway behind her. A low handheld camera glides along the edge of the mattress, with soft bedding partially obscuring the foreground. Shot 2 (2–4s): The house cat jumps onto the bed, compressing the blanket as it walks toward her forearm and lets out a single meow. The camera gives a slight jolt from the impact, then quickly reframes from the phone to an over-the-shoulder view that includes her face, hands, phone, and cat. Shot 3 (4–6s): She turns warmly toward the cat, stroking it once from its forehead to its shoulders. In a gentle, affectionate tone she says: "우리 애기 왔어?" ("Did my baby come?"). Lip sync is clear and natural. The camera pushes in briefly through the cat's softly blurred foreground toward her face and hand. Shot 4 (6–8s): A friend's ringtone sounds. She glances down, confirms the caller, swipes to answer, and brings the phone to her ear. The camera moves in a loose semicircle from a slightly tilted overhead angle, capturing the entire answering motion. Shot 5 (8–11s): Her friend begins talking. She responds first with confusion, then disbelief: "어, 왜? 진짜 거짓말." ("Huh? Why? No way, you're kidding."). Immediately afterward she naturally bursts into laughter. The cat kneads the blanket beside her. A close handheld side angle keeps both her phone and profile within the same focal plane. Shot 6 (11–13s): A clear doorbell interrupts the laughter. Both she and the cat turn toward the bedroom door. She braces one hand on the mattress to get up. The camera reacts a fraction too late, briefly turning toward the door before naturally correcting back to her rising movement. Shot 7 (13–15s): Still on the phone, she hurriedly says: "어, 나 밥 왔다. 끊어." ("My food's here. I'll hang up."). She ends the call with her thumb, gets off the bed, and walks toward the door while the cat follows across the blankets. The camera slides low beneath her elbow and tilts upward as she stands in one continuous motion. Sound: No music. Only raw production sound: bedding rustling, quiet breathing, finger taps on the phone, the soft impact of the cat jumping onto the bed, one natural meow, faint purring, a single short incoming ringtone, call connection tone, the friend's voice through the phone, the woman's four Korean lines with accurate lip sync, a genuine brief laugh, one clear doorbell, the tap ending the call, and the soft cat footsteps. No subtitles, no on-screen text, no logo, no watermark. Never render a reference sheet or duplicate the subject.
by oggii_0
Handheld DV morning dog-walk vlog
CAMERA: Handheld DV 16mm daily vlog footage. The video MUST begin with her holding the camera at arm's length in selfie mode while stepping outside her apartment building with her dog. The first 20–30 seconds are entirely handheld. Later she occasionally places the camera on a park bench, low stone wall, picnic table, or the ground for wider shots. Keep subtle handheld shake, drifting composition, autofocus hunting, rushed reframing, uneven zooms, exposure breathing, brief accidental face cropping, and imperfect framing throughout. The camera itself is never visible. LOOK: Warm analog tape texture with gentle film grain, slightly softened sharpness, subtle halation around sunlight, realistic skin tones, low contrast, tiny exposure shifts, and natural motion blur. It should feel like authentic footage from someone's everyday life rather than a polished commercial. STYLE: A relaxed morning lifestyle vlog. Quiet, cozy, and spontaneous. She occasionally laughs at her dog, pauses to look around, adjusts the leash, brushes hair away from her face, and speaks naturally in short sentences with comfortable pauses. CHARACTER: EMMA — a beautiful white woman in her mid-20s. Long light brown hair in a messy ponytail, green eyes, minimal makeup, oversized gray hoodie, black biker shorts, white sneakers, and a small crossbody bag. She is walking a happy golden retriever. SETTING: A peaceful suburban neighborhood on a sunny morning. Tree-lined sidewalks, quiet residential streets, birds singing, a small park with benches, green grass, and soft golden morning light. Very few people are around. SCENES: The vlog opens in selfie mode. Emma holds the camera while leaving her apartment building with the golden retriever excitedly pulling on the leash. "Good morning." She smiles. "Someone couldn't wait." The dog eagerly sniffs everything as they walk down the sidewalk. She laughs quietly. "He has to inspect every single tree." Still holding the camera, she walks into a small neighborhood park. The dog suddenly stops and stares at a squirrel. "Oh... there we go." She smiles and gently shakes her head. She places the camera on a nearby bench for a wider angle while throwing a tennis ball. The dog happily chases after it. "Worth waking up early." She picks the camera back up. Walking slowly through the park, she looks up at the trees for a moment. "It's actually really peaceful out here." The dog returns with the ball but drops it halfway. She laughs. "Close enough." She sits on the bench while the dog lies beside her. She scratches behind his ears. "I think he's happier than I am." She stands up and continues walking. The camera stays in selfie mode as they head toward home. "Coffee is definitely next." She smiles into the lens. "See you later." She gives a small wave and ends the recording.
by maxxmalist
Thirty seconds of hand-to-hand combat in a rain-flooded metro maintenance hall
生成一段完整30秒、16:9横屏、24fps、写实电影级质感的现代近身格斗视频。使用我上传的两组人物参考图:第一名人物固定为"主角",第二名人物固定为"敌人"。严格继承参考图中两人的面部、年龄、发型、体型、身高比例、服装、鞋子、配饰和整体气质。全程不得交换身份,不得变脸、改变服装颜色、改变体型或生成第三名参战者。 【核心风格】 参考《一个人的武林》所呈现的现代硬派武术电影气质:动作迅猛、贴身、凶狠,攻防转换极快,拳脚具有真实接触感和明确受力反馈,但不照搬电影中的人物、场景和具体镜头。以高质量ACT动作游戏的第三人称战斗镜头诠释两人对决,摄影机主要跟随主角,位于主角肩后、腰后或侧后方,使观众产生操控主角迎战强敌的沉浸感;关键格挡和重击时,可以短暂切入侧面中景或近景。 视频不要求真正一镜到底,可以在30秒内进行多次自然切镜,但整场战斗必须像一个连续长镜头般流畅。利用人物身体遮挡、立柱擦镜、快速摇镜、撞击震动和动作匹配完成隐形切镜。切镜后必须延续上一镜的动作惯性、人物位置、身体朝向、伤势和攻击方向,禁止瞬移、换位错误、人物突然恢复站姿或摄影机无理由越轴。 【场景】 深夜暴雨,一座已经停运的高架地铁检修站。站厅呈狭长矩形,地面是被雨水浸湿的深灰色防滑地砖,分布少量浅水洼,能够反射冷白顶灯和红色维修警示灯。左侧是一排金属检票闸机和关闭的卷帘门,右侧是半透明钢化玻璃护栏,护栏外能够看到高架轨道、雨幕、车灯和模糊城市建筑。站厅中段有三根粗大的混凝土立柱,柱脚带黄黑警示条,尽头是一段通往废弃站台的短楼梯。 冷白灯管是主要光源,其中两盏接触不良、偶尔闪烁;红色警示灯在墙面提供少量轮廓光。斜吹进来的雨水、潮湿薄雾、地面积水、立柱、护栏和闸机都要参与战斗。场景空间在所有镜头中保持一致,不得随切镜改变门、立柱、闸机和玻璃护栏的位置。 开场时主角位于站厅中段偏左,背后靠近检票闸机,面朝右前方;敌人位于右前方约三米处,背后是玻璃护栏。前半段战斗向站厅深处移动,中段绕过第二根立柱,后半段反向压回闸机区域。除非通过环绕镜头明确展示换位,否则敌人始终位于主角前方,保持清楚的行动轴线。 【动作原则】 动作由现代实战拳法、肘击、膝撞、低位踢击、近身控制、擒拿拆解和短距离摔法组成。主角冷静、精准、动作幅度小,善于格挡后立即反击;敌人力量更大、进攻性更强,擅长连续压迫和贴身撞击。敌人不是等待挨打的木桩,主角也不能全程无伤碾压。主角必须经历一次明显失势和一次危险化解,最后依靠判断和重心控制夺回主动。 每次进攻都必须呈现起点、运行路径、接触点和结果;格挡必须真正改变攻击方向;受击者的头部、肩膀、躯干、脚步和重心按照力量方向产生反馈。禁止隔空挥拳、拳脚穿透身体、没有接触却自行后退、无故旋转、连续空翻、悬浮和夸张飞行。动作速度快,但关键接触点必须清楚,不得用混乱残影掩盖动作。 【0—4秒:敌人抢攻】 第一帧已经处于战斗中,不进行对视、摆造型或绕圈试探。摄影机位于主角右后方约一米,接近肩部高度,采用轻微广角的第三人称ACT跟随视角。主角位于画面左侧偏中央,敌人从右前方快速逼近。前景是主角潮湿的肩膀和抬起防守的手臂,中景清楚呈现双方身体,背景是玻璃护栏、雨幕和城市灯光。 敌人右脚踩入水洼,水花向后飞散,同时右直拳攻击主角面部。主角小幅向左偏头,使拳锋擦过脸侧,左手从内侧拍开敌人手腕,右前臂立即架住敌人跟进的左摆拳。敌人不停顿,顺势用右肩和身体重量向前撞击。主角被迫后退半步,鞋底在湿地面短暂滑动,但立刻稳住。 敌人抬膝攻击主角腹部,主角双肘向内压住膝部,将冲击引向身体左侧,同时右脚切入敌人支撑腿外侧,为反击建立位置。摄影机跟随主角后退,敌人始终处于正前方。动作快速但不使用慢镜头。 【4—8秒:第一次反击】 敌人的膝部被压开,主角手臂从镜头前掠过形成遮挡,自然切换到两人右侧的平行移动机位。采用略低的中景,清楚展示脚步、髋部和拳肘路径。 主角左手控制敌人的右前臂,右拳从胸前短距离击中敌人右侧肋部。敌人的上半身因受力向左收缩,呼吸短暂中断,但立刻用左肘横击主角太阳穴。主角抬右臂挡住肘尖,前臂在碰撞中明显震动;随后转动肩膀,以左肘短击敌人胸口,再用右脚低位扫踢敌人左侧小腿。 敌人的左脚被扫得横移,但用右脚及时撑住,没有夸张腾空。他抓住主角上臂,将主角猛烈推向第一根混凝土立柱。摄影机随主角急速后退。主角后背接近立柱时用左掌撑住柱面卸力,沿柱边旋转躲避。敌人紧随而来的重拳擦过主角肩膀,砸中柱面,产生沉闷撞击、少量灰尘和疼痛反馈,但立柱不得碎裂。 敌人击中立柱时,镜头产生一次短促撞击震动。主角快速绕到立柱另一侧,摄影机紧贴主角背后绕柱跟随。立柱短暂遮满画面,完成隐形切镜。遮挡结束后,主角位于立柱右侧,敌人从左侧追出,两人沿站厅纵深向楼梯方向移动。 敌人连续使用左右直拳压迫主角。主角边后撤边格挡:右掌向外拨开第一拳,左前臂贴住第二拳。敌人第三次攻击改为右腿中段踢击。主角来不及完全避开,只能收紧左臂保护肋部。踢击真实命中护臂,力量让主角向右踉跄两步,肩膀撞上玻璃护栏。玻璃发出沉闷震颤却没有破碎,护栏外的雨夜灯光因撞击产生短暂拖影。 敌人跨步逼近,左手压住主角肩膀,右肘攻击面部。主角在最后一刻低头,肘部从头顶掠过;主角左手抓住敌人上臂,右手控制手腕,尝试将其压向玻璃护栏。敌人凭借力量强行抽臂,以额头短促撞击主角眉骨。主角头部向后偏移,额角出现轻微擦伤,目光短暂失焦。 【13—17秒:主角失势】 头撞形成明确节奏重音。快速切入主角侧面近景,表现接触瞬间、飞散雨水和真实眩晕,随即拉回能够看清全身动作的中景,不使用长时间慢动作。 敌人抓住主角上身,将他从玻璃护栏方向拽回站厅中央,右腿绊住主角左腿,试图完成向前摔投。主角身体倾斜,在倒地前用右脚重新踩稳,并以左手勾住敌人后颈,双方进入短暂贴身角力。 摄影机从主角侧后方环绕约九十度,明确展示两人换位:敌人转到画面左侧,主角转到画面右侧,主角背后变为第二根立柱和闸机方向。敌人以力量将主角压在立柱边缘,进行一次肩撞和一次短膝击。主角用大腿与前臂挡住主要冲击,但仍因疼痛弯腰。敌人抬起右肘,准备从上方向下砸击主角后背。镜头靠近主角肩后,使抬起的肘部在画面中形成强烈威胁。敌人的右肘快速下砸。主角不能凭空躲避,而是在肘击启动时向敌人身体内侧贴近,使肘尖从背后落空。主角左手托住敌人肘关节,右手控制其手腕,以肩膀作为支点向前旋转,破坏敌人的肩线。摄影机跟随主角完成半环绕,背景中的立柱、闸机和警示灯提供清楚的空间参照。 敌人为避免手臂被锁,顺势转身,以左拳反击。主角放开控制,低头闪过拳头,起身时用右肩撞击敌人胸口,将其推离立柱。主角立即以左直拳攻击敌人面部防守,迫使敌人抬高手臂;随后右拳从手臂下方击中腹部,再用左肘短击锁骨区域。三次攻击快速递进,但必须分别清楚,不能生成多条手臂。 敌人后退两步,脚后跟撞到检票闸机底座,随即用右腿横扫主角腰部。主角向前贴近,压缩踢击距离,以左臂承受大腿近端冲击,同时右手环住敌人腰侧,左脚踏到敌人支撑脚外侧。低机位镜头清楚交代主角如何控制敌人重心。 【22—26秒:环境摔投】 主角转动髋部并向斜前方发力,将失去支撑的敌人摔向闸机。敌人的肩背先撞在金属闸机侧面,闸机挡板被撞开并发出刺耳金属声,随后敌人落在湿地面,水花向外扩散。敌人的身体必须表现真实重量,不能轻飘飘飞行或在空中翻转多圈。 摄影机沿敌人落地方向快速俯冲,在触地时转为略高的斜俯视ACT视角。主角没有停下摆姿势,而是跨过被撞开的挡板继续追击。敌人翻身用右脚蹬向主角膝部,主角侧移避开。敌人借蹬腿动作翻滚起身,左手撑地,右脚踩稳,还未完全站直便挥出右摆拳。 主角从拳头内侧进入,左臂贴住敌人右臂,右掌推击敌人下颌。敌人向后仰头卸掉部分力量,同时左拳击中主角侧腹。主角明显收腹受击,却保持贴身距离,用右膝短促撞击敌人大腿前侧,进一步破坏其稳定性。两人呼吸加重,衣服逐渐湿污,但动作不能突然变得疲软。 【26—30秒:高潮反制】 摄影机回到主角左后方约一米的第三人称ACT跟随位置。主角位于画面中央偏左,敌人在正前方,背景保持被撞开的闸机、闪烁顶灯和红色警示灯。 敌人最后一次猛烈前冲,先用左拳虚晃主角面部,真正攻击是右肘横扫。主角没有被虚招骗出大幅动作,而是用右掌压下敌人左手,在右肘接近时向前踏入攻击内圈,左前臂架住敌人上臂,使肘尖从主角脑后掠过。主角抓住敌人后颈,右脚勾住其前脚踝,以肩膀向前下方施压。 敌人重心向后倒去,却抓住主角衣服试图将主角一同拖倒。主角立刻松开后颈,身体向左转出抓握,以右肘从极短距离击中敌人胸口偏肩位置,再用左掌推开敌人上胸。敌人向后撞在已经打开的闸机挡板上,挡板进一步弯曲。敌人半跪落地,一只手撑住湿地,另一只手仍保持防守,没有昏迷,也没有飞出画面。 主角受惯性影响向前半步,迅速重新稳住重心,双手保持防守,没有背对敌人、庆祝或摆胜利造型。摄影机缓慢移动到主角右侧,以主角肩膀作为前景,将焦点落在半跪喘息、暂时失去进攻能力的敌人身上。红色警示灯闪烁一次,远处传来雷声。最后一帧保持双方继续互相锁定的危险状态,让主角暂时占据优势,但为下一段战斗留下自然衔接。 【摄影、画面与声音】 保持电影级实拍质感与ACT游戏镜头的融合。主要采用肩后跟随、侧后跟随和腰部高度移动机位;关键命中使用短暂侧面中景或近景。摄影机移动快速但具有真实重量,不得漂浮、瞬移、疯狂旋转或持续无规则抖动。撞击时只使用极短、低幅度震动,震动方向与力量方向一致。 动作中允许适量运动模糊,但人物面部、手脚和接触点在关键时刻保持清晰。禁止鱼眼、长时间慢动作、子弹时间、冻结画面、速度线、游戏血条、准星、按钮提示和文字UI。ACT感来自主角中心构图、第三人称跟随、空间推进和及时的动作反馈。 超写实皮肤、肌肉牵拉、衣料褶皱、湿润头发、汗水和雨水反射;真实金属、混凝土、玻璃与湿地砖材质。冷蓝灰色为主色,冷白顶灯作为主光,红色警示灯作为局部轮廓光。高对比但保留暗部细节,高光不过曝,地面反射自然。禁止蜡像皮肤、塑料CG质感、噪点、脏污颗粒、锐化白边和压缩色块。 生成严格同步的现场声音:雨击顶棚、远处雷声、脚踩积水、湿地摩擦、衣料摆动、急促呼吸、拳掌击中皮肉、前臂格挡、肘膝碰撞、身体撞柱、玻璃震颤和闸机金属回响。无需对白。音乐仅保留极其克制的低频脉冲,不能盖过格斗音效。 【负面约束】 禁止身份互换、面孔漂移、服装变化、身体比例突变、多余人物、复制人物、多手多脚、关节反折、身体穿透和人物粘连;禁止隔空受击、动作没有惯性、敌人静止等待、无故滑行、瞬移、悬浮、武侠轻功、气功、能量特效、夸张冲击波和墙体爆炸;禁止主角全程无伤碾压,禁止敌人只会挨打;禁止无意义空镜、英雄登场、长时间对视、重复招式和结尾突然黑屏。 最终效果必须呈现清晰的动作结构:敌人强势抢攻、主角短暂反击、绕柱追击、主角失势、危险拆解、重新夺回主动、利用闸机完成摔投和最终反制。30秒全程高强度、无废秒,刺激而不混乱,凶狠但遵守人体力学,摄影机始终以主角为核心。
by lansenai
We curate copy-ready prompt libraries — every entry shows its full text and a sample result, ready to adapt.