
minimax-h3MiniMax H3 is a general-purpose multimodal video model rather than a stack of single-purpose tools. It reads text, images, video and audio in the same pass, works out what the brief is actually asking for, and returns one coherent clip with picture, motion and sound already in agreement — no separate dubbing, lip-sync or motion-transfer step to stitch together afterwards. Practically, that means one endpoint covers the whole range: send a prompt on its own for text-to-video, or attach up to 9 reference images, 3 reference videos and 3 reference audio clips to pin down character, camera movement and voice at once. Every reference input is optional. Choose any whole-second length from 5 to 15 and an aspect ratio of adaptive, 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16. Billed per second at $0.145/s (5s = $0.725, 15s = $2.175); the first 5 reference images are free and each additional image is $0.03. Only successful generations are charged.
At $0.145/s it costs 29.5% of our own Seedance 2.0 1080p tier ($0.492/s) — and outputs native 2K rather than 1080p. Both figures are published on this site, so you can check the comparison yourself
Text, image, video and audio are understood together rather than by separate specialist tools — visuals, motion and sound arrive already in sync, in a single pass
Every generation renders at 2K with synchronized audio — no upscaling step and no lower-resolution tier to pick between
Prompt alone behaves as text-to-video; add images, videos or audio to steer character, camera motion and voice in the same request
Whole-second control rather than fixed buckets, billed per second at $0.145/s so you pay for exactly the length you asked for
Up to 9 reference images, 3 reference videos and 3 reference audio clips — the first 5 images are free, extras are $0.03 each
Leave the references below empty for pure text-to-video — every reference input is optional.
Steer character, style or composition. The first 5 are free; each image past the 5th adds $0.03.
MP4 / MOV, ≤50MB and 2-15s each. Use these to carry camera movement or motion style into the output.
MP3 / WAV, ≤15MB and 2-15s each. Used to match voice or drive lip movement.
Generated video will appear here
Provide URLs and click Generate
MiniMax Hailuo H3 是由 Minimax 提供的视频生成 API。MiniMax H3 is a general-purpose multimodal video model rather than a stack of single-purpose tools. It reads text, images, video and audio in the same pass, works out what the brief is actually asking for, and returns one coherent clip with picture, motion and sound already in agreement — no separate dubbing, lip-sync or motion-transfer step to stitch together afterwards. Practically, that means one endpoint covers the whole range: send a prompt on its own for text-to-video, or attach up to 9 reference images, 3 reference videos and 3 reference audio clips to pin down character, camera movement and voice at once. Every reference input is optional. Choose any whole-second length from 5 to 15 and an aspect ratio of adaptive, 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16. Billed per second at $0.145/s (5s = $0.725, 15s = $2.175); the first 5 reference images are free and each additional image is $0.03. Only successful generations are charged. 通过 APIMODELS 平台,您可以使用统一的 API 接口调用该模型,按量计费、价格透明。当前定价:2K · per second: $0.145, 2K · 5s: $0.725, 2K · 10s: $1.45, 2K · 15s: $2.175, Extra reference image (past 5): $0.03。
快速生成品牌宣传视频,适用于广告投放和社交媒体推广。
为抖音、小红书等平台创建引人注目的短视频内容。
生成产品功能演示和使用教程视频,提升用户转化率。
制作课程讲解、知识科普等教育类视频,降低视频制作门槛。
MiniMax Hailuo H3 通过 APIMODELS 平台调用,当前定价:2K · per second: $0.145, 2K · 5s: $0.725, 2K · 10s: $1.45, 2K · 15s: $2.175, Extra reference image (past 5): $0.03。按量计费,用多少付多少。
在 APIMODELS 注册账号并获取 API Key,然后通过我们的统一 API 端点调用即可。我们提供详细的 API 文档和 cURL、Python、Node.js 代码示例。
APIMODELS 通过聚合平台提供与官方相同的 MiniMax Hailuo H3 模型。我们提供统一的 API 接口,无需分别注册各平台账号,一个 API Key 即可调用所有模型。
MiniMax 海螺 H3 是 MiniMax 的多模态视频生成模型,一次请求同时产出原生 2K 画面与同步音频,不需要另外配音。它和多数视频模型最大的不同在于「一个接口打通全部用法」:参考图、参考视频、参考音频全部是可选输入,只填提示词就是文生视频,加上素材就变成图生视频、动作参考或音色参考,不必在不同端点之间切换。
在 APIMODELS 上按秒计费 $0.145/秒:5 秒 $0.725、10 秒 $1.45、15 秒 $2.175。参考图前 5 张免费,第 6 张起每张 $0.03(最多 9 张,即最高 +$0.12);参考视频和参考音频不额外收费。没有起充门槛、没有订阅,只对成功的请求计费——生成失败不扣费。
走统一的视频端点 POST /api/v1/video/generations,把 model 设成 minimax-h3(也接受 hailuo-h3 这个别名),必填 prompt 和 duration(5-15 整秒)。参考素材按需传:images(≤9)、video_list(≤3)、audio_list(≤3),都支持公网 URL 或 base64。想要指定画幅就传 ratio。任务创建后返回 taskId,用 GET /api/v1/video/generations?task_id=xxx 查询,或者在创建时传 callback_url 让我们在完成时回调你。同一个 API Key 可以调用站内全部模型,国内可直连、无需组织认证。
分辨率固定原生 2K(这个渠道只提供 2K,没有更低档);时长 5-15 秒任意整秒,不是固定档位;画幅可选 adaptive、21:9、16:9、4:3、1:1、3:4、9:16。参考素材上限:图片最多 9 张(每张 ≤30MB)、视频最多 3 段(MP4/MOV,每段 ≤50MB 且 2-15 秒)、音频最多 3 段(MP3/WAV,每段 ≤15MB 且 2-15 秒)。提示词最长 20480 字符。生成结果我们会转存到自有存储,链接保留 7 天,请及时下载或转存。
它的强项是「原生 2K + 同步音频 + 多模态参考」三件事同时要:需要人物一致性的连续镜头、要复刻特定运镜的广告片、需要指定音色或对口型的口播视频,以及要把图、视频、音频混合参考的复杂创意。性价比上它其实很硬:$0.145/秒是我们自己 Seedance 2.0 1080p 档($0.492/秒)的 29.5%,而且输出的是原生 2K 不是 1080p —— 两个价格都公布在站内,可以自己核对。如果你只要短平快的竖屏素材且预算敏感,LTX-2.3(480p $0.02/s 起)或 Seedance 2.0 Fast($0.11/s 起)仍然更省;如果需要 4K 或视频编辑,可以看 Omni Flash。在 APIMODELS 上这些模型共用一个 API Key 和同一套端点,换模型只改 model 这一个字段。
在 APIMODELS,MiniMax Hailuo H3 与 60+ 个模型共用一个 API Key、一个余额,所以选型只看合不合适,不存在锁定。它支持 Native 2K、Multimodal Reference、5-15s、Up to 9 Ref Images、3 Ref Videos + 3 Audios、Synchronized Audio,你可以在价格和能力上把它和其它视频生成模型对比,换模型只需改一个模型名字符串——无需新账号、无需重新对接。所有视频生成模型与实时价格见 apimodels.app/models。
MiniMax Hailuo H3 支持:Native 2K、Multimodal Reference、5-15s、Up to 9 Ref Images、3 Ref Videos + 3 Audios、Synchronized Audio。完整参数与调用方式见 APIMODELS 的 API 文档。
可以。APIMODELS 提供可直接访问的统一 API,一个 API Key 即可调用 MiniMax Hailuo H3,无需分别注册官方账号、也无需自行处理官方接口的网络访问。
我们支持 Stripe(Visa、Mastercard 等国际信用卡)和支付宝付款。充值后积分即时到账。
社区作者公开的提示词,可直接复制修改。每条都标注作者并回链原帖。
【MiniMax H3】Hailuo3のコストの件💰 @Hailuo_AI #MiniMaxH3 アーリーアクセス開始直後は10クレジット/秒でしたが、今は12クレジット/秒と20%増になっています📈 あと、動画参照がある場合と、画像参照が6枚以上ある場合は別途加算されるようです(生成秒数とは別で計算) ────────── ①【出力】生成秒数依存部分 生成秒数×12クレジット ②【入力】動画参照秒数依存部分 参照動画秒数×12クレジット ③【入力】画像参照数依存部分 6枚目以降、1枚あたり+3クレジット(5枚目までは無料) ────────── 例1)①生成秒数10秒、②動画参照なし、③画像5枚以内 ①12×10 = 120クレジット(12/秒) 例2)①生成秒数10秒、②動画参照10秒、③画像5枚以内 ①12×10 + ②12×10 = 240クレジット(24/秒) 例3)①生成秒数4秒、②動画参照15秒、③画像参照9枚(そんなことある?😇) 12×4 + 12×15 + 3×4 = 240クレジット(60/秒) ────────── 動画参照する場合、必要な部分にだけちゃんとトリミングした方がいいですね 幸いUIが優秀で長尺のトリミングも簡単です ちなみに音源参照はクレジット影響なしです ────────── という内容をH3にそのまんまプロンプトとして貼り付けて、パピヨンのキャラクターシートを参照させつつ「この内容を説明して」的な指示でできた動画が👇です ①内容をわかりやすく要約してゆっくり聞き取りやすい日本語で ②無駄に大きい動作でハイテンションで早口で15秒以内に全部読んで 指示通りがんばってくれました🤣 ②は必死過ぎてもう最後の「影響なしです!」以外ほぼ何言ってるかわからないですね😂でもなんかめっちゃがんばってるから繰り返し見ちゃう……笑
作者 @projectmuse_ai

🚧MiniMaxH3でオーディオ参照&リップシンク🚧 添付は無編集の出たまんま動画 まず参照オーディオの維持性能がすごく高い そしてリップシンクも綺麗 リリックモーションは私自身がまだ探求できてないので試行錯誤の余地有り 可愛いエフェクトたくさん出してくれるとこも好きよ @Hailuo_AI #MiniMaxH3 https://t.co/gjF9o0kv5v
作者 @Ushizaru_LAB
#MiniMaxH3 Hailuoの新モデルMiniMaxH3の早期アクセスをさせていただいたので早速試してみました❣️ 悪くはないけどもう少し激しく躍らせてみたいな✨ 色々試してみます☺️ @Hailuo_AI https://t.co/ib0fff7mmZ
作者 @mugi_AI_Art