An enhanced build of the open-source MiniMax H3 video model. One model id, three modes picked from what you send: text-to-video, first/last-frame image-to-video, and multi-reference video (up to 9 reference images and 3 reference videos). Every mode can be driven by your own audio for lip sync, and audio_mode decides how the soundtrack is handled. 480P / 768P / 1080P, 4-15 seconds, billed for the seconds you request: 480P $0.05/s, 768P $0.07/s, 1080P $0.10/s.
| Resolution | Per second | 5s | 10s | 15s |
|---|---|---|---|---|
| 480P | $0.05 | $0.25 | $0.50 | $0.75 |
| 768P | $0.07 | $0.35 | $0.70 | $1.05 |
| 1080P (default) | $0.1 | $0.50 | $1.00 | $1.50 |
Billed for the seconds you request, 4 to 15. Reference images, reference videos, driving audio and voice references cost nothing extra. Charged only on success, refunded on failure.
All requests carry the API key in the header:
Authorization: Bearer YOUR_API_KEY/api/v1/video/generations·GET/api/v1/video/generations?task_id=# Text-to-video (prompt only)
curl -X POST https://api.apimodels.app/v1/video/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "minimax-h3-rh-enhanced",
"prompt": "A potter shapes a clay bowl on a spinning wheel, she looks up and smiles at the camera, warm window light",
"duration": 8,
"resolution": "768p",
"aspect_ratio": "16:9"
}'
# Image-to-video with a driving audio track (lip sync)
curl -X POST https://api.apimodels.app/v1/video/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "minimax-h3-rh-enhanced",
"prompt": "She keeps working and talks to the camera with a warm smile",
"first_frame_url": "https://example.com/first.jpg",
"audio_url": "https://example.com/voice.mp3",
"duration": 6,
"resolution": "1080p",
"aspect_ratio": "3:4"
}'
# Multi-reference music video: keep your soundtrack untouched
curl -X POST https://api.apimodels.app/v1/video/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "minimax-h3-rh-enhanced",
"prompt": "The singer from image 1 performs on the stage from image 2, moving like the dancer in the reference video",
"reference_image_urls": ["https://example.com/singer.jpg", "https://example.com/stage.jpg"],
"reference_video_urls": ["https://example.com/dance.mp4"],
"audio_url": "https://example.com/song.mp3",
"audio_mode": "lock_source",
"duration": 10,
"resolution": "1080p",
"aspect_ratio": "9:16"
}'
# Poll until completed (result files are kept 30 days)
curl "https://api.apimodels.app/v1/video/generations?task_id=TASK_ID" \
-H "Authorization: Bearer YOUR_API_KEY"A first or last frame cannot be combined with reference images or videos — that request returns a 400 at create time. All three modes take audio_url (driving audio for lip sync) and audio_mode.
| Field | Required | Type | Description |
|---|---|---|---|
| model | Yes | string | minimax-h3-rh-enhanced |
| prompt | Text-to-video | string | Subject, action, camera, dialogue and sound; up to 20,000 characters. |
| duration | Yes | number | 4 to 15 seconds; billed for this many seconds. |
| resolution | No | string | 480p ($0.05/s) / 768p ($0.07/s) / 1080p ($0.10/s, default). |
| aspect_ratio | No | string | 1:1 / 2:3 / 3:2 / 3:4 / 4:3 / 9:16 / 16:9 / 21:9 |
| first_frame_url | No | string | First frame (public URL or base64). image is also accepted. |
| last_frame_url | No | string | Last frame. |
| reference_image_urls | No | string[] | Reference images, up to 9. images is also accepted. |
| reference_video_urls | No | string[] | Reference videos, up to 3. video_list is also accepted. |
| reference_video_audio_urls | No | string[] | Audio of the reference videos, up to 2. |
| audio_url | No | string | Driving audio for lip sync. drive_audio_url is also accepted. |
| reference_audio_urls | No | string[] | Voice references, up to 3 (image-to-video and multi-reference). audio_list is also accepted. |
| audio_mode | No | string | native (default) / lock_source (keep the driving audio untouched, for music videos) / remix_source / reference_only. |
| callback_url | No | string | We POST the result to this URL on completion; omit it and poll the GET endpoint instead. |
Other models in the family: minimax-h3 (native 2K), minimax-h3-lite, minimax-h3-max-turbo (clips in seconds).
Try it in the Playground
Per second of the length you request on apimodels.app: 480P $0.05/s, 768P $0.07/s, 1080P $0.10/s, from 4 to 15 seconds — a 768P 10-second clip is $0.70, a 1080P 15-second clip $1.50. Reference images, reference videos and audio cost nothing extra. Only successful generations are charged; failures are refunded.
POST /api/v1/video/generations with model minimax-h3-rh-enhanced, a duration (4-15 seconds) and a resolution (480p, 768p or 1080p; default 1080p). A prompt alone is text-to-video; add first_frame_url / last_frame_url for image-to-video, or reference_image_urls / reference_video_urls for multi-reference. audio_url sets the driving audio for lip sync and audio_mode controls the soundtrack. The call returns a taskId — poll GET /api/v1/video/generations?task_id= or pass callback_url.
It is picked from the inputs. Reference images, reference videos or reference-video audio select multi-reference; otherwise a first or last frame (or voice references alone) selects image-to-video; otherwise it is text-to-video and the prompt is required. A first or last frame cannot be combined with reference images or videos — that request returns a clear 400.
It decides how the soundtrack is handled: native (default) lets the model create the audio; lock_source keeps your driving audio exactly as it is, which is what music videos need; remix_source and reference_only use your audio more loosely. Pass a driving track with audio_url.