Wan 2.7 Spicy image-to-video: give it a first frame and a prompt and it returns a 2 to 15 second clip with a generated stereo audio track. You can also pass your own audio file to drive the motion. Output is 720p or 1080p, billed per second.
Authorization: Bearer YOUR_API_KEY| Resolution | Per second | 5-second clip | 10-second clip | 15-second clip | Frame |
|---|---|---|---|---|---|
| 720p | $0.169 | $0.845 | $1.69 | $2.535 | 1280x720 |
| 1080p (default) | $0.26 | $1.30 | $2.60 | $3.90 | 1920x1080 |
Billed on the duration you request (a whole number of seconds from 2 to 15), not on the exact length of the returned file. Audio, a driving audio file, negative prompts and prompt rewriting cost nothing extra. Failed requests are not charged.
| A first frame is required, and it is the only frame | This is image-to-video; there is no text-only mode and omitting image returns a 400. It does not take a last frame either: last_image, image_tail and last_frame_url return a 400 rather than being ignored. Use wan-2.2-i2v-spicy for first-to-last interpolation. |
| First-frame format and size | JPEG, PNG (no transparency), BMP or WebP; each side between 240 and 8000 pixels, aspect ratio between 1:8 and 8:1, at most 20MB. A public URL or base64 is accepted. |
| Duration is a whole number from 2 to 15 seconds | Default 5. Passing 1, 16 or 5.5 returns a 400 stating the legal range — it is never rounded to the nearest legal value and billed anyway. |
| Resolution is 720p or 1080p, default 1080p | Omitting resolution renders and bills at 1080p — pass 720p explicitly to pay less. There is no 480p, 2K or 4K tier. |
| The output has an audio track | By default the MP4 carries a stereo AAC track (44.1kHz, 2 channels in our test) generated to match the scene. Pass audio: false for a silent file, or audio_url to drive the clip with your own sound — see the section below. |
| Generation takes about 45 seconds and up | A 1080p, 2-second clip took 45 seconds in our test; longer and higher-resolution clips take longer. Design your polling around an async task. |
| No watermark | We checked all four corners of the first and last frames of a test clip: no watermark or corner mark. |
# ─── Image-to-video (first frame + prompt, audio generated) ─────
curl -X POST https://api.apimodels.app/v1/video/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "wan-2.7-i2v-spicy",
"image": "https://example.com/first-frame.jpg",
"prompt": "slow cinematic push-in, water flowing, wind in the trees, natural ambient sound",
"duration": 5,
"resolution": "1080p"
}'
# Response (async — a task is created):
# { "code": 200, "data": { "taskId": "clxxx", "state": "pending" } }
# ─── Drive the clip with your own audio (https WAV/MP3, 2-30 s) ───
curl -X POST https://api.apimodels.app/v1/video/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "wan-2.7-i2v-spicy",
"image": "https://example.com/character.png",
"prompt": "she speaks to the camera and gestures naturally",
"audio_url": "https://example.com/speech.mp3",
"negative_prompt": "blurry, distorted face, extra limbs",
"duration": 8,
"resolution": "720p",
"seed": 42
}'
# ─── Poll for the result ──────────────────────────────────
curl "https://api.apimodels.app/v1/video/generations?task_id=TASK_ID" \
-H "Authorization: Bearer YOUR_API_KEY"
# { "code": 200, "data": { "state": "completed", "resultUrls": ["https://r2.apimodels.app/..."] } }| Field | Required | Type | Description |
|---|---|---|---|
| model | Yes | string | wan-2.7-i2v-spicy |
| image | Yes | string | First frame: a public image URL or base64. image_url and first_frame_url are accepted too. |
| prompt | Yes | string | How the shot should move and what it should sound like |
| negative_prompt | No | string | What to keep out of the clip, e.g. blurry, extra limbs |
| audio_url | No | string | Driving audio: an https link to a WAV or MP3, 2 to 30 seconds, at most 15MB. audios[0] is accepted too. |
| audio | No | boolean | Whether the output has an audio track, default true. When false the file is silent even if audio_url was given. generate_audio is accepted too. |
| duration | No | number | A whole number from 2 to 15, default 5. Anything else is a 400 — it is never silently changed to a legal value. |
| resolution | No | string | 720p or 1080p, default 1080p. There is no 480p, 2K or 4K tier. |
| seed | No | number | Random seed (0 to 2147483647); random when omitted |
| prompt_extend | No | boolean | Automatic prompt rewriting, on by default, no extra charge. See the section below. |
| callback_url | No | string | Webhook URL called when the task completes |
Without audio_url the model generates sound to match the scene. With audio_url it drives the motion from your audio instead, and that audio is the track in the output. The upstream example use is a character speaking and gesturing.
Constraints: an https link to a WAV or MP3, 2 to 30 seconds, at most 15MB. audio: false takes priority — set it and the output is silent even if audio_url was given.
When enabled, your prompt is first expanded with camera, lighting, motion and sound detail, and the expanded version is what gets generated. It helps most with one-line prompts. The switch is on by default and costs nothing extra.
Pass prompt_extend: false to turn it off and have the model follow your text exactly. Turn it off when you have already written a long, precise prompt, when the exact scene in the first frame must be preserved, or when you need reproducible comparisons alongside seed — each rewrite differs, so a fixed seed alone will not line up.
The Spicy line is two complementary models. Wan 2.2 Spicy is the volume engine: silent, 5 or 8 seconds, from $0.03/s, for game sprites, chat avatars, social loops and animated feed covers. Wan 2.7 Spicy is the one you sell: native audio, 2 to 15 seconds, 1080p — the four scenarios below run on it.
| Premium subscription and pay-per-view content | Film-grade clips that drive a subscription or a one-off purchase: the native audio carries the immersion without a post step, 15 seconds is enough for a complete micro-story, and 1080p supports premium pricing. A 15-second 1080p clip costs $3.90; a 10-second 720p one $1.69. |
| AI role-play and interactive narrative | Branching stories where the user's choice decides where the video goes. Audio gives immediate emotional feedback, 2 to 15 seconds fits every narrative beat, and each branch is one API call, so the architecture stays simple. |
| Creator tools and SaaS platforms | One-click professional video for creators: built-in prompt rewriting removes the learning curve, native audio means clips publish without an edit, and the 720p / 1080p tiers cover different platform requirements. |
| AIGC content marketplaces | Marketplaces where creators generate and sell premium AI video. 1080p plus sound is a quality bar cheaper models cannot reach, and prompt rewriting keeps output consistent across creators. |
Worth knowing from the latest upstream build: fuller audio, smoother motion and better frame consistency; anime-style and Asian-face scenes added; prompt rewriting now runs in half the time.
Per second of the duration you request: $0.169 at 720p and $0.26 at 1080p. A 5-second clip is $0.845 or $1.30, a 15-second clip $2.535 or $3.90. Audio, a driving audio file, negative prompts and prompt rewriting cost nothing extra, and failed requests are not charged.
Any whole number of seconds from 2 to 15 (default 5) and a resolution of 720p (1280x720) or 1080p (1920x1080, the default). A fractional or out-of-range duration, or any other resolution, returns a 400 stating the legal values rather than being silently adjusted and billed.
Yes. By default the MP4 comes back with a stereo AAC track the model generates to match the scene. Pass audio_url with an https link to a 2–30 second WAV or MP3 to drive the clip with your own sound, or audio: false to get a silent file. This is the main difference from Wan 2.2 Spicy, which has no audio.
No. Wan 2.7 Spicy takes exactly one first frame; last_image, image_tail and last_frame_url are rejected with a 400 rather than ignored, so you never pay for a clip that differs from what you asked for. For first-to-last interpolation use wan-2.2-i2v-spicy on the same endpoint.