Alibaba's Wan 3.0 all-in-one video model: one model name covers text-to-video, first-frame and first+last-frame image-to-video, and omni-reference generation from up to 10 reference images, 5 reference video clips and 5 audio clips — or straight from a PPTX / DOCX / XLSX / PDF or a public web page. Output is a native single shot up to 30 seconds at 30fps with a synced audio track. Billing is per second: with reference images, audio, documents or web pages you pay for OUTPUT seconds only and those inputs cost nothing extra; **once you send a reference VIDEO (video edit, video extend, character swap) you pay for INPUT video seconds + OUTPUT seconds**. This matches the upstream, which also caps such jobs at input + output ≤ 30 seconds. It runs on the same unified video endpoint as VEO, Kling and Seedance, so switching models is a one-string change and one API key covers everything — with no Alibaba Cloud account and no mainland real-name verification needed.
Billed per second. Without a reference video you pay for output seconds. With a reference video (video edit, extend, character swap) you pay for INPUT video duration + OUTPUT duration — matching the upstream, which also caps these jobs at input + output ≤ 30 seconds. Reference images, audio, documents and web pages cost nothing extra.
| Tier / resolution | Rate | 5s | 10s | 30s |
|---|---|---|---|---|
| standard · 480P | $0.045/s | $0.225 | $0.45 | $1.35 |
| standard · 720P | $0.09/s | $0.45 | $0.90 | $2.70 |
| standard · 1080P | $0.18/s | $0.90 | $1.80 | $5.40 |
| prime · 480P | $0.068/s | $0.34 | $0.68 | $2.04 |
| prime · 720P | $0.14/s | $0.70 | $1.40 | $4.20 |
| prime · 1080P | $0.28/s | $1.40 | $2.80 | $8.40 |
POST /api/v1/video/generations — async: create, then poll the same endpoint for the result.
# Step 1: Create task (text-to-video)
curl -X POST https://api.apimodels.app/v1/video/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "wan-3.0-video",
"mode": "prime",
"prompt": "A paper airplane gliding over a calm lake at sunrise, camera drifting alongside",
"resolution": "720P",
"ratio": "16:9",
"duration": 5
}'
# Omni-reference: address materials positionally in the prompt
curl -X POST https://api.apimodels.app/v1/video/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "wan-3.0-video",
"prompt": "The person in video 1 picks up the object from figure 2 and places it on the table",
"reference_image_urls": ["https://example.com/object.png"],
"reference_video_urls": ["https://example.com/person.mp4"],
"resolution": "720P",
"duration": 8
}'
# Document to video: feed a deck straight in
curl -X POST https://api.apimodels.app/v1/video/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "wan-3.0-video",
"prompt": "A polished product launch film based on this deck",
"file_url": "https://example.com/launch-deck.pptx",
"resolution": "1080P",
"duration": 10
}'
# Step 2: Poll status
curl "https://api.apimodels.app/v1/video/generations?task_id=TASK_ID" \
-H "Authorization: Bearer YOUR_API_KEY"| Mode | What you send |
|---|---|
| Text to video | prompt only |
| First / first+last frame | first_frame_url [+ last_frame_url] |
| Omni reference | reference_image_urls / reference_video_urls / reference_audio_urls |
| Document / web to video | file_url or link_url |
⚠️ Frame mode (first_frame_url / last_frame_url) and reference mode (reference_* / file_url / link_url) are mutually exclusive — mixing them makes upstream fail the task, so we reject it locally first and name both colliding sides. file_url and link_url are likewise one or the other.
In omni-reference mode the prompt addresses materials positionally — "figure 1", "figure 2", "video 1", "audio 1", with images, videos and audio numbered independently. For example: "the person in video 1 holds the object from figure 3 and plays guitar on the chair from figure 4." That is how several characters, props and environments stay consistent inside one shot.
| Field | Required | Type | Description |
|---|---|---|---|
| model | Yes | string | "wan-3.0-video" |
| mode | No | string | "standard" (default) or "prime". Prime is the high-speed tier — 58s versus 356s on the identical job — at 51-56% more per second depending on resolution. Capabilities are identical per Alibaba; only speed and price differ |
| prompt | No | string | Text description. Send at least one of prompt / reference material; up to 20,000 characters |
| first_frame_url | No | string | First frame, used strictly as frame one. Mutually exclusive with reference material |
| last_frame_url | No | string | Last frame, used strictly as the final frame |
| reference_image_urls | No | string[] | Reference images, up to 10. Addressed in the prompt as "figure 1", "figure 2", … |
| reference_video_urls | No | string[] | Reference videos, up to 5 clips, 1-15s each and 15s total |
| reference_audio_urls | No | string[] | Reference audio, up to 5 clips, 15s total (wav / mp3) |
| file_url | No | string | Document: docx/doc/xlsx/xls/pptx/ppt/pdf/txt/key/pages/numbers/md, ≤50 pages and ≤100MB. Mutually exclusive with link_url |
| link_url | No | string | Public web page URL (must be readable without login) |
| resolution | No | string | "480P" / "720P" / "1080P". ⚠️ We default to 720P when omitted (Alibaba defaults to 1080P, which costs double) |
| ratio | No | string | "adaptive" (default — inferred from your inputs), "16:9", "4:3", "1:1", "3:4", "9:16" |
| duration | No | number | 2-30 seconds, default 5; pass -1 for smart duration (the model chooses). With reference video: input video length + output length must stay within 30s |
| audio | No | boolean | Whether the output carries an audio track, default true. Turning it off does not reduce the price |
| prompt_extend | No | boolean | Smart prompt rewriting, default true. Helps short prompts noticeably but adds latency |
| seed | No | number | Random seed, 0-2147483647, for reproducible results |
| watermark | No | boolean | Whether to add a watermark, default false |
| callback_url | No | string | Webhook URL called when the task completes |
The create call returns data.taskId right away and the video renders in the background. GET the same endpoint with that task_id and read data.state: pending means still running, completed puts the MP4 URL in data.resultUrls[0], and failed puts the reason in data.failMsg. To skip polling, pass callback_url and we POST you when the task finishes.
Yes. On apimodels.app you POST to /api/v1/video/generations with model "wan-3.0-video" (add mode:"prime" for the high-speed tier), a prompt and/or reference material, then poll the same endpoint with the returned task_id until data.state is "completed" — data.resultUrls[0] is the MP4. It is the same request shape as VEO, Kling and Seedance here, so switching models is a one-string change, and one API key covers all of them.
Because most developers outside China cannot open that account. Wan 3.0 is served only from Alibaba Model Studio's China (Beijing) and Singapore regions — we probed US/Virginia and it does not carry the model at all, returning AccessDenied.Unpurchased while that workspace lists 92 models with not one video model among them. The China station requires mainland Chinese real-name verification to register. Through apimodels.app you need no Alibaba Cloud account, no real-name check and no separate billing relationship — same key and endpoint as every other model here.
Per second. The standard tier is 10% below Alibaba's official USD list price: 480P $0.045/s, 720P $0.09/s, 1080P $0.18/s — a 5-second 720P clip is $0.45, a 30-second 1080P is $5.40. mode=prime is at list price, 51-56% more per second than standard ($0.068 / $0.14 / $0.28). Billable seconds depend on what you send in: with reference images, audio, documents or web pages you pay for OUTPUT seconds only and those inputs cost nothing extra; with a reference VIDEO (video edit, extend, character swap) you pay for INPUT video duration + OUTPUT duration, matching the upstream, which also caps such jobs at input + output ≤ 30 seconds. The audio-track toggle does not change the price. Pass duration -1 and we hold the 30-second worst case, then settle on the real length and refund the difference. Failed tasks are not billed.
Speed and price, not quality. Alibaba states Prime's capabilities match the standard model — same omni-reference inputs, same 30-second 30fps ceiling, same native audio, same aspect ratios. On an identical prompt and identical parameters (480P, 9:16, 5s, audio off) Prime returned in 58 seconds where the standard tier took 356 — about 6x faster, measured back-to-back here. Prime costs 36% more per second at 480P and 40% more at 720P and 1080P. Use Prime for interactive tools and iteration loops; use standard for batch jobs where six extra minutes cost nothing.
Yes. duration accepts any integer from 2 to 30 and the output is a native single continuous shot at 30fps, not stitched segments — that is the biggest difference from models capped at 5-10 seconds. Pass -1 for smart-duration mode and the model picks a length between 2 and 30 from your prompt and materials. One constraint: when you supply reference video, input video length plus output length must stay within 30 seconds.
Four modes under one model name. Text-to-video (prompt only). First-frame or first+last-frame image-to-video. Omni-reference: up to 10 reference images, 5 reference video clips (15s total) and 5 audio clips (15s total), which the prompt addresses positionally as "figure 1", "video 1", "audio 1". And document/web-to-video: a PPTX, DOCX, XLSX, PDF, TXT or MD file up to 50 pages, or a publicly readable web page. One hard rule: frame mode and reference mode are mutually exclusive — mixing them fails upstream, so we reject it locally first and name both colliding sides rather than letting you burn a round trip.