
wan-3.0-videoWan 3.0 is Alibaba's all-in-one video model: a single model name that covers text-to-video, first-frame and first+last-frame image-to-video, and omni-reference generation driven by up to 10 reference images, 5 reference video clips (15s total) and 5 audio clips — or a document (PPTX, DOCX, XLSX, PDF, TXT, MD, up to 50 pages) or a public web page. Output is native single-shot up to 30 seconds at 30fps with a synced audio track you can switch off at no price difference, in 480P / 720P / 1080P and six aspect ratios including an adaptive mode that picks the ratio from your inputs. In omni-reference mode the prompt can address materials by position — "figure 1 holds the object from figure 2, walking past video 1" — which is how you keep several characters, props and environments consistent across a shot. Billing is per second at 480p $0.045/s, 720p $0.09/s, 1080p $0.18/s. Reference images, audio, documents and web pages cost nothing extra, but a reference VIDEO (video edit, video extend, character swap) is billed on its own duration too — input video seconds + output seconds, matching the upstream, which also caps those jobs at input + output ≤ 30 seconds. So a 5-second 720p clip is $0.45 and a 30-second 1080p is $5.40; pass duration -1 and the model picks a length between 2 and 30 seconds, in which case we hold the 30-second worst case and refund the difference once the real length is known. Measured on this platform: a 480p 5s clip returned in 356 seconds on this standard tier — reach for Wan 3.0 Video Prime if that matters, it did the identical job in 58 seconds. Async create-then-poll on the shared video endpoint; failed tasks are never billed. We do not add a content-moderation layer of our own: prompts and reference media pass through as sent. The model's own safety systems still apply upstream and do fire in practice — both inputs and outputs are screened, and real human faces in reference material are commonly refused — so this is not an unfiltered endpoint. Reviewing what you generate, and complying with the law and your own platform's rules, is your responsibility; use only with subjects who have consented.
Prompt only
Document-to-video (PPTX / DOCX / XLSX / PDF) and web-page-to-video are API-only for now — see /docs/wan-3-0-video.
Generated video will appear here
Provide URLs and click Generate
Single-shot up to 30s at 30fps — or pass duration -1 and let the model choose
Up to 10 images + 5 videos + 5 audio clips; the prompt addresses them as "figure 1", "video 1"
Feed a PPTX / DOCX / PDF or a public URL and the model reads it into the shot
From $0.045/s at 480p. Reference images, audio, documents and web pages cost nothing extra — but a reference VIDEO (video edit, extend, character swap) bills INPUT video seconds + OUTPUT seconds, matching the upstream
mode:"prime" returns in 58s where standard took 356s on the identical job — 51% more per second at 480p, 56% at 720p and 1080p
When Alibaba's Tongyi Wanxiang team launched Wan 3.0 they published a creator handbook with 64 prompt-and-clip pairs. We re-hosted every clip, kept every prompt verbatim, and measured the files: 64 clips, 1,112 seconds of video in total, every one of them with a synced audio track. Seventeen clips run the full 30 seconds Wan 3.0 allows in a single shot; fifteen are 5-second clips (edits and video-reference jobs). Forty-nine clips fall in the 720P tier (1280×720 or its portrait twin) and fifteen in the 1080P tier (1920×1072 and similar). The full set, sorted by job type, lives in our Wan 3.0 prompt library; the table below is what the set tells you about where the model is used.
The split matters because it is the publisher's own view of what Wan 3.0 is for. Text-to-video and image reference account for half of the cases, but video editing alone is fifteen of them, which is as many as image reference. Document-to-video, the feature no other model on this site has, gets five cases, and three of those are 30 seconds long. On apimodels.app all eight job types run through one model name, wan-3.0-video, on the same video endpoint as VEO, Kling and Seedance.
| Job type | Cases | Clip length | Resolution tier | What the prompt points at | API fields |
|---|---|---|---|---|---|
| Wan 3.0 · Text to video | 17 | 15–30 s (12 at 30 s) | 9 × 720P, 8 × 1080P | nothing but the prompt | prompt, duration, resolution, aspect_ratio |
| Wan 3.0 · Image reference | 15 | 8–30 s | 10 × 720P, 5 × 1080P | 图1 / 图2 … (up to 10 images) | reference_image_urls |
| Wan 3.0 · Video editing | 15 | 5–15 s | 14 × 720P, 1 × 1080P | 视频1 = the clip being edited | reference_video_urls |
| Wan 3.0 · Document / web page | 5 | 23–30 s | 5 × 720P | a DOCX, PPTX, XLSX or a public URL | file_url or link_url |
| Wan 3.0 · Audio reference | 3 | 15 s | 3 × 720P | 音频1 / 音频2 for voice or beat | reference_audio_urls (+ images / video) |
| Wan 3.0 · Video reference | 3 | 5 s | 3 × 720P | 视频1 for motion, camera move or VFX | reference_video_urls (+ image) |
| Wan 3.0 · Video extension | 3 | 10–20 s | 3 × 720P | 视频1 plus the direction to extend | reference_video_urls |
| Wan 3.0 · First frame | 3 | 15–30 s | 2 × 720P, 1 × 1080P | @Image1 as frame one | first_frame_url |
One case per job type, chosen so that each shows a different input the prompt has to address. Every prompt below is the publisher's original Chinese text, unedited; the clip next to it is the output the handbook published for that prompt, not something we regenerated. Under each clip you get the numbers you need to budget the same job on Wan 3.0 here: clip length, resolution tier, billed seconds (input plus output when a reference video is involved) and the cost at our standard tier. Prime mode, the faster tier, costs 51–56% more per second and delivered the same job in 58 seconds instead of 356 in our measurement.
Read the prompts for structure rather than vocabulary. The handbook writes long jobs as a timeline, addresses reference material by slot number, tags sound inline, and ends every edit with a sentence saying what must not change. Those four habits carry over to any Wan 3.0 job you write, in Chinese or in English.
A single 241-character prompt drives a full 30-second Wan 3.0 shot: a pale-green filtered hospital corridor under cold fluorescent light, low saturation, high contrast, the camera starting behind the nurse station and following a figure in dark green. Nothing is attached to the request; the look, the pace and the sound come entirely from the description. Note how much of the text is about colour and light rather than plot: that single sentence about the green filter decides the whole grade.
This is the most affordable kind of Wan 3.0 job to reproduce. Send the prompt with duration 30 and resolution 720P and you pay for 30 output seconds, $2.70 at the standard tier.
一段现代悬疑风格的电影感长镜头,整个画面呈现出强烈的淡绿色滤镜风格,营造出一种基于淡绿色系的清冷氛围。在散发着冰冷白光的荧光灯下,空无一人的医院走廊呈现出低饱和的色彩与极高对比的光影关系,锐利且非人情化。镜头从护士台后方开始,一名身穿墨绿色夹克的女子猛地从走廊尽头的门后冲入,沿着对称的中心线向镜头快速跑来,摄影机流畅地向后移动。她的脚步声急促而响亮,口中焦急地呼喊。最终,她冲到台前,双手用力拍在台面上,绝望地环顾四周,发出几近崩溃的求助呐喊,声音在只有微弱设备嗡鸣的寂静中回荡。The prompt opens with 参考 @Image1 的分镜图 — use the storyboard in image 1 — and then describes three shots: a wide view of a small rainy platform with a girl under a blue umbrella, a mid shot of her facing a boy with a backpack across the tracks, and a final shot through the rain. Wan 3.0 reads the storyboard for layout and staging and takes the beat structure from the text; the 20-second clip follows the three shots in order.
On the API this is one entry in reference_image_urls. Images cost nothing extra on Wan 3.0, so the bill is the output only: 20 seconds at 720P, $1.80.
参考 @Image1 的分镜图,镜头1:雨夜小火车站台全景,长发女孩撑蓝色雨伞独自站在站台上,身后小站房透出暖光,远处青山隐没在雨雾中;镜头2:中景,女孩撑伞与背书包的短发男孩在站台上面对面站立,大雨倾盆,铁轨在两人之间延伸;镜头3:透过雨滴滑落的透明伞面拍摄女孩面部特写,她眼眶泛红,轻声说{到了那边记得发消息};镜头4:从男孩背后越肩拍向女孩方向,远处一列火车亮着车头灯缓缓驶来,雨幕中光晕扩散;镜头5:双手特写,女孩将一张折叠的纸条递到男孩手中,指尖微微颤抖;镜头6:女孩独自站在站台边,目送铁轨延伸向远方,男孩已不在身旁,雨势渐弱。This case shows the slot grammar at its clearest. The prompt defines 图片1 as subject 1 and 图片2 as subject 2, assigns 音频1 as the voice of subject 1 and 音频2 as the voice of subject 2, places both by a treadmill, and writes the dialogue in curly braces: subject 1 turns and asks which muscle group today, subject 2 answers while towelling off. Two short audio clips, 5 and 6 seconds, carry the two voices; Wan 3.0 clones the timbre and lip-syncs the lines.
Images and audio ride in reference_image_urls and reference_audio_urls, and neither adds to the bill. Fifteen seconds at 720P: $1.35.
将图片1中的女人定义为主体1,将图片2中的女人定义为主体2,主体1台词音色参考音频1,主体2台词音色参考音频2,在健身房内,主体1和主体2并排站在跑步机旁,主体1转头微笑对主体2说{今天练哪个部位},主体2一边用毛巾擦汗一边回应{练背吧,昨天腿还酸着呢},主体1笑着点头说{行,我陪你,先热身十分钟},主体2拿起水杯喝一口回应{走吧},背景是器械区和落地镜,写实拍摄,画面明亮通透,冷白色自然饱和度色调A 5-second skateboarding clip is the motion source; a still image is the identity source. The prompt tells Wan 3.0 to replace the man in video 1 with the woman and her outfit from image 1, and then spends its remaining words on the lock: only the person changes, the skateboard action, the trajectory and the skate-park background stay exactly as they are. The result keeps the original timing frame for frame.
Because a reference video is involved, billing counts input plus output: 5 seconds in, 5 seconds out, 10 billed seconds at 720P, $0.90. Reference videos on Wan 3.0 can be 1–15 seconds each and 15 seconds in total across up to five clips.
参考视频,视频1 @Video 中的滑板男人替换为图片 @Image1 中的女人和衣服,注意:只替换人物,保留原视频中滑板动作、运动轨迹和滑板场背景完全保持不变。The input is an XLSX file of cross-border monthly GMV and the prompt is the longest of the eight, 2,677 characters, because it writes a 22-second edit as four numbered segments on a white background with soft colour blocks and clean line charts, and it insists on a background track that starts at second zero and never drops out. Wan 3.0 reads the figures from the spreadsheet and animates them; the prompt owns pacing, typography and sound.
Send the file as file_url (DOCX, PPTX, XLSX, PDF, TXT and MD are accepted, up to 50 pages and 100 MB) or a public page as link_url; the two are mutually exclusive. Documents cost nothing extra: 23 seconds at 720P is $2.07.
总时长22秒,4个片段顺序拼接。全片纯白底,柔色色块与简洁线条图表,Apple Keynote式极简商务风格,无衬线深灰字体。BGM(必须有,贯穿全片,音量清晰可闻):一段轻快的商务背景音乐,节奏稳定、旋律简洁,从第一秒开始就有音乐,全程不间断,旁白出现时音乐不消失、音量略低于人声。音效层:色块滑入时'嗖'的whoosh声、数字跳动时轻键盘'咔嗒'声、关键数据弹出时清脆的'叮'一声bell。旁白:英文男声或女声,明快自信,语速约2.5词/秒,像在季度会上做presentation的语气——专业但不冰冷,有节奏感,关键数字处微微加重。
H1-2026跨境电商月度GMV.xlsx
片段1:全景,中心构图,纯白底,柔光,极简商务风格,视觉参考参考图1的干净白底+柔和色块。纯白画面从中央展开,无数金色光粒子从画面四周向中心汇聚聚拢,粒子碰撞闪烁后凝结成标题文字'H1 2026 · Wan',文字表面泛着流动的数据流光效,边缘有微光粒子持续环绕飘散,整体炫酷而高级。停留片刻后标题化作粒子流散消失,下方同时滑入两组大号深灰数字——左侧是第7行B7单元格的值$1,580(1月GMV)、右侧是G7单元格的值$3,460(6月GMV),中间一个浅灰色箭头连接,箭头下方浮出浅灰小字'Monthly GMV Growth'。底部弹出一枚珊瑚色圆角标签,显示第8行H8单元格的值+45.7%(H1整体增幅),弹出时伴随清脆的'叮'一声。旁白明快自信地说道:'In the first half of 2026, monthly GMV grew from 15.8 to 34.6 million dollars.' BGM从画面第一秒起就清晰响起,轻快节奏贯穿,色块滑入时带轻微的'嗖'whoosh声。片段末尾色块向两侧平滑滑开,露出下一段画面。
片段2:全景,左侧Y轴+底部X轴构图,纯白底+浅灰点状网格,柔光,极简商务风格,视觉参考参考图2的趋势线图风格。画面底部X轴从左到右依次清晰标注六个月份标签:Jan、Feb、Mar、Apr、May、Jun,等距排列,浅灰色无衬线字体。左侧Y轴标注数值刻度。画面从左向右同时绘制三条趋势线——珊瑚色代表东南亚(数据取自第4行:B4=$580, C4=$640, D4=$750, E4=$880, F4=$1,050, G4=$1,280)、青蓝色代表北美(第5行:B5=$420, C5=$480, D5=$560, E5=$680, F5=$860, G5=$1,260)、琥珀色代表欧洲(第6行:B6=$580, C6=$610, D6=$680, E6=$750, F6=$830, G6=$920),三条线从Jan同一位置出发,随节奏向右上方延伸至Jun。线条下方各带30%透明度渐变填充。每个月份节点到达时,数据点轻弹一下,正上方浮出该点的精确数值标签——标签只显示美元数值,不显示百分比、不显示其他文字,数值必须与上述单元格完全一致。右上角滑出图例:三个色块配市场英文名。六个点全部到达后画面定格片刻。旁白节奏稳定地说道:'All three markets — Southeast Asia, North America, and Europe — showed consistent month-over-month growth, with no single dip.' BGM持续可闻,数字弹出时伴随细碎的'咔嗒'键盘音。6月终点标签放大填满画面后缩小重组,过渡至下一段。
片段3:全景,左60%柱状图+右40%数据列构图,纯白底,柔光,极简商务风格,视觉参考参考图1的柱状图风格。白底上从左到右依次升起6根青蓝色圆角柱体,分别对应1月至6月——必须完整展示6根柱子,柱体高度严格对应第5行North America各月数值:B5=$420, C5=$480, D5=$560, E5=$680, F5=$860, G5=$1,260,从低到高递增排列,呈现明显的逐月增长趋势。每根柱体升起时带轻微的'嗖'弹性音效,柱顶浮出对应的精确美元数值标签。柱体全部就位后,右侧纵向排列滑出5个北美逐月环比增长率,分别由相邻两月数值计算得出:Feb +14.3%(C5相对B5)、Mar +16.7%(D5相对C5)、Apr +21.4%(E5相对D5)、May +26.5%(F5相对E5)、Jun +46.5%(G5相对F5)。每个增长率显示为:翠绿色向上箭头图标 + 百分比数值(如↑+14.3%),整体呈阶梯状从上到下依次排列,绿色箭头强调增长态势,字体略大确保清晰可读。每个数字滑入时带一声细碎的'咔嗒'。最后,北美整体增幅用珊瑚色大号字体从中央弹出:+200%(由G5=$1,260相对B5=$420计算),伴随清脆的'叮'一声bell,停留片刻,数字周围散发轻微光晕。旁白在关键数据处微微加重语气:'North America was our growth engine — up two hundred percent in just six months, accelerating every single month.' BGM持续清晰,节奏逐渐推向高点。整体增幅数字居中放大,周围元素柔和淡出,过渡至尾段。
片段4:全景,中心构图偏上,纯白底,柔光,极简商务风格,视觉参考参考图3的环形图风格。白底中央浮现一个简约环形图(甜甜圈图),只有三段弧形,分别用珊瑚色、青蓝色、琥珀色渲染三个市场的H1占比——东南亚占比由H4=5180计算、北美由H5=4260计算、欧洲由H6=4370计算,三者总和H7=13810。三段弧形之间用白色间隙分隔,环形图整体干净简洁,图表上没有任何百分比符号、没有任何数字标注、没有任何散落的标签,只有纯粹的三段色弧。环形图以约10°/s的速度缓缓旋转。环形图正中央显示一个大号深灰数字,该数字从0快速跳动至H7单元格的值$138.1M(即13810万美元),跳动时伴随连续的细碎'咔嗒'声,最终数字落定时一声清脆的'叮'。下方淡入一行浅灰副标题'Three Markets · All Growing · Zero Downturn'。旁白沉稳收尾,说完留1秒静音:'138 million in total — all markets growing, zero downturns.' BGM持续可闻,最后缓缓衰减,1秒静音后全片结束。Same skate footage as the motion-reference case, different job. The prompt starts with 编辑视频 — edit the video — and turns the man in video 1 into a short-haired woman in a sports top and loose cargo trousers, while keeping the baseball cap and the pads from the original, the skateboard action, the trajectory and the skate park. The before and after play side by side above; only the subject changed. Fifteen of the 64 handbook cases are edits of this shape: add an element, replace one, delete one, relight, restyle, rewrite a line of dialogue, or reshape the plot.
Edits keep the input length, so a 5-second source returns a 5-second result and bills 10 seconds: $0.90 at 720P.
编辑视频,视频1中的滑板男人替换为一位短发女人,穿运动背心和宽松工装裤,保留原视频中的棒球帽、护膝护肘等护具,滑板动作、运动轨迹和滑板场背景完全保持不变。The prompt begins 将视频1延长15s — extend video 1 by 15 seconds — and then does what most people forget: it names the two characters from the existing footage by their clothes and props (the man in the dark grey Chinese gown with the prayer beads, the man in the dark gown), and writes the new beats with their dialogue, down to the line about when the goods reach port. Wan 3.0 continues the scene from the last frame with the same two people and the same room, and the new 15 seconds carry the written dialogue.
A 5-second source extended to a 20-second result bills 25 seconds (5 in, 20 out): $2.25 at 720P. Upstream caps input plus output at 30 seconds, which is why the handbook extends 5-second clips rather than 15-second ones.
将视频1延长15s。角色A为梳背头、穿深灰色中式长衫、左手腕佩戴棕色佛珠手串的男人,角色B为穿深色中式长衫的男人。角色B缓缓放下茶碗,目光微沉,用指尖轻叩桌面两下,不紧不慢说"价钱好讲,但货要几时到埠,你要畀个准数我。"角色A笑容不减,端起盖碗刮了刮茶沫。角色A呷一口茶,将茶碗搁回茶托,身体前倾压低声音,佛珠随手势轻晃,语气笃定"最迟月底,码头嗰边我已经打点好晒,一船货齐齐整整运到你仓库门口。"说罢右手掌心朝下往桌面一按,目光直视角色B。角色B沉吟片刻,视线扫向窗外"萬安當"牌匾,竹帘被街上微风轻轻掀起。他转回头,嘴角浮起笑意,拿起盖碗朝角色A微微一举,"好,月底交货,我等你好消息。"两人相视而笑。保持民国商战片的沉稳调性,对话节奏从容,茶馆环境音低沉衬底。A portrait 720×1280 case. The prompt is written like a shot list with timecodes: 0:00–0:03 mid shot, the swordsman's silhouette in the ink-wash bamboo grove from @Image1, hand resting on the hilt, leaves falling; then the next beats escalate. Because the image is passed as the first frame rather than as a reference, Wan 3.0 treats it as frame one and animates out of it, and everything the prompt describes is motion away from that picture.
first_frame_url cannot be combined with reference_* fields, file_url or link_url on Wan 3.0; the request is rejected before it runs. Twenty seconds at 720P: $1.80.
(0:00 - 0:03) 运镜:中景。
画面:沿用 @Image1 的水墨竹林场景。头戴斗笠、身着墨色长袍的刀客,其剪影伫立在浓雾环绕、竹影婆娑的中心。
动作:刀客左手轻按刀柄,竹叶随风飘落。
氛围:寂静,压抑,水墨晕染的颗粒感清晰可见。
(0:03 - 0:08) 剧情引入,运镜:特写。
剧情:刀客察觉杀气。
画面:镜头快速推进到刀客斗笠下的特写。
细节:斗笠边缘露出一双凌厉、决绝的眼睛,眼神如刀。水墨晕染的“汗珠”从额头滑落。
动作:刀客右手缓缓握紧刀柄,指关节泛白。竹叶飘落的速度加快。
(0:08 - 0:13) 运镜:快速环绕。
画面:镜头以刀客为中心快速环绕。
剧情:两名同样是水墨剪影的刺客从竹林深处杀出,手持双钩和短刃。
细节:刺客的动作也是水墨笔触构成的,如墨迹流动。镜头快速掠过刺客的身影,强调速度感。
(0:13 - 0:18) 运镜:手持感,高速捕捉。
打斗:刀客猛然拔刀,刀光是一道纯白色的水墨裂痕。刺客双钩攻来。
细节:
一击:刀客回身一斩,刀刃与刺客双钩碰撞,迸发出无数黑白相间、带有书法感的火花和墨点。
二击:刀客虚晃一枪,利用竹竿弹射,从高处劈下。刺客翻滚躲避。竹竿被打断,化作飞散的墨痕。
三击:刀客刀尖直指另一刺客咽喉,刺客以短刃格挡,镜头捕捉刀刃与短刃摩擦时细微的墨迹崩裂细节。
**(0:18 - 0:20) 运镜:慢镜头至定格。**
* **画面:** 刀客的刀停在刺客颈侧,刺客凝固。
* **细节:** 刀客的斗笠在打斗中略微倾斜。竹林归于平静,更多竹叶飞舞。
* **氛围:** 镜头慢慢拉远,回到中景,定格在刀客收刀、刺客倒地的画面。
* **文字(可选):** 画面下方显现水墨风文字:“江湖,不过一笔勾销。”
**风格说明:** 整个过程保持 `image_0.png` 的高度黑白、水墨画风格,动作必须有毛笔触感,速度要快,细节要足(墨滴、火花、墨痕裂口)。The 64 prompts range from 18 to 3,423 characters, and the long ones are not padded: they are timelines. Below are the eight habits that recur across the set, each with a snippet quoted from a real case so you can see the exact wording Wan 3.0 responded to. The snippets are Chinese because the handbook is; the structure is language-independent and works in English on the same model.
Two of these habits matter more than the rest. Numbering the material — 图1, 视频1, 音频1 — is how Wan 3.0 knows which file plays which role when several are attached, and the closing lock — the sentence that says what must stay unchanged — is what separates a clean edit from a re-imagined frame.
| Habit | What it does for Wan 3.0 | Quoted from a case |
|---|---|---|
| Wan 3.0 · Number the material | reference images, videos and audio are addressed by slot, counted separately | 将图片1中的女人定义为主体1,将图片2中的女人定义为主体2,主体1台词音色参考音频1 |
| Wan 3.0 · Write a timeline | any job over 10 seconds is split into timed segments the model follows in order | 镜头1 [0秒-30秒] 近景,平视视角,连续快速摇镜,一镜到底 |
| Wan 3.0 · State length, ratio and frame rate up front | the first sentence carries the delivery spec and matches the duration parameter | 生成一支10秒、16:9横版、1920×1080、30fps、单镜头连续完成的微缩奇幻高速追逐短片 |
| Wan 3.0 · Tag sound inline | effects in angle brackets, dialogue in braces or quotes; Wan 3.0 renders both | 用抹布仔细擦拭吧台<抹布摩擦声> |
| Wan 3.0 · Lock what must not change | every edit ends with the elements that stay as they are | 女人的动作和画面其余部分保持不变 |
| Wan 3.0 · Name characters by their clothes | extensions and edits re-identify people from the footage before adding beats | 角色A为梳背头、穿深灰色中式长衫、左手腕佩戴棕色佛珠手串的男人 |
| Wan 3.0 · Give references a single job | a clip is named as the source of motion, camera move or effect, not everything | 参考视频的运镜方式,场景更改为一杯咖啡放在天台咖啡桌上 |
| Wan 3.0 · Describe motion out of a still | first-frame prompts write what moves away from the picture, with timecodes | (0:00 - 0:03) 运镜:中景。 画面:沿用 @Image1 的水墨竹林场景 |
Twelve of the seventeen text-to-video cases run the full 30 seconds, and they share a skeleton. A one-line brief states genre, length and look. A palette-and-light sentence fixes the grade. Then the timeline: numbered shots or second ranges, each with camera, action and the sound that plays under it. Long cases add a continuity paragraph, listing the character's face, hair and wardrobe once so every segment reuses it. The order is consistent across the handbook: spec and look first, plot after, so a reader (or the model) gets the delivery constraints before the story.
The same skeleton scales down. A 5-second edit is a brief, an action and a lock. A 15-second product clip is a brief, a light sentence and three timed beats. What never appears in the handbook is a wall of adjectives with no time structure — that is the prompt that reads as sluggish because Wan 3.0 spreads the action evenly across the clip when nobody tells it when things happen.
| Block | Purpose | Seen in the cases as |
|---|---|---|
| Wan 3.0 · Brief | genre, length, ratio, look in one line | 这是一段30秒的真实电影质感海上巨兽登场场景,整体氛围参考好莱坞灾难怪兽大片 |
| Wan 3.0 · Palette and light | one sentence that decides the grade | 整个画面呈现出强烈的淡绿色滤镜风格,营造出一种基于淡绿色系的清冷氛围 |
| Wan 3.0 · Continuity sheet | face, hair, wardrobe stated once, reused by every segment | 严格保持男性角色的黑发高发髻、鬓角碎发被风吹起、锐利深色瞳 |
| Wan 3.0 · Timeline | numbered shots or second ranges with camera, action, sound | 分段1[0-5秒]:一个连续的、照片级写实的电影镜头 |
| Wan 3.0 · Sound | ambience, effects, dialogue placed where they happen | BGM(必须有,贯穿全片,音量清晰可闻) |
| Wan 3.0 · Lock | for edits and extensions: what stays unchanged | 滑板动作、运动轨迹和滑板场背景完全保持不变 |
Every case above maps to one request against POST /api/v1/video/generations with model wan-3.0-video. The prompt goes in as written; the material the prompt numbers goes into the matching array in the same order, because 图1 is the first entry of reference_image_urls and 视频1 the first of reference_video_urls. Wan 3.0 has two input modes that do not mix: frame mode (first_frame_url, optionally last_frame_url) and reference mode (reference_image_urls, reference_video_urls, reference_audio_urls, file_url or link_url). We reject a mixed request locally and name both sides, so it never reaches upstream.
Two parameters change what you get without changing the prompt. mode: prime switches to the fast tier at 51–56% more per second, identical capabilities. prompt_extend, on by default, lets Wan 3.0 rewrite short prompts; the handbook prompts are long enough that you can turn it off and get exactly what you wrote.
| Case type | Fields to send | Limits on Wan 3.0 |
|---|---|---|
| Wan 3.0 · Text to video | prompt, duration 2–30 (or -1 to let the model choose), resolution 480P / 720P / 1080P, aspect_ratio | single shot up to 30 s, 30 fps, audio on by default |
| Wan 3.0 · First frame / first + last | first_frame_url, last_frame_url | exclusive with every reference field |
| Wan 3.0 · Image reference | reference_image_urls | up to 10 images, no extra charge |
| Wan 3.0 · Video reference, edit, extend | reference_video_urls | up to 5 clips, 1–15 s each, 15 s total; billed input + output; input + output ≤ 30 s |
| Wan 3.0 · Audio reference | reference_audio_urls | up to 5 clips, 15 s total, wav or mp3, no extra charge |
| Wan 3.0 · Document / web page | file_url or link_url (one of the two) | DOCX, PPTX, XLSX, PDF, TXT, MD… ≤ 50 pages, ≤ 100 MB; public URL only |
| Wan 3.0 · Speed tier | mode: standard (default) or prime | prime is 51–56% more per second, same capabilities |
Wan 3.0 is billed per output second, by resolution: $0.045 at 480P, $0.09 at 720P, $0.18 at 1080P on the standard tier; prime is $0.068, $0.14 and $0.28. Reference images, audio, documents and web pages are free; a reference video adds its own length to the billed seconds. Only successful jobs are charged. The table prices each showcased case at the resolution its clip was actually delivered in, so a 720P clip is quoted at 720P even where the handbook author could have asked for 1080P.
For planning: the whole 64-case set is 1,112 output seconds plus 147 input seconds from the 22 jobs that carried a reference video, 1,258 billed seconds in all, which comes to about $113 at the standard 720P rate. That is the cost of reproducing an entire launch handbook on Wan 3.0.
| Case | Output | Billed seconds | Standard tier | Prime tier |
|---|---|---|---|---|
| Wan 3.0 · Hospital corridor, text to video | 30 s · 720P | 30 | $2.70 | $4.20 |
| Wan 3.0 · Storyboard farewell, image reference | 20 s · 720P | 20 | $1.80 | $2.80 |
| Wan 3.0 · Gym conversation, audio reference | 15 s · 720P | 15 | $1.35 | $2.10 |
| Wan 3.0 · Skater swap, video reference | 5 s · 720P (5 s in) | 10 | $0.90 | $1.40 |
| Wan 3.0 · GMV spreadsheet, document to video | 23 s · 720P | 23 | $2.07 | $3.22 |
| Wan 3.0 · Skater to woman, video editing | 5 s · 720P (5 s in) | 10 | $0.90 | $1.40 |
| Wan 3.0 · Tea-table negotiation, extension | 20 s · 720P (5 s in) | 25 | $2.25 | $3.50 |
| Wan 3.0 · Ink-wash swordsman, first frame | 20 s · 720P | 20 | $1.80 | $2.80 |
Thirty seconds is a hard ceiling per job, and with a reference video it is a shared one: input plus output may not exceed 30 seconds, which is why every extension in the handbook starts from a 5-second clip. Reference video is capped at 15 seconds total across at most five clips, reference audio at 15 seconds total, images at ten, documents at 50 pages. Frame mode and reference mode cannot be combined, and file_url and link_url cannot both be sent. None of these are our limits; they are Wan 3.0's, and we enforce them before the request leaves so a rejected job never costs a round trip.
What the handbook does not show is as informative as what it does. No case uses a photograph of an identifiable real person as a reference, and no case relies on text rendering inside the frame beyond a logo or a caption already present in the input. Treat both as unverified on Wan 3.0 rather than as supported. The reference images and documents behind the reference cases were not published with the handbook, so those cards carry the prompt and the result only.
Three models on apimodels.app take reference video and audio and generate 15 seconds or more in one shot. The figures are the ones each model page states; prices are our per-second list prices at the tier named.
| Wan 3.0 | Seedance 2.5 | MiniMax H3 | |
|---|---|---|---|
| Single-shot length | 2–30 s | 4–30 s | 5–15 s |
| Reference inputs per job | 10 images + 5 videos + 5 audio, or a document / web page | 30 images + 10 videos + 10 audio | 9 images + 3 videos + 3 audio |
| Resolutions | 480P / 720P / 1080P | 480p / 720p | 768P / 2K |
| Price per second here | $0.045 / $0.09 / $0.18 | $0.120 / $0.270 | $0.097 / $0.145 |
| Video edit and extend | yes, via reference video | yes, dedicated task types | generative editing |
| Document or web page as input | yes | no | no |
| Native synced audio | yes | yes | yes |
Wan 3.0 Video is a Video Generation API provided by Alibaba. Wan 3.0 is Alibaba's all-in-one video model: a single model name that covers text-to-video, first-frame and first+last-frame image-to-video, and omni-reference generation driven by up to 10 reference images, 5 reference video clips (15s total) and 5 audio clips — or a document (PPTX, DOCX, XLSX, PDF, TXT, MD, up to 50 pages) or a public web page. Output is native single-shot up to 30 seconds at 30fps with a synced audio track you can switch off at no price difference, in 480P / 720P / 1080P and six aspect ratios including an adaptive mode that picks the ratio from your inputs. In omni-reference mode the prompt can address materials by position — "figure 1 holds the object from figure 2, walking past video 1" — which is how you keep several characters, props and environments consistent across a shot. Billing is per second at 480p $0.045/s, 720p $0.09/s, 1080p $0.18/s. Reference images, audio, documents and web pages cost nothing extra, but a reference VIDEO (video edit, video extend, character swap) is billed on its own duration too — input video seconds + output seconds, matching the upstream, which also caps those jobs at input + output ≤ 30 seconds. So a 5-second 720p clip is $0.45 and a 30-second 1080p is $5.40; pass duration -1 and the model picks a length between 2 and 30 seconds, in which case we hold the 30-second worst case and refund the difference once the real length is known. Measured on this platform: a 480p 5s clip returned in 356 seconds on this standard tier — reach for Wan 3.0 Video Prime if that matters, it did the identical job in 58 seconds. Async create-then-poll on the shared video endpoint; failed tasks are never billed. We do not add a content-moderation layer of our own: prompts and reference media pass through as sent. The model's own safety systems still apply upstream and do fire in practice — both inputs and outputs are screened, and real human faces in reference material are commonly refused — so this is not an unfiltered endpoint. Reviewing what you generate, and complying with the law and your own platform's rules, is your responsibility; use only with subjects who have consented. Through APIMODELS platform, you can access this model via a unified API with transparent pay-as-you-go pricing. Current pricing: 480p: $0.045, 720p: $0.09, 1080p: $0.18, 480p-prime: $0.068, 720p-prime: $0.14, 1080p-prime: $0.28.
Quickly generate brand promotion videos for ad campaigns and social media marketing.
Create compelling short-form video content for platforms like TikTok, Instagram, and YouTube.
Generate product feature demonstrations and tutorials to improve user conversion.
Produce course explanations, knowledge explainers, and training videos at an affordable price.
Wan 3.0 Video is available through APIMODELS at: 480p: $0.045, 720p: $0.09, 1080p: $0.18, 480p-prime: $0.068, 720p-prime: $0.14, 1080p-prime: $0.28. Billing is pay-as-you-go — you only pay for what you generate.
Sign up at APIMODELS, get your API key, and call our unified API endpoint. We provide detailed API documentation with code examples in cURL, Python, and Node.js.
APIMODELS offers the same Wan 3.0 Video model through our aggregation platform. We provide a unified API interface so you do not need separate accounts for each provider - one API key to access all models.
Wan 3.0 (万相 3.0) is Alibaba's all-in-one video model, released August 2026. One model name covers text-to-video, first-frame and first+last-frame image-to-video, and "omni-reference" generation driven by up to 10 reference images, 5 reference video clips and 5 audio clips (15 seconds total each for video and audio) — or a document (PPTX, DOCX, XLSX, PDF, TXT, MD, up to 50 pages) or a public web page. On apimodels.app you POST to /api/v1/video/generations with model "wan-3.0-video", then poll the same endpoint with the returned task_id. Same request shape as VEO, Kling and Seedance here, so switching models is a one-string change and one API key covers all of them.
Because most people cannot open the account. Wan 3.0 is currently served only from Alibaba Model Studio's China (Beijing) and Singapore regions — we probed US/Virginia and it does not carry the model at all (that workspace lists 92 models and not one of them is a video model). The China station requires mainland Chinese real-name verification to register, which is where non-Chinese developers stop. Going through apimodels.app needs no Alibaba Cloud account, no real-name check and no separate billing relationship: same key and same endpoint as every other model you already call here.
Billed per second: 480p $0.045/s, 720p $0.09/s, 1080p $0.18/s. Reference images, audio, documents and web pages are free, but a reference VIDEO (video edit, extend, character swap) is billed on its own duration too — input video seconds + output seconds, matching the upstream, which also caps those jobs at input + output ≤ 30 seconds. So with no input video a 5-second 720p clip is about $0.45, a 10-second 1080p about $1.80, and a 30-second 1080p about $5.40; feed in a 10-second reference video and that 10-second 1080p bills as 20 seconds, about $3.60. Turning the audio track off does not change the price. Failed tasks are not billed. Passing mode:"prime" switches to the high-speed tier at 51–56% more per second ($0.068 / $0.14 / $0.28) — it buys latency, not quality.
Yes — duration takes any integer from 2 to 30 and the output is a native single shot at 30fps, not stitched segments. Passing -1 selects smart-duration mode, where the model picks a length between 2 and 30 seconds from your prompt and materials. For -1 we hold the 30-second worst case up front and then charge the real output length once upstream reports it, refunding the difference — so a model that decides on 5 seconds never costs you 30. Note that when you supply reference video, input video length + output length must stay within 30 seconds.
Send the materials in order and the prompt can address them positionally: "figure 1", "figure 2", "video 1", "audio 1" — images, videos and audio are numbered independently. For example: "the person in video 1 holds the object from figure 3 and plays guitar on the chair from figure 4." That is how you keep several characters, props and environments consistent inside one shot. One hard rule: reference materials (reference_image / video / audio, plus document and web page) are mutually exclusive with first_frame / last_frame. Mixing them makes upstream fail the task, so we reject it locally first and tell you which two sides collided — you never burn a round trip on it.
Pass the public URL of a PPTX, DOCX, XLSX, PDF, TXT or Markdown file (under 50 pages and 100MB) as file_url, and the model reads the content and generates video from it — you do not have to flatten it into a prompt first. Launch decks, courseware and reports go straight into the video workflow. Web pages work the same way through link_url, for publicly readable pages that need no login. file_url and link_url are mutually exclusive; send one or the other.
A synced audio track (dialogue, sound effects, ambience) is generated by default; pass audio: false to drop it, at no price difference. Resolutions are 480P / 720P / 1080P; aspect ratios are adaptive (the default — inferred from your inputs and intent), 16:9, 4:3, 1:1, 3:4 and 9:16. One deliberate difference from upstream: if you omit resolution we default to 720P rather than Alibaba's 1080P, because 1080P is double the price and nobody should land on the priciest tier by not typing a parameter. Ask for 1080P explicitly and you get it.
Measured on the standard tier: a 480P 5-second clip came back in about 356 seconds (Alibaba quotes 1-5 minutes). With mode:"prime" the identical prompt and parameters returned in 58 seconds — roughly 6x faster. It is an async create-then-poll endpoint; poll every 10-15 seconds, or pass callback_url and we POST you when the task finishes. Result files are kept for 30 days and then deleted automatically, so download or re-host anything you need to keep.
On APIMODELS, Wan 3.0 Video runs alongside 60+ models on one API key and one balance, so choosing is about fit, not lock-in. It supports Text to Video, First / Last Frame, Omni Reference, Doc & Web to Video, Up to 30s @ 30fps, Native Audio, and you can weigh it on price and capability against other Video Generation models, then switch by changing a single model-name string — no new account or integration. Browse every Video Generation option with live pricing at apimodels.app/models.
Wan 3.0 Video supports: Text to Video, First / Last Frame, Omni Reference, Doc & Web to Video, Up to 30s @ 30fps, Native Audio. See the APIMODELS docs for full parameters and call examples.
Yes. APIMODELS exposes Wan 3.0 Video through a single unified API and one key — no separate provider accounts, and no need to handle each provider's regional network access yourself.
We support Stripe (Visa, Mastercard, and other international cards) and Alipay. Credits are available instantly after payment.
Prompts shared by their authors — copy and adapt them. Each one credits its author and links back to the original post.

Graduation-gown close-up: a worried monologue to camera
这是一个正面的近景镜头,画面主要呈现一位年轻黑人女孩的头部和上半身,她穿着一件蓝色的毕业袍。镜头稍微偏一点点,可以看到左边还有一个人,但背景被虚化处理了,看起来很模糊。女孩正对着镜头说话,嘴巴在动,眉头稍微皱着,表情看起来有点严肃和忧虑。光线从上方照下来,把她脸上的轮廓衬托得很自然,甚至能看到皮肤真实的质感。镜头虽然固定,但带着一点点人手持拍摄时的轻微呼吸感和晃动,视线从看着前面慢慢变成了低头看手里。
by Wan 3.0 视频创作者手册 · 通义万相

Suspense long take down an empty hospital corridor
一段现代悬疑风格的电影感长镜头,整个画面呈现出强烈的淡绿色滤镜风格,营造出一种基于淡绿色系的清冷氛围。在散发着冰冷白光的荧光灯下,空无一人的医院走廊呈现出低饱和的色彩与极高对比的光影关系,锐利且非人情化。镜头从护士台后方开始,一名身穿墨绿色夹克的女子猛地从走廊尽头的门后冲入,沿着对称的中心线向镜头快速跑来,摄影机流畅地向后移动。她的脚步声急促而响亮,口中焦急地呼喊。最终,她冲到台前,双手用力拍在台面上,绝望地环顾四周,发出几近崩溃的求助呐喊,声音在只有微弱设备嗡鸣的寂静中回荡。
by Wan 3.0 视频创作者手册 · 通义万相

One-take comedy: director versus leading lady on set
镜头1 [0秒-30秒] 近景,平视视角,连续快速摇镜,一镜到底(单镜头连续无缝运镜): 画面左侧是身穿黑色潮流导演马甲、戴着高档专业工作耳麦的导演,其左下方放置着一台巨大的专业电影摄像机。右侧是妆容精致、留着时尚韩式逗号刘海、身穿高定制西装的关系户小鲜肉。背景是宽敞的专业影棚拍摄现场,可见巨大的绿色幕布、数个高耸的影视级C型灯架、缠绕的电线,以及在背景中穿梭忙碌的灯光师和化妆师等工作人员。 专业影棚人工光。柔和的LED柔光箱光源从右侧洒入,将三位主角的面部均匀照亮,场景色彩饱满。背景的摄影器材与绿幕在充足的影棚灯光下投射出淡淡的阴影,呈现出真实而忙碌的现代化大制作片场质感。 镜头手持拍摄一镜到底,近景,镜头一开始,拍摄导演,导演无所谓的表情,瞥了一眼画面右侧斜前方画外的一眼关系户小鲜肉,对他下达指令,说道:“媒体故意黑你,做个表情看看”。说完后镜头快速右摇到关系户面部近景,镜头不要切镜,关系户小鲜肉,面部朝向画面左侧斜前方画外导演方向,神色自若,眼神真挚而自信。他嘴唇微动,神情专业地向对方阐述到““就这种事情来说,情绪可以有好几种”,说话时,头部伴有极其轻微的自然晃动,说完后镜头快速左摇到导演面部近景,镜头不要切镜,导演思索了一下,一边挑衅似地歪了歪头,向关系户小鲜肉说到““狗仔爆料你深夜幽会”,导演说完台词后,镜头迅速右摇到小鲜肉面部近景,不要切镜,关系户小鲜肉,面部朝向画面左侧斜前方画外导演方向,并快速进入表演状态,他眉头紧锁,眼神慌乱、焦急地快速游移,随后他狠狠地咬紧下唇,闭上双眼,整张脸(虽带着精致妆容)因极度的焦虑而紧绷,完美呈现出面对偶像生涯毁灭时,内心惊恐、紧张交织的挣扎状态。表演完后, 镜头再次迅速左摇到导演近景,不要切镜,导演面部表情微变,语速加快,说到““媒体又帮你洗白了,说那是你亲姐!”导演说完后,镜头再次迅速右摇到小鲜肉近景,镜头不要切镜,镜头中,关系户小鲜肉,面部朝向画面左侧斜前方画外导演方向,他的表情在千分之一秒内完成神级转换。他先是面部一滞,随后双眼猛地眯起,嘴角极大程度地咧开,露出一排白皙的牙齿,展露出一个极其夸张、甚至有些变形的狂喜笑容。这种“星途尽毁的绝望”与“彻底摆脱束缚”交织的疯魔情绪在他脸上扭曲呈现。角色表演完后,镜头再次迅速左摇到导演近景,镜头不要切镜,导演说到“你亲姐向媒体透露,你是领养的”,导演说完后,镜头再次迅速右摇到小鲜肉近景。镜头不要切镜,镜头内,关系户小鲜肉面部朝向画面左侧斜前方画外导演方向,只见他双眼暴睁,眼球凸出,精心修剪的眉毛高高扬起,嘴巴震惊地张成“O”形。紧接着,他的嘴角再次无法克制地向上拉扯,露出一副不可思议、惊魂未定却的复杂神态。 角色表演完后,镜头再次快速左摇到导演近景,镜头不要切镜,紧接着导演露出有点惊讶的表情说到“其实你是千亿豪门真少爷”。导演说完,镜头再次随即迅速右摇到小鲜肉近景,镜头不要切镜,镜头内,关系户小鲜肉面部朝向画面左侧斜前方画外导演方向,他瞬间爆发出极度魔性的狂喜。他整个人激动得微微颤抖,双眼笑得眯成了一条缝,双手甚至抬起至胸前,双手小幅度快速连续鼓掌,身体前倾摇晃,将一夜暴富、无法自持的癫狂喜悦演绎得淋漓尽致。 角色表演完后,镜头再次快速左摇到导演近景,镜头不要切镜,导演说到“你姐卷款跑路”导演说完,镜头再次迅速右摇到小鲜肉近景。小鲜肉面部朝向画面左侧斜前方画外导演方向,只见他的双眼再次瞪大,瞳孔颤抖,嘴角向下耷拉,神情在一秒钟之内从天堂坠入地狱,写满了惊愕、痛苦与不知所措。画面最后定格在这个表情上。
by Wan 3.0 视频创作者手册 · 通义万相
We curate copy-ready prompt libraries — every entry shows its full text and a sample result, ready to adapt.
How to get access, regional availability, and how this model compares with its alternatives.