
wan-3.0-videoWan 3.0 is Alibaba's all-in-one video model: a single model name that covers text-to-video, first-frame and first+last-frame image-to-video, and omni-reference generation driven by up to 10 reference images, 5 reference video clips (15s total) and 5 audio clips — or a document (PPTX, DOCX, XLSX, PDF, TXT, MD, up to 50 pages) or a public web page. Output is native single-shot up to 30 seconds at 30fps with a synced audio track you can switch off at no price difference, in 480P / 720P / 1080P and six aspect ratios including an adaptive mode that picks the ratio from your inputs. In omni-reference mode the prompt can address materials by position — "figure 1 holds the object from figure 2, walking past video 1" — which is how you keep several characters, props and environments consistent across a shot. Billing is per OUTPUT second only: reference images, video and audio you feed in cost nothing. 480p $0.0596/s, 720p $0.1191/s, 1080p $0.2382/s, so a 5-second 720p clip is $0.60 and a 30-second 1080p is $7.15; pass duration -1 and the model picks a length between 2 and 30 seconds, in which case we hold the 30-second worst case and refund the difference once the real length is known. Measured on this platform: a 480p 5s clip returned in 356 seconds on this standard tier — reach for Wan 3.0 Video Prime if that matters, it did the identical job in 58 seconds. Async create-then-poll on the shared video endpoint; failed tasks are never billed.
Single-shot up to 30s at 30fps — or pass duration -1 and let the model choose
Up to 10 images + 5 videos + 5 audio clips; the prompt addresses them as "figure 1", "video 1"
Feed a PPTX / DOCX / PDF or a public URL and the model reads it into the shot
From $0.0596/s at 480p; every reference file you send in is free
Prompt only
Document-to-video (PPTX / DOCX / XLSX / PDF) and web-page-to-video are API-only for now — see /docs/wan-3-0-video.
Generated video will appear here
Provide URLs and click Generate
Wan 3.0 Video は Alibaba の動画生成 API です。Wan 3.0(万相 3.0)はアリババのオールインワン動画モデルです。モデル名 1 つで text-to-video、先頭フレーム/先頭+末尾フレームの image-to-video、そして最大 10 枚の参照画像・5 本の参照動画・5 本の参照音声を使うオムニ参照生成までカバーし、PPTX / DOCX / XLSX / PDF などの文書(50 ページまで)や公開ウェブページをそのまま入力にすることもできます。出力はワンショットで最長 30 秒・30fps、同期した音声トラック付き(オフにしても価格は同じ)。解像度は 480P / 720P / 1080P、アスペクト比は 6 種類です。オムニ参照モードではプロンプトから素材を位置で指し示せる(「図1 が 図2 の物を持つ」)ため、複数のキャラクターや小道具を 1 カット内で一貫させられます。課金は【出力秒数】のみで、送り込む参照素材は無料です。480p $0.0596/秒、720p $0.1191/秒、1080p $0.2382/秒。 APIMODELS のプラットフォーム経由なら、統一 API と明朗な従量課金でこのモデルを呼び出せます。 現在の料金: 480p: $0.0596, 720p: $0.1191, 1080p: $0.2382。
広告キャンペーンや SNS 施策に向けたブランド動画を短時間で作ります。
TikTok、Instagram、YouTube 向けの縦型ショート動画を量産できます。
機能紹介やチュートリアル動画を作り、コンバージョンにつなげます。
講座の解説、知識の説明、研修用の動画を低コストで継続的に作れます。
Wan 3.0 Video は APIMODELS 経由で 480p: $0.0596, 720p: $0.1191, 1080p: $0.2382 で利用できます。課金は従量制で、生成した分だけの支払いです。
APIMODELS に登録して API キーを取得し、統一エンドポイントを呼ぶだけです。cURL / Python / Node.js のサンプルを含む詳しいドキュメントを用意しています。
APIMODELS は同じ Wan 3.0 Video を集約プラットフォーム経由で提供します。統一された API インターフェースなので、プロバイダごとにアカウントを作る必要はありません。キー 1 本ですべてのモデルに届きます。
Wan 3.0 (万相 3.0) is Alibaba's all-in-one video model, released August 2026. One model name covers text-to-video, first-frame and first+last-frame image-to-video, and "omni-reference" generation driven by up to 10 reference images, 5 reference video clips and 5 audio clips (15 seconds total each for video and audio) — or a document (PPTX, DOCX, XLSX, PDF, TXT, MD, up to 50 pages) or a public web page. On apimodels.app you POST to /api/v1/video/generations with model "wan-3.0-video", then poll the same endpoint with the returned task_id. Same request shape as VEO, Kling and Seedance here, so switching models is a one-string change and one API key covers all of them.
Because most people cannot open the account. Wan 3.0 is currently served only from Alibaba Model Studio's China (Beijing) and Singapore regions — we probed US/Virginia and it does not carry the model at all (that workspace lists 92 models and not one of them is a video model). The China station requires mainland Chinese real-name verification to register, which is where non-Chinese developers stop. Going through apimodels.app needs no Alibaba Cloud account, no real-name check and no separate billing relationship: same key and same endpoint as every other model you already call here.
Billed per OUTPUT second — reference images, video, audio and documents you send in are free: 480p $0.0596/s, 720p $0.1191/s, 1080p $0.2382/s. So a 5-second 720p clip is about $0.60, a 10-second 1080p about $2.38, and a 30-second 1080p about $7.15. Turning the audio track off does not change the price. Failed tasks are not billed. The high-speed tier, wan-3.0-video-prime, costs 50% more per second (from $0.0893/s at 480p) and buys latency, not quality.
Yes — duration takes any integer from 2 to 30 and the output is a native single shot at 30fps, not stitched segments. Passing -1 selects smart-duration mode, where the model picks a length between 2 and 30 seconds from your prompt and materials. For -1 we hold the 30-second worst case up front and then charge the real output length once upstream reports it, refunding the difference — so a model that decides on 5 seconds never costs you 30. Note that when you supply reference video, input video length + output length must stay within 30 seconds.
Send the materials in order and the prompt can address them positionally: "figure 1", "figure 2", "video 1", "audio 1" — images, videos and audio are numbered independently. For example: "the person in video 1 holds the object from figure 3 and plays guitar on the chair from figure 4." That is how you keep several characters, props and environments consistent inside one shot. One hard rule: reference materials (reference_image / video / audio, plus document and web page) are mutually exclusive with first_frame / last_frame. Mixing them makes upstream fail the task, so we reject it locally first and tell you which two sides collided — you never burn a round trip on it.
Pass the public URL of a PPTX, DOCX, XLSX, PDF, TXT or Markdown file (under 50 pages and 100MB) as file_url, and the model reads the content and generates video from it — you do not have to flatten it into a prompt first. Launch decks, courseware and reports go straight into the video workflow. Web pages work the same way through link_url, for publicly readable pages that need no login. file_url and link_url are mutually exclusive; send one or the other.
A synced audio track (dialogue, sound effects, ambience) is generated by default; pass audio: false to drop it, at no price difference. Resolutions are 480P / 720P / 1080P; aspect ratios are adaptive (the default — inferred from your inputs and intent), 16:9, 4:3, 1:1, 3:4 and 9:16. One deliberate difference from upstream: if you omit resolution we default to 720P rather than Alibaba's 1080P, because 1080P is double the price and nobody should land on the priciest tier by not typing a parameter. Ask for 1080P explicitly and you get it.
Measured on the standard tier: a 480P 5-second clip came back in about 356 seconds (Alibaba quotes 1-5 minutes). The high-speed tier, wan-3.0-video-prime, returned the identical prompt and parameters in 58 seconds — roughly 6x faster. It is an async create-then-poll endpoint; poll every 10-15 seconds, or pass callback_url and we POST you when the task finishes. Result files are kept for 7 days and then deleted automatically, so download or re-host anything you need to keep.
APIMODELS では Wan 3.0 Video が 60 以上のモデルと同じ API キー・同じ残高の上に並んでいるので、選択は「相性」の問題であって「囲い込み」の問題ではありません。Text to Video、First / Last Frame、Omni Reference、Doc & Web to Video、Up to 30s @ 30fps、Native Audio に対応しており、他の動画生成モデルと価格・機能を並べて評価できます。乗り換えはモデル名の文字列を 1 つ書き換えるだけ。新しいアカウントも追加の実装も要りません。動画生成の選択肢と最新価格は apimodels.app/models で確認できます。
Wan 3.0 Video は次に対応しています: Text to Video、First / Last Frame、Omni Reference、Doc & Web to Video、Up to 30s @ 30fps、Native Audio。パラメータの全一覧と呼び出し例は APIMODELS のドキュメントをご覧ください。
はい。APIMODELS は Wan 3.0 Video を単一の統一 API とキー 1 本で提供します。プロバイダごとのアカウントも、各社の地域ごとのネットワーク経路を自分で面倒みる必要もありません。
Stripe(Visa、Mastercard などの国際カード)と Alipay に対応しています。支払い後、残高はすぐ反映されます。