
grok-imagine-video-1.5Grok Imagine Video 1.5 は xAI の動画モデルで、競合の多くがまだ持っていない特徴があります —— 映像と一緒に同期した音声トラック(セリフ、効果音、環境音)を生成するため、ダウンロードした MP4 が最初から H.264 映像 + AAC 音声で完結し、別途アフレコ工程が要りません。モデル名 1 つで text-to-video と image-to-video(参照画像は最大 7 枚、URL でも base64 でも可)をカバーし、尺は 1〜15 秒。秒課金は 3 段階で、480p $0.0294/s、720p $0.0529/s、1080p $0.0882/s。4 秒のドラフトなら約 $0.12 です。出力にウォーターマークはなく、速度も実測で 4 秒クリップが 30〜70 秒ほどで返ってきます。作成と照会に分かれた非同期エンドポイントで、失敗したタスクは課金されません。
原生音轨:生成画面的同时自动产出同步音频(对话 / 音效 / 环境音),成片无水印,MP4 = H.264 + AAC。
Generated video will appear here
Provide URLs and click Generate
Dialogue, SFX and ambience generated with the frames
From $0.0294/s at 480p — pay for exactly the seconds you use
One model name; up to 7 reference images
4s clip in ~30-70s measured, clean output
Grok Imagine Video 1.5 は xAI の動画生成 API です。Grok Imagine Video 1.5 は xAI の動画モデルで、競合の多くがまだ持っていない特徴があります —— 映像と一緒に同期した音声トラック(セリフ、効果音、環境音)を生成するため、ダウンロードした MP4 が最初から H.264 映像 + AAC 音声で完結し、別途アフレコ工程が要りません。モデル名 1 つで text-to-video と image-to-video(参照画像は最大 7 枚、URL でも base64 でも可)をカバーし、尺は 1〜15 秒。秒課金は 3 段階で、480p $0.0294/s、720p $0.0529/s、1080p $0.0882/s。4 秒のドラフトなら約 $0.12 です。出力にウォーターマークはなく、速度も実測で 4 秒クリップが 30〜70 秒ほどで返ってきます。作成と照会に分かれた非同期エンドポイントで、失敗したタスクは課金されません。 APIMODELS のプラットフォーム経由なら、統一 API と明朗な従量課金でこのモデルを呼び出せます。 現在の料金: 480p: $0.0294, 720p: $0.0529, 1080p: $0.0882。
All six clips were generated on this exact model with native audio — four at the 480p / 8s tier ($0.24 each) and two at 1080p (4s $0.35, 10s $0.88). The prompt under each clip is exactly what produced it, and the tier tag shows what it cost. The four 480p prompts are adapted from our MiniMax H3 prompt library.
Style: Cinematic lifestyle commercial. SCENE: Red-haired young woman in a white tank top, loose pink trousers, and sunglasses. A classic yellow convertible. Palm-tree-lined coastal highway, sandy beach, and ocean. SHOTS: 0-4s low angle tracks the yellow convertible down a sunny, palm-lined street, cut to close-up of the smiling woman driving, hair blowing in the wind. 4-8s sunset wide shot: she sits relaxed on the parked car's hood on the beach. CAMERA: smooth dynamic tracking, low-angle rolling mounts, intimate driver-seat close-ups. LIGHTING: sun-drenched daylight into golden-hour sunset, warm glow, vibrant yellow, pastel pink, deep ocean blue. AUDIO: rushing wind, classic car engine hum, distant ocean waves, upbeat carefree indie-pop instrumental.
Style: Handheld action sports parkour videography. SCENE: Athletic young man, short black hair, black t-shirt, olive cargo pants, white sneakers. Urban alleyways and high-rise rooftops. SHOTS: 0-4s man executes a parkour vault over a metal barricade fence in a narrow alley, then a dynamic leap over a concrete parapet against clear blue sky, low tracking angles. 4-8s man lands in a balanced crouch on a roof parapet edge and rolls; camera pulls back to a high aerial wide shot revealing a sprawling skyline. CAMERA: wide-angle action lens, handheld energetic tracking, extreme low and high angles. LIGHTING: bright midday sun, harsh high-contrast shadows, concrete grays, olive greens, sky blue. AUDIO: whooshing wind, heavy landing thuds, rapid footsteps, fast-paced electronic action instrumental.
Style: High-end cosmetics commercial, photorealistic. SCENE: East Asian woman, long dark hair, white silk robe. Marble vanity countertop, illuminated ring mirror, gold and rose-gold cosmetic bottles, bright white room. SHOTS: 0-4s medium shot of woman holding a frosted bottle and glass dropper, cut to close-up dispensing a drop onto her cheek, then macro shot of clear liquid touching the skin. 4-8s extreme close-up of a beige teardrop sponge blending foundation on her cheek, then a fluffy rose-gold brush applying blush. CAMERA: macro beauty lenses, shallow depth of field, static positioning with rhythmic cuts. LIGHTING: soft diffused studio light, high-key, ring-light reflections in the eyes, crisp whites, marble grays, metallic golds, glossy pastel pinks. AUDIO: pristine room tone, distinct liquid droplet sound, soft sponge tapping, serene lo-fi background music.
Style: Realistic everyday vlog video. SCENE: Three young Asian women in casual summer clothes standing outside a suburban house. A large grey monitor lizard clinging to a black metal driveway gate, residential street with houses. SHOTS: 0-4s medium shot of the three women by the open gate, laughing and reacting as the large lizard climbs the metal bars, one records it on a smartphone. 4-8s medium side-profile shot of the women watching; the woman in the green shirt holds her phone close to film the reptile clinging to the vertical bars. CAMERA: handheld digital camera, natural depth of field, eye-level medium framing. LIGHTING: bright midday outdoor sun, crisp shadows on concrete, foliage greens, neutral grays, matte black metal. AUDIO: ambient street atmosphere, faint scooter hum in the distance, casual background laughter.
A chef flambeing a pan in a busy restaurant kitchen, flames leaping up, steam and sparks, handheld close-up, warm tungsten light
A hot air balloon drifting over terraced rice fields at sunrise, thin morning mist, slow aerial glide
広告キャンペーンや SNS 施策に向けたブランド動画を短時間で作ります。
TikTok、Instagram、YouTube 向けの縦型ショート動画を量産できます。
機能紹介やチュートリアル動画を作り、コンバージョンにつなげます。
講座の解説、知識の説明、研修用の動画を低コストで継続的に作れます。
Grok Imagine Video 1.5 は APIMODELS 経由で 480p: $0.0294, 720p: $0.0529, 1080p: $0.0882 で利用できます。課金は従量制で、生成した分だけの支払いです。
APIMODELS に登録して API キーを取得し、統一エンドポイントを呼ぶだけです。cURL / Python / Node.js のサンプルを含む詳しいドキュメントを用意しています。
APIMODELS は同じ Grok Imagine Video 1.5 を集約プラットフォーム経由で提供します。統一された API インターフェースなので、プロバイダごとにアカウントを作る必要はありません。キー 1 本ですべてのモデルに届きます。
xAI の動画モデルで、際立った特徴はネイティブ音声です。映像と一緒にセリフ・効果音・環境音の同期トラックを生成するので、ダウンロードした MP4 が最初から H.264 映像 + AAC 音声で完結します。モデル名 1 つで text-to-video と image-to-video をカバーし、尺は 1〜15 秒、ウォーターマークはありません。apimodels.app では秒課金で、480p の $0.0294/s からです。
解像度別の秒課金です: 480p $0.0294/s、720p $0.0529/s、1080p $0.0882/s。480p の 4 秒ドラフトなら約 $0.12、1080p の 15 秒クリップで約 $1.32。text-to-video と image-to-video は同額で、参照画像を渡しても追加料金はありません。失敗したタスクは課金されません。
apimodels.app の /api/v1/video/generations に POST し、model に "grok-imagine-video-1.5"、それに prompt と images(参照画像は最大 7 枚、URL でも base64 でも可)のどちらか一方または両方を渡します。あとは返ってきた task_id で同じエンドポイントを照会するだけです。形式は VEO、Kling、Seedance と同一なので、モデルを乗り換えるのは文字列を 1 つ変えるだけ。API キー 1 本ですべてカバーできます。4 秒クリップは実測で 30〜70 秒ほどで返ってきます。
APIMODELS では Grok Imagine Video 1.5 が 60 以上のモデルと同じ API キー・同じ残高の上に並んでいるので、選択は「相性」の問題であって「囲い込み」の問題ではありません。Text to Video、Image to Video、Native Audio、1-15s、480p / 720p / 1080p、No Watermark に対応しており、他の動画生成モデルと価格・機能を並べて評価できます。乗り換えはモデル名の文字列を 1 つ書き換えるだけ。新しいアカウントも追加の実装も要りません。動画生成の選択肢と最新価格は apimodels.app/models で確認できます。
Grok Imagine Video 1.5 は次に対応しています: Text to Video、Image to Video、Native Audio、1-15s、480p / 720p / 1080p、No Watermark。パラメータの全一覧と呼び出し例は APIMODELS のドキュメントをご覧ください。
はい。APIMODELS は Grok Imagine Video 1.5 を単一の統一 API とキー 1 本で提供します。プロバイダごとのアカウントも、各社の地域ごとのネットワーク経路を自分で面倒みる必要もありません。
Stripe(Visa、Mastercard などの国際カード)と Alipay に対応しています。支払い後、残高はすぐ反映されます。