
grok-imagine-video-1.5Grok Imagine Video 1.5는 xAI의 영상 모델이며, 경쟁 모델 상당수가 아직 갖추지 못한 특징이 있습니다. 영상과 함께 동기화된 오디오 트랙(대사, 효과음, 환경음)을 생성하기 때문에 내려받은 MP4가 처음부터 H.264 영상과 AAC 오디오로 완결되어 별도의 더빙 공정이 필요 없습니다. 모델 이름 하나로 text-to-video와 image-to-video(참조 이미지 최대 7장, URL 또는 base64)를 모두 다루고 길이는 1~15초입니다. 초당 요금은 세 구간으로 480p $0.0294/s, 720p $0.0529/s, 1080p $0.0882/s이며 4초 드래프트는 약 $0.12입니다. 출력에 워터마크가 없고 속도도 실측 기준 4초 클립이 30~70초 만에 돌아옵니다. 생성과 조회가 나뉜 비동기 엔드포인트이고 실패한 작업은 청구되지 않습니다.
原生音轨:生成画面的同时自动产出同步音频(对话 / 音效 / 环境音),成片无水印,MP4 = H.264 + AAC。
Generated video will appear here
Provide URLs and click Generate
Dialogue, SFX and ambience generated with the frames
From $0.0294/s at 480p — pay for exactly the seconds you use
One model name; up to 7 reference images
4s clip in ~30-70s measured, clean output
Grok Imagine Video 1.5은(는) xAI의 영상 생성 API입니다. Grok Imagine Video 1.5는 xAI의 영상 모델이며, 경쟁 모델 상당수가 아직 갖추지 못한 특징이 있습니다. 영상과 함께 동기화된 오디오 트랙(대사, 효과음, 환경음)을 생성하기 때문에 내려받은 MP4가 처음부터 H.264 영상과 AAC 오디오로 완결되어 별도의 더빙 공정이 필요 없습니다. 모델 이름 하나로 text-to-video와 image-to-video(참조 이미지 최대 7장, URL 또는 base64)를 모두 다루고 길이는 1~15초입니다. 초당 요금은 세 구간으로 480p $0.0294/s, 720p $0.0529/s, 1080p $0.0882/s이며 4초 드래프트는 약 $0.12입니다. 출력에 워터마크가 없고 속도도 실측 기준 4초 클립이 30~70초 만에 돌아옵니다. 생성과 조회가 나뉜 비동기 엔드포인트이고 실패한 작업은 청구되지 않습니다. APIMODELS 플랫폼을 거치면 통합 API와 투명한 종량 과금으로 이 모델을 호출할 수 있습니다. 현재 가격: 480p: $0.0294, 720p: $0.0529, 1080p: $0.0882.
All six clips were generated on this exact model with native audio — four at the 480p / 8s tier ($0.24 each) and two at 1080p (4s $0.35, 10s $0.88). The prompt under each clip is exactly what produced it, and the tier tag shows what it cost. The four 480p prompts are adapted from our MiniMax H3 prompt library.
Style: Cinematic lifestyle commercial. SCENE: Red-haired young woman in a white tank top, loose pink trousers, and sunglasses. A classic yellow convertible. Palm-tree-lined coastal highway, sandy beach, and ocean. SHOTS: 0-4s low angle tracks the yellow convertible down a sunny, palm-lined street, cut to close-up of the smiling woman driving, hair blowing in the wind. 4-8s sunset wide shot: she sits relaxed on the parked car's hood on the beach. CAMERA: smooth dynamic tracking, low-angle rolling mounts, intimate driver-seat close-ups. LIGHTING: sun-drenched daylight into golden-hour sunset, warm glow, vibrant yellow, pastel pink, deep ocean blue. AUDIO: rushing wind, classic car engine hum, distant ocean waves, upbeat carefree indie-pop instrumental.
Style: Handheld action sports parkour videography. SCENE: Athletic young man, short black hair, black t-shirt, olive cargo pants, white sneakers. Urban alleyways and high-rise rooftops. SHOTS: 0-4s man executes a parkour vault over a metal barricade fence in a narrow alley, then a dynamic leap over a concrete parapet against clear blue sky, low tracking angles. 4-8s man lands in a balanced crouch on a roof parapet edge and rolls; camera pulls back to a high aerial wide shot revealing a sprawling skyline. CAMERA: wide-angle action lens, handheld energetic tracking, extreme low and high angles. LIGHTING: bright midday sun, harsh high-contrast shadows, concrete grays, olive greens, sky blue. AUDIO: whooshing wind, heavy landing thuds, rapid footsteps, fast-paced electronic action instrumental.
Style: High-end cosmetics commercial, photorealistic. SCENE: East Asian woman, long dark hair, white silk robe. Marble vanity countertop, illuminated ring mirror, gold and rose-gold cosmetic bottles, bright white room. SHOTS: 0-4s medium shot of woman holding a frosted bottle and glass dropper, cut to close-up dispensing a drop onto her cheek, then macro shot of clear liquid touching the skin. 4-8s extreme close-up of a beige teardrop sponge blending foundation on her cheek, then a fluffy rose-gold brush applying blush. CAMERA: macro beauty lenses, shallow depth of field, static positioning with rhythmic cuts. LIGHTING: soft diffused studio light, high-key, ring-light reflections in the eyes, crisp whites, marble grays, metallic golds, glossy pastel pinks. AUDIO: pristine room tone, distinct liquid droplet sound, soft sponge tapping, serene lo-fi background music.
Style: Realistic everyday vlog video. SCENE: Three young Asian women in casual summer clothes standing outside a suburban house. A large grey monitor lizard clinging to a black metal driveway gate, residential street with houses. SHOTS: 0-4s medium shot of the three women by the open gate, laughing and reacting as the large lizard climbs the metal bars, one records it on a smartphone. 4-8s medium side-profile shot of the women watching; the woman in the green shirt holds her phone close to film the reptile clinging to the vertical bars. CAMERA: handheld digital camera, natural depth of field, eye-level medium framing. LIGHTING: bright midday outdoor sun, crisp shadows on concrete, foliage greens, neutral grays, matte black metal. AUDIO: ambient street atmosphere, faint scooter hum in the distance, casual background laughter.
A chef flambeing a pan in a busy restaurant kitchen, flames leaping up, steam and sparks, handheld close-up, warm tungsten light
A hot air balloon drifting over terraced rice fields at sunrise, thin morning mist, slow aerial glide
광고 캠페인과 SNS 마케팅에 쓸 브랜드 영상을 빠르게 만듭니다.
TikTok, Instagram, YouTube에 맞는 세로형 짧은 영상을 대량으로 뽑습니다.
기능 소개와 사용법 영상을 만들어 전환율을 끌어올립니다.
강의 해설과 지식 설명, 사내 교육 영상을 낮은 비용으로 꾸준히 제작합니다.
Grok Imagine Video 1.5은(는) APIMODELS를 통해 480p: $0.0294, 720p: $0.0529, 1080p: $0.0882에 이용할 수 있습니다. 과금은 종량제라서 생성한 만큼만 냅니다.
APIMODELS에 가입해 API 키를 받고 통합 엔드포인트를 호출하면 됩니다. cURL / Python / Node.js 예제를 담은 상세 문서를 제공합니다.
APIMODELS는 같은 Grok Imagine Video 1.5을(를) 집약 플랫폼을 통해 제공합니다. API 인터페이스가 통합되어 있어 공급자마다 계정을 만들 필요가 없고, 키 하나로 모든 모델에 닿습니다.
xAI의 영상 모델이며 두드러진 특징은 네이티브 오디오입니다. 영상과 함께 대사와 효과음, 환경음의 동기 트랙을 생성하므로 내려받은 MP4가 처음부터 H.264 영상과 AAC 오디오로 완결됩니다. 모델 이름 하나로 text-to-video와 image-to-video를 다루고 길이는 1~15초이며 워터마크가 없습니다. apimodels.app에서는 초당 과금이고 480p의 $0.0294/s부터 시작합니다.
해상도별 초당 과금입니다: 480p $0.0294/s, 720p $0.0529/s, 1080p $0.0882/s. 480p 4초 드래프트는 약 $0.12, 1080p 15초 클립은 약 $1.32입니다. text-to-video와 image-to-video가 같은 값이고 참조 이미지를 넘겨도 추가 요금이 없습니다. 실패한 작업은 청구되지 않습니다.
apimodels.app의 /api/v1/video/generations로 POST하면서 model에 "grok-imagine-video-1.5"를 넣고 prompt와 images(참조 이미지 최대 7장, URL 또는 base64) 중 하나 또는 둘 다를 함께 보냅니다. 그다음 돌려받은 task_id로 같은 엔드포인트를 조회하면 됩니다. 형식이 VEO, Kling, Seedance와 동일해서 모델을 바꾸는 일은 문자열 하나 바꾸는 일이고, API 키 하나로 전부 커버됩니다. 4초 클립은 실측 기준 30~70초 만에 돌아옵니다.
APIMODELS에서는 Grok Imagine Video 1.5이(가) 60개가 넘는 모델과 같은 API 키, 같은 잔액 위에 나란히 놓입니다. 그래서 선택은 궁합의 문제이지 종속의 문제가 아닙니다. Text to Video、Image to Video、Native Audio、1-15s、480p / 720p / 1080p、No Watermark을(를) 지원하며 다른 영상 생성 모델과 가격·성능을 나란히 놓고 따져볼 수 있습니다. 갈아타기는 모델 이름 문자열 하나만 바꾸면 되고 새 계정도 추가 작업도 필요 없습니다. 영상 생성 선택지와 실시간 가격은 apimodels.app/models에서 볼 수 있습니다.
Grok Imagine Video 1.5은(는) 다음을 지원합니다: Text to Video、Image to Video、Native Audio、1-15s、480p / 720p / 1080p、No Watermark. 전체 파라미터와 호출 예제는 APIMODELS 문서를 참고하세요.
네. APIMODELS는 Grok Imagine Video 1.5을(를) 하나의 통합 API와 키 한 개로 제공합니다. 공급자별 계정도 필요 없고, 각 공급자의 지역별 네트워크 경로를 직접 챙길 필요도 없습니다.
Stripe(Visa, Mastercard 등 해외 카드)와 Alipay를 지원합니다. 결제 후 잔액은 즉시 반영됩니다.