2026 年 9 月,信息流里全是同一个镜头:有人开着一辆黑色豪车驶过空旷的高速,低角度阳光、缓慢推近、对着镜头一瞥。几乎没有一条是真拍的。创作者找一段拍得好的驾驶视频,上传一张自己的照片,让视频模型把自己放进驾驶座。大家复制的那段提示词之所以那么长,是因为它必须列出场景里每一样不能变的东西。
如果你在做一个 AI 开车大片生成器(AI car driving video generator)—— 一个把用户自拍变成大片的 App、机器人或批量流水线 —— 每条视频只需要一次 Wan 3.0 调用,本页底部的 Python 函数就是能直接用的起点。
下面的提示词沿用了在 apimodels.app 生产环境里最稳的写法 —— 我们的客户用同一套结构成功跑了几百条:把视频称作「video 1」、照片称作「figure 1」「figure 2」;说明几张照片是同一个人的不同角度;逐项列出保持不变的东西;要求去掉画面上的文字。
基础版,一张照片:Replace the driver in video 1 with the person shown in figure 1. Keep the person's face, hair, body and outfit exactly as in figure 1. The person sits in exactly the same seat, position and posture as the original driver. Preserve everything else from video 1 exactly: the same car, interior, steering wheel, dashboard, road, location, background, lighting, reflections and shadows, the same camera angle and camera movement, the same driving motion, hand movements and timing. The output must look like the original clip with only the driver swapped. If there is any text on the video, remove it.
多角度版,最稳定(正面、四分之三侧、侧面三张照片):Replace the driver in video 1 with the person shown in figure 1, figure 2 and figure 3. All figures show the same person from different angles — use them together to keep the face, hair, body and outfit consistent when the head turns. Preserve everything else from video 1 exactly: the same car, interior, steering wheel, road, background, lighting, reflections, camera angle, camera movement, driving motion and timing. The output must look like the original clip with only the driver swapped. If there is any text on the video, remove it.
保留看镜头的动作(原片司机有看镜头的镜头时用):Replace the driver in video 1 with the person shown in figure 1, keeping their face, hair and outfit exactly as in figure 1. Follow the original driver's performance frame by frame: the steering gestures, the head turns, the moment they glance at the camera and any smile or nod. Keep the car, the interior, the road, the lighting and the camera movement identical to video 1. If there is any text on the video, remove it.
顺便换衣服(figure 2 是服装图):Replace the driver in video 1 with the person shown in figure 1, and dress them in the outfit shown in figure 2. Keep the person's face and hair exactly as in figure 1. Preserve everything else from video 1 exactly: the car, interior, road, lighting, reflections, camera angle, camera movement, driving motion and timing. If there is any text on the video, remove it.
成片有多大片感取决于原片,所以先选视频:5 到 12 秒、驾驶座上只有一个人、大部分时间能看到脸、镜头稳、没有大字幕。奔驰、宝马、保时捷或别的车都行 —— 模型会保留原片里的车。Wan 3.0 接受 1 到 15 秒的视频。
照片:清晰、光线均匀、脸不被遮挡,取景和原片司机一致(坐姿镜头用头肩照就够)。只要司机会转头,同一个人两三个角度的照片就比一张好。照片里戴墨镜,视频里通常也会戴。
带参考视频时,Wan 3.0 按「原片秒数 + 输出秒数」计费。duration 设为 -1 时输出与原片等长,所以 9 秒原片计费 18 秒。Prime 档快约 6 倍、价格约 1.5 倍。照片不收费。
| 9 秒原片 | 标准档 | Prime |
|---|---|---|
| 480P | $0.81 | $1.22 |
| 720P | $1.62 | $2.52 |
| 1080P | $3.24 | $5.04 |
cURL
curl -X POST https://api.apimodels.app/v1/video/generations \
-H "Authorization: Bearer $APIMODELS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "wan-3.0-video",
"prompt": "Replace the driver in video 1 with the person shown in figure 1 and figure 2. Both figures show the same person from different angles. Preserve everything else from video 1 exactly: the car, interior, steering wheel, road, lighting, camera angle, camera movement, driving motion and timing. If there is any text on the video, remove it.",
"reference_video_urls": ["https://your-bucket.example/luxury-car-driving.mp4"],
"reference_image_urls": ["https://your-bucket.example/me-front.jpg", "https://your-bucket.example/me-side.jpg"],
"resolution": "720P",
"duration": -1
}'Python
import requests, time, os
API = "https://api.apimodels.app/v1/video/generations"
H = {"Authorization": f"Bearer {os.environ['APIMODELS_API_KEY']}"}
def car_trend(driving_clip_url, photo_urls, resolution="720P", mode="standard"):
figs = ", ".join(f"figure {i + 1}" for i in range(len(photo_urls)))
same = " All figures show the same person from different angles." if len(photo_urls) > 1 else ""
prompt = (
f"Replace the driver in video 1 with the person shown in {figs}.{same} "
"Keep the person's face, hair, body and outfit exactly as in the figures. "
"Preserve everything else from video 1 exactly: the car, interior, steering wheel, "
"road, background, lighting, reflections, camera angle, camera movement, driving "
"motion and timing. If there is any text on the video, remove it."
)
task = requests.post(API, headers=H, json={
"model": "wan-3.0-video", "prompt": prompt,
"reference_video_urls": [driving_clip_url], "reference_image_urls": photo_urls,
"resolution": resolution, "mode": mode, "duration": -1,
}).json()["data"]["taskId"]
while True:
s = requests.get(API, headers=H, params={"task_id": task}).json()["data"]
if s["state"] in ("completed", "failed"):
return s
time.sleep(15)2026 年 9 月走红的短视频玩法:找一段电影感的豪车驾驶视频,用 AI 视频模型把司机换成你照片里的人,车、路、光线和镜头运动完全不变。本质是「参考视频 + 人物替换」。
有。apimodels.app 上的 Wan 3.0 一次调用就能完成:POST /v1/video/generations,驾驶视频放 reference_video_urls,用户照片放 reference_image_urls,提示词写把 video 1 里的司机换成 figure 1 里的人,轮询任务后交付 MP4。9 秒 720P 约 $1.62,失败不收钱。
任何支持「参考视频 + 参考图」的视频模型都能做。Wan 3.0 一次调用就行:驾驶视频放 reference_video_urls、照片放 reference_image_urls,提示词写把 video 1 里的司机换成 figure 1 里的人。
在提示词里把它们列为要保留的:车、内饰、方向盘、道路、背景、光线、反光、镜头角度、镜头运动、驾驶动作和节奏。结尾加一句「成片必须和原片一样,只换了司机」。
在 apimodels.app 上用 Wan 3.0 标准档,9 秒原片 480P 约 $0.81、720P 约 $1.62、1080P 约 $3.24,因为原片秒数和输出秒数都计费。失败不收钱。
模型只知道你照片里那个角度。上传两三张不同角度的自己,并在提示词里说明是同一个人,这是最有效的办法。