In September 2026 feeds filled with the same shot: someone behind the wheel of a black luxury car on an empty highway, low sun, a slow push-in, a glance at the camera. Almost none of them were filmed. Creators take one well-shot driving clip, upload a photo of themselves, and have a video model put them in the seat. The prompt people copy for it is long because it has to be: it lists every part of the scene that must not change.
If you are building an AI car driving video generator — an app, a bot or a batch pipeline that turns a user’s selfie into the clip — the whole job is one Wan 3.0 call per video, and the Python function at the bottom of this page is a working starting point.
Our prompts below follow the structure that holds up best in production on apimodels.app — the same pattern our customers used for hundreds of successful swaps: refer to the clip as "video 1" and the photos as "figure 1", "figure 2"; say the photos show one person from different angles; list what stays the same; ask for any on-screen text to be removed.
Basic swap, one photo: Replace the driver in video 1 with the person shown in figure 1. Keep the person's face, hair, body and outfit exactly as in figure 1. The person sits in exactly the same seat, position and posture as the original driver. Preserve everything else from video 1 exactly: the same car, interior, steering wheel, dashboard, road, location, background, lighting, reflections and shadows, the same camera angle and camera movement, the same driving motion, hand movements and timing. The output must look like the original clip with only the driver swapped. If there is any text on the video, remove it.
Multi-angle, most stable (front, three-quarter and side photos of you): Replace the driver in video 1 with the person shown in figure 1, figure 2 and figure 3. All figures show the same person from different angles — use them together to keep the face, hair, body and outfit consistent when the head turns. Preserve everything else from video 1 exactly: the same car, interior, steering wheel, road, background, lighting, reflections, camera angle, camera movement, driving motion and timing. The output must look like the original clip with only the driver swapped. If there is any text on the video, remove it.
Keep the glance at the camera (for clips where the driver looks into the lens): Replace the driver in video 1 with the person shown in figure 1, keeping their face, hair and outfit exactly as in figure 1. Follow the original driver's performance frame by frame: the steering gestures, the head turns, the moment they glance at the camera and any smile or nod. Keep the car, the interior, the road, the lighting and the camera movement identical to video 1. If there is any text on the video, remove it.
Swap the outfit too (figure 2 is the clothing): Replace the driver in video 1 with the person shown in figure 1, and dress them in the outfit shown in figure 2. Keep the person's face and hair exactly as in figure 1. Preserve everything else from video 1 exactly: the car, interior, road, lighting, reflections, camera angle, camera movement, driving motion and timing. If there is any text on the video, remove it.
The clip decides how cinematic the result is, so choose it first: 5 to 12 seconds, one person in the driver’s seat, face visible for most of the shot, steady camera, no heavy captions. Mercedes, BMW, Porsche or any other car works — the model keeps whatever car is in the clip. Wan 3.0 accepts clips of 1 to 15 seconds.
For the photos: sharp, evenly lit, face unobstructed, framed like the driver in the clip (head and shoulders is enough for a seated shot). Two or three angles of the same person beat one photo whenever the driver turns their head. Sunglasses in the photo will usually carry into the video.
With a reference clip, Wan 3.0 bills the clip’s seconds plus the output’s seconds. Set duration to -1 and the output matches the clip, so a 9-second clip is 18 billed seconds. Prime is about six times faster for roughly 1.5 times the price. Photos are free.
| 9-second clip | Standard | Prime |
|---|---|---|
| 480P | $0.81 | $1.22 |
| 720P | $1.62 | $2.52 |
| 1080P | $3.24 | $5.04 |
cURL
curl -X POST https://api.apimodels.app/v1/video/generations \
-H "Authorization: Bearer $APIMODELS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "wan-3.0-video",
"prompt": "Replace the driver in video 1 with the person shown in figure 1 and figure 2. Both figures show the same person from different angles. Preserve everything else from video 1 exactly: the car, interior, steering wheel, road, lighting, camera angle, camera movement, driving motion and timing. If there is any text on the video, remove it.",
"reference_video_urls": ["https://your-bucket.example/luxury-car-driving.mp4"],
"reference_image_urls": ["https://your-bucket.example/me-front.jpg", "https://your-bucket.example/me-side.jpg"],
"resolution": "720P",
"duration": -1
}'Python
import requests, time, os
API = "https://api.apimodels.app/v1/video/generations"
H = {"Authorization": f"Bearer {os.environ['APIMODELS_API_KEY']}"}
def car_trend(driving_clip_url, photo_urls, resolution="720P", mode="standard"):
figs = ", ".join(f"figure {i + 1}" for i in range(len(photo_urls)))
same = " All figures show the same person from different angles." if len(photo_urls) > 1 else ""
prompt = (
f"Replace the driver in video 1 with the person shown in {figs}.{same} "
"Keep the person's face, hair, body and outfit exactly as in the figures. "
"Preserve everything else from video 1 exactly: the car, interior, steering wheel, "
"road, background, lighting, reflections, camera angle, camera movement, driving "
"motion and timing. If there is any text on the video, remove it."
)
task = requests.post(API, headers=H, json={
"model": "wan-3.0-video", "prompt": prompt,
"reference_video_urls": [driving_clip_url], "reference_image_urls": photo_urls,
"resolution": resolution, "mode": mode, "duration": -1,
}).json()["data"]["taskId"]
while True:
s = requests.get(API, headers=H, params={"task_id": task}).json()["data"]
if s["state"] in ("completed", "failed"):
return s
time.sleep(15)A viral short-video format from September 2026: take a cinematic clip of someone driving a luxury car and have an AI video model replace the driver with the person in your photo, keeping the car, road, lighting and camera movement exactly the same. It is character replacement with a reference video.
Yes. Wan 3.0 on apimodels.app does the swap in one call: POST /v1/video/generations with the driving clip in reference_video_urls, the user’s photos in reference_image_urls and a prompt that replaces the driver in video 1 with the person in figure 1. Poll the task and deliver the MP4. A 9-second 720P clip costs about $1.62, and failed jobs are free.
Any video model that accepts a reference video plus reference images can do it. Wan 3.0 does it in one call: send the driving clip as reference_video_urls, your photos as reference_image_urls, and a prompt that replaces the driver in video 1 with the person in figure 1.
List them in the prompt as things to preserve: the car, interior, steering wheel, road, background, lighting, reflections, camera angle, camera movement, driving motion and timing. End with "the output must look like the original clip with only the driver swapped".
On apimodels.app with Wan 3.0 standard: about $0.81 at 480P, $1.62 at 720P and $3.24 at 1080P for a 9-second clip, because the clip’s seconds and the output’s seconds are both billed. Failed jobs are not charged.
The model only knows the angle in your photo. Upload two or three photos of yourself from different angles and say in the prompt that they show the same person; that is the single biggest fix.