The request usually sounds like this: take the driver in this clip and put my person in the seat, same position, same posture, and leave the car, the road, the light and the camera exactly as they are. That is character replacement, and it is different from generating a new video from a photo: the original clip supplies the motion, the framing and the timing, and the model only rebuilds the person.
Wan 3.0 does this in one call through its omni-reference mode. You pass the clip as a reference video and the new person as reference images, and refer to them in the prompt as "video 1" and "figure 1", "figure 2" and so on. No masks, no tracking, no separate face-swap step.
Say three things, in this order: who is replaced (the person in video 1, or the driver, the presenter, the dancer), who replaces them (the person in figure 1), and what must not change. The last part does most of the work. List the things you care about explicitly — vehicle, interior, background, lighting, camera angle and movement, the motion and its timing — because anything left unsaid is fair game for the model to "improve".
If you have several photos of the same person, pass them all and say so: "figure 1, figure 2 and figure 3 are the same person from different angles". Front, three-quarter and side views keep the face and outfit consistent when the head turns. Match the framing to the clip where you can — a full-body photo for a full-body shot.
If the source has burned-in captions or watermarks you do not want carried over, add a final sentence asking for any on-screen text to be removed.
The most-copied version of this technique right now is the AI car driving video: a cinematic clip of someone driving a luxury car, with the driver replaced by the person in your photo and the car, road and camera left untouched. It went viral on Instagram and YouTube in September 2026. Four ready prompts, how to choose the driving clip and the per-clip cost are on the dedicated page.
With a reference video, Wan 3.0 bills the input clip’s seconds plus the output’s seconds at the same per-second rate, and the two together are capped at 30 seconds. With duration set to -1 the output matches the clip, so the bill is roughly twice the clip length times the rate. Prime runs the same job about six times faster for roughly 1.5 times the price. Reference images cost nothing extra.
| Resolution | Standard, per second | Prime, per second | 9 s clip, standard (18 billed s) |
|---|---|---|---|
| 480P | $0.045 | $0.068 | $0.81 |
| 720P | $0.09 | $0.14 | $1.62 |
| 1080P | $0.18 | $0.28 | $3.24 |
The reference clip must be 1 to 15 seconds, and input plus output at most 30 seconds, so trim long footage first. Standard mode typically takes four to ten minutes for a 720P clip; prime brings that close to a minute. One person per clip gives the most reliable result — in a crowd the model has to guess who you meant.
Photos of real people can be rejected by the upstream content check, and so can revealing outfits. Hands touching objects — the steering wheel, a cup, a microphone — are the weakest part of the output. Only swap people who have agreed to it, and never use this to make someone appear to do or say something they did not.
If you only need the person to follow the clip’s motion in your own background, or you want a cheaper per-second rate, Wan 2.2 Animate’s replace mode ($0.0742/s at 480P, billed on the reference clip only) is the lighter option.
cURL
curl -X POST https://api.apimodels.app/v1/video/generations \
-H "Authorization: Bearer $APIMODELS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "wan-3.0-video",
"prompt": "Replace the driver in video 1 with the person in figure 1 and figure 2 (the same person from two angles). Keep their face, hair, build and clothing exactly as in the figures. Keep everything else from video 1 unchanged: the car, the interior, the road, the lighting, the camera angle and movement, the driving motion and the timing.",
"reference_video_urls": ["https://your-bucket.example/driving-shot.mp4"],
"reference_image_urls": ["https://your-bucket.example/person-front.jpg", "https://your-bucket.example/person-side.jpg"],
"resolution": "720P",
"mode": "standard",
"duration": -1
}'
# -> {"data": {"taskId": "..."}} then GET /v1/video/generations?task_id=...Python
import requests, time, os
API = "https://api.apimodels.app/v1/video/generations"
H = {"Authorization": f"Bearer {os.environ['APIMODELS_API_KEY']}"}
TEMPLATE = (
"Replace the person in video 1 with the person shown in {figs}"
"{same}. Keep their face, hair, build and clothing exactly as in the "
"reference. Keep everything else from video 1 unchanged: movements, gestures "
"and timing, camera framing and motion, lighting, background and setting."
)
def swap(video_url, image_urls, resolution="720P", mode="standard"):
figs = ", ".join(f"figure {i + 1}" for i in range(len(image_urls)))
same = " (the same person from different angles)" if len(image_urls) > 1 else ""
r = requests.post(API, headers=H, json={
"model": "wan-3.0-video",
"prompt": TEMPLATE.format(figs=figs, same=same),
"reference_video_urls": [video_url],
"reference_image_urls": image_urls,
"resolution": resolution, "mode": mode,
"duration": -1, # output length follows the reference clip
}).json()
task = r["data"]["taskId"]
while True:
s = requests.get(API, headers=H, params={"task_id": task}).json()["data"]
if s["state"] in ("completed", "failed"):
return s
time.sleep(15)Yes. With Wan 3.0 you send the video as a reference video and photos of the new person as reference images, and the prompt says to replace the person in video 1 with the person in figure 1 while keeping everything else. The motion, camera and scene come from the original clip.
List them in the prompt as things that must stay unchanged: the vehicle, interior, road, location, background, lighting, camera angle, camera movement and the timing of the motion. The model preserves what you name much more reliably than what you leave implicit.
At 720P standard on apimodels.app, about $1.80: 10 seconds of input plus 10 seconds of output at $0.09 per second. At 480P it is about $0.90 and at 1080P about $3.60. Prime costs roughly 1.5 times more and is about six times faster.
Several, if you have them. Two to four photos of the same person from different angles, with the prompt saying they show the same person, keep the face and outfit stable when the subject turns. One clear front-facing photo works for shots where the person mostly faces the camera.
No. Face swap changes only the face and keeps the original body and clothing. Character replacement rebuilds the whole person — face, hair, build and outfit — from your reference, which is what you want when the new person should look completely different.