digital-humanDigital Human turns a video of a person plus a driving audio clip into a lip-synced talking video: the person keeps their original look, motion and background while the mouth is re-animated to speak the audio. If the audio runs longer than the video, the footage is extended automatically — the output always follows the audio length. Send two URLs (video_url: mp4/mov of the person; audio_url: mp3/wav/m4a, up to 30 minutes) and get an mp4 back at the source video's resolution. Billing is $0.01 per second of the driving audio — a 60-second voiceover costs $0.60, half the per-second price of image-based lip-sync. Use AI Lip-Sync instead when you only have a still portrait image. Typical turnaround is about 2 minutes for short clips. Only successful generations are charged; results are kept for 7 days.
$0.01 per second of audio vs $0.02/s for the image-based ai-lipsync — when you already have footage of the person, this is the cheaper path
Lighting, motion, background and camera work come from your footage; only the mouth is re-animated to match the new audio
Audio longer than the video? The footage is extended automatically — verified with a 6.2s voiceover over a 5s clip
Output matches the source video resolution (tested at 1440×2560) — no forced downscale
POST /api/v1/video/generations with model digital-human, video_url and audio_url — poll the task id for the mp4
$0.01 per second of audio · if the audio outlasts the video, the footage is extended automatically
Your talking video will appear here
Digital Human 是由 APIMODELS 提供的视频生成 API。Digital Human turns a video of a person plus a driving audio clip into a lip-synced talking video: the person keeps their original look, motion and background while the mouth is re-animated to speak the audio. If the audio runs longer than the video, the footage is extended automatically — the output always follows the audio length. Send two URLs (video_url: mp4/mov of the person; audio_url: mp3/wav/m4a, up to 30 minutes) and get an mp4 back at the source video's resolution. Billing is $0.01 per second of the driving audio — a 60-second voiceover costs $0.60, half the per-second price of image-based lip-sync. Use AI Lip-Sync instead when you only have a still portrait image. Typical turnaround is about 2 minutes for short clips. Only successful generations are charged; results are kept for 7 days. 通过 APIMODELS 平台,您可以使用统一的 API 接口调用该模型,按量计费、价格透明。当前定价:per second of audio: $0.01。
快速生成品牌宣传视频,适用于广告投放和社交媒体推广。
为抖音、小红书等平台创建引人注目的短视频内容。
生成产品功能演示和使用教程视频,提升用户转化率。
制作课程讲解、知识科普等教育类视频,降低视频制作门槛。
Digital Human 通过 APIMODELS 平台调用,当前定价:per second of audio: $0.01。按量计费,用多少付多少。
在 APIMODELS 注册账号并获取 API Key,然后通过我们的统一 API 端点调用即可。我们提供详细的 API 文档和 cURL、Python、Node.js 代码示例。
APIMODELS 通过聚合平台提供与官方相同的 Digital Human 模型。我们提供统一的 API 接口,无需分别注册各平台账号,一个 API Key 即可调用所有模型。
在 APIMODELS,Digital Human 与 60+ 个模型共用一个 API Key、一个余额,所以选型只看合不合适,不存在锁定。它支持 Video + Audio、Lip-Sync、Keeps Original Motion、Audio up to 30min、Per-Second Pricing,你可以在价格和能力上把它和其它视频生成模型对比,换模型只需改一个模型名字符串——无需新账号、无需重新对接。所有视频生成模型与实时价格见 apimodels.app/models。
Digital Human 支持:Video + Audio、Lip-Sync、Keeps Original Motion、Audio up to 30min、Per-Second Pricing。完整参数与调用方式见 APIMODELS 的 API 文档。
可以。APIMODELS 提供可直接访问的统一 API,一个 API Key 即可调用 Digital Human,无需分别注册官方账号、也无需自行处理官方接口的网络访问。
我们支持 Stripe(Visa、Mastercard 等国际信用卡)和支付宝付款。充值后积分即时到账。