digital-humanDigital Human turns a video of a person plus a driving audio clip into a lip-synced talking video: the person keeps their original look, motion and background while the mouth is re-animated to speak the audio. If the audio runs longer than the video, the footage is extended automatically — the output always follows the audio length. Send two URLs (video_url: mp4/mov of the person; audio_url: mp3/wav/m4a, up to 30 minutes) and get an mp4 back at the source video's resolution. Billing is $0.01 per second of the driving audio — a 60-second voiceover costs $0.60, half the per-second price of image-based lip-sync. Use AI Lip-Sync instead when you only have a still portrait image. Typical turnaround is about 2 minutes for short clips. Only successful generations are charged; results are kept for 7 days.
$0.01 per second of audio · if the audio outlasts the video, the footage is extended automatically
Your talking video will appear here
$0.01 per second of audio vs $0.02/s for the image-based ai-lipsync — when you already have footage of the person, this is the cheaper path
Lighting, motion, background and camera work come from your footage; only the mouth is re-animated to match the new audio
Audio longer than the video? The footage is extended automatically — verified with a 6.2s voiceover over a 5s clip
Output matches the source video resolution (tested at 1440×2560) — no forced downscale
POST /api/v1/video/generations with model digital-human, video_url and audio_url — poll the task id for the mp4
Digital Human は APIMODELS の動画生成 API です。Digital Human turns a video of a person plus a driving audio clip into a lip-synced talking video: the person keeps their original look, motion and background while the mouth is re-animated to speak the audio. If the audio runs longer than the video, the footage is extended automatically — the output always follows the audio length. Send two URLs (video_url: mp4/mov of the person; audio_url: mp3/wav/m4a, up to 30 minutes) and get an mp4 back at the source video's resolution. Billing is $0.01 per second of the driving audio — a 60-second voiceover costs $0.60, half the per-second price of image-based lip-sync. Use AI Lip-Sync instead when you only have a still portrait image. Typical turnaround is about 2 minutes for short clips. Only successful generations are charged; results are kept for 7 days. APIMODELS のプラットフォーム経由なら、統一 API と明朗な従量課金でこのモデルを呼び出せます。 現在の料金: per second of audio: $0.01。
広告キャンペーンや SNS 施策に向けたブランド動画を短時間で作ります。
TikTok、Instagram、YouTube 向けの縦型ショート動画を量産できます。
機能紹介やチュートリアル動画を作り、コンバージョンにつなげます。
講座の解説、知識の説明、研修用の動画を低コストで継続的に作れます。
Digital Human は APIMODELS 経由で per second of audio: $0.01 で利用できます。課金は従量制で、生成した分だけの支払いです。
APIMODELS に登録して API キーを取得し、統一エンドポイントを呼ぶだけです。cURL / Python / Node.js のサンプルを含む詳しいドキュメントを用意しています。
APIMODELS は同じ Digital Human を集約プラットフォーム経由で提供します。統一された API インターフェースなので、プロバイダごとにアカウントを作る必要はありません。キー 1 本ですべてのモデルに届きます。
APIMODELS では Digital Human が 60 以上のモデルと同じ API キー・同じ残高の上に並んでいるので、選択は「相性」の問題であって「囲い込み」の問題ではありません。Video + Audio、Lip-Sync、Keeps Original Motion、Audio up to 30min、Per-Second Pricing に対応しており、他の動画生成モデルと価格・機能を並べて評価できます。乗り換えはモデル名の文字列を 1 つ書き換えるだけ。新しいアカウントも追加の実装も要りません。動画生成の選択肢と最新価格は apimodels.app/models で確認できます。
Digital Human は次に対応しています: Video + Audio、Lip-Sync、Keeps Original Motion、Audio up to 30min、Per-Second Pricing。パラメータの全一覧と呼び出し例は APIMODELS のドキュメントをご覧ください。
はい。APIMODELS は Digital Human を単一の統一 API とキー 1 本で提供します。プロバイダごとのアカウントも、各社の地域ごとのネットワーク経路を自分で面倒みる必要もありません。
Stripe(Visa、Mastercard などの国際カード)と Alipay に対応しています。支払い後、残高はすぐ反映されます。