
rh-lip-sync-videoKling Lip-Sync Video does frame-level lip synchronization, aligning an audio track to the mouth movements of a character in a video — real humans, 3D and 2D animated characters — with local audio upload or online TTS and minute-level duration. The typical flow is to run Kling Face Recognition first to get a faceId, then align audio (uploaded or from Kling Lip-Sync TTS) to that face. Ideal for digital-human voiceover, dub-to-lip-sync and talking animated characters.
Precise lip-audio alignment
Real human, 3D, 2D support
Upload or online TTS
Minute-level video generation
Kling lip-sync is a 3-step flow — run them in order; intermediate values carry forward automatically.
Upload or paste a public video URL; recognition returns sessionId + faceId.
Trimmed audio must be ≥2s; the insert window must overlap the face window by ≥2s.
Kling Lip-Sync Video は Kling の動画生成 API です。Kling Lip-Sync Video does frame-level lip synchronization, aligning an audio track to the mouth movements of a character in a video — real humans, 3D and 2D animated characters — with local audio upload or online TTS and minute-level duration. The typical flow is to run Kling Face Recognition first to get a faceId, then align audio (uploaded or from Kling Lip-Sync TTS) to that face. Ideal for digital-human voiceover, dub-to-lip-sync and talking animated characters. APIMODELS のプラットフォーム経由なら、統一 API と明朗な従量課金でこのモデルを呼び出せます。 現在の料金: per 5s: $0.065。
広告キャンペーンや SNS 施策に向けたブランド動画を短時間で作ります。
TikTok、Instagram、YouTube 向けの縦型ショート動画を量産できます。
機能紹介やチュートリアル動画を作り、コンバージョンにつなげます。
講座の解説、知識の説明、研修用の動画を低コストで継続的に作れます。
Kling Lip-Sync Video は APIMODELS 経由で per 5s: $0.065 で利用できます。課金は従量制で、生成した分だけの支払いです。
APIMODELS に登録して API キーを取得し、統一エンドポイントを呼ぶだけです。cURL / Python / Node.js のサンプルを含む詳しいドキュメントを用意しています。
APIMODELS は同じ Kling Lip-Sync Video を集約プラットフォーム経由で提供します。統一された API インターフェースなので、プロバイダごとにアカウントを作る必要はありません。キー 1 本ですべてのモデルに届きます。
It does frame-level lip synchronization — aligning an audio track to the mouth movements of a character in a video, for real humans, 3D and 2D animated characters, with local audio upload or online TTS and minute-level duration. Good for digital-human voiceover, dub-to-lip-sync, and talking animated characters.
Typical flow: first run Kling Face Recognition (kling-identify-face) to detect a face in the video and get a faceId, then align audio (uploaded or generated via Kling Lip-Sync TTS) to that face to produce the lip-synced video.
APIMODELS では Kling Lip-Sync Video が 60 以上のモデルと同じ API キー・同じ残高の上に並んでいるので、選択は「相性」の問題であって「囲い込み」の問題ではありません。Lip Sync、Multi-Character、Audio Alignment、Minute-Level Duration に対応しており、他の動画生成モデルと価格・機能を並べて評価できます。乗り換えはモデル名の文字列を 1 つ書き換えるだけ。新しいアカウントも追加の実装も要りません。動画生成の選択肢と最新価格は apimodels.app/models で確認できます。
Kling Lip-Sync Video は次に対応しています: Lip Sync、Multi-Character、Audio Alignment、Minute-Level Duration。パラメータの全一覧と呼び出し例は APIMODELS のドキュメントをご覧ください。
はい。APIMODELS は Kling Lip-Sync Video を単一の統一 API とキー 1 本で提供します。プロバイダごとのアカウントも、各社の地域ごとのネットワーク経路を自分で面倒みる必要もありません。
Stripe(Visa、Mastercard などの国際カード)と Alipay に対応しています。支払い後、残高はすぐ反映されます。