Every model below is callable through one unified API and a single key. Click a model for its full playground and docs.
| Model | model (API) | Price | Details |
|---|---|---|---|
| Kling Custom Voice | kling-custom-voice | $0.006 / per call | Create custom voice profiles from audio samples. Upload .mp3/.wav/.mp4/.mov (5-30s) or reference a video ID. |
| Kling Face Recognition | kling-identify-face | $0.01 / per call | Identify faces in a video and return a session ID and face IDs for Kling lip-sync video generation. |
| Kling Sound Effects | kling-sound-effects | $0.040 / per call | Generate sound effects from text descriptions. 3-10 second audio with natural quality. |
| Kling Video-to-Audio | kling-video-to-audio | $0.003 / per call | Auto-generate sound effects and background music for videos. Supports ASMR mode for immersive content. |
| Kling TTS | kling-tts | $0.01 / per call | Text-to-speech with multiple voice options. Adjustable speed and multi-language support. |
| Minimax Speech 2.8 HD | minimax-speech-2.8-hd | $0.07 / 1K chars | Latest high-fidelity TTS by MiniMax (海螺). Predicts emotion and intonation from context for ultra-natural, expressive, personalized speech. Supports voice clone and voice design. |
| Minimax Speech 2.8 Turbo | minimax-speech-2.8-turbo | $0.04 / 1K chars | Latest fast, cost-effective async TTS by MiniMax (海螺). Great quality-to-price for high-volume synthesis. Supports voice clone and voice design. |
| Minimax Speech 2.6 HD | minimax-speech-2.6-hd | $0.07 / 1K chars | High-definition async TTS by Minimax (海螺). Rich expressiveness with natural prosody. Supports voice clone and voice design. |
| Minimax Speech 02 HD | minimax-speech-02-hd | $0.07 / 1K chars | High-fidelity TTS by MiniMax (海螺). Predicts emotion and intonation from context to produce ultra-natural, expressive, personalized speech — built for social, podcasts, audiobooks, news, education and digital humans. Supports voice clone and voice design. |
| Minimax Speech 02 Turbo | minimax-speech-02-turbo | $0.04 / 1K chars | Fast and cost-effective async TTS by Minimax (海螺). Supports voice clone, voice design, and pronunciation dictionaries. |
| Eleven Flash v2.5 | eleven-tts-flash | $0.0425 / 1K chars | Ultra low latency model in 32 languages. Ideal for real-time conversational use cases. |
| Eleven Turbo v2.5 | eleven-tts-turbo | $0.0425 / 1K chars | High quality, low latency model in 32 languages. Best for developer use cases where speed matters. |
| Eleven Multilingual v2 | eleven-tts-multilingual | $0.085 / 1K chars | Most life-like, emotionally rich mode in 29 languages. Best for voice overs, audiobooks, post-production. |
| Suno Music (v4 & v5) | suno-v5 | $0.26 / per song (2 variants) | AI music generation — full songs with vocals + lyrics from a one-line idea (inspiration), your own lyrics (custom), or instrumental only. Also sound effects, continue, cover, and voice personas. Each run returns 2 variants. |
| Eleven v3 | eleven-tts-v3 | $0.085 / 1K chars | Most expressive model with 70+ languages. Supports audio tags like [laughs], [whispers] for emotional control. |
| ElevenLabs Dialogue | eleven-dialogue | $0.085 / 1K chars | Multi-speaker dialogue generation with natural conversation flow. Perfect for podcasts and audiobooks. |
| Voice Isolator | eleven-isolator | $0.102 / min | Extract speech from background noise, music and ambient sounds. Clean audio extraction. |
| AI Dubbing | eleven-dubbing | $0.2805 / min | Translate audio/video while preserving emotion, timing and tone. Automatic lip-sync. |
ElevenLabs and MiniMax TTS, Suno v5 music generation (dedicated page), plus Kling TTS, sound effects, video-to-audio and dubbing — all via one endpoint and key.
Use model suno-v5 on POST /api/v1/audio/generations; the mode (inspiration / custom / cover / hum-to-song / …) is auto-detected from your fields. See the dedicated Suno docs page.