xAI
xAI's newest frontier model, built for long-running agents and multi-step work — it matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index. 500K context, text + image input, four reasoning-effort levels, and function calling verified at 12/12 under concurrency with nested schemas. 12% below xAI official ($2 / $6 per 1M).
xAI
xAI's Image 2.0, built for images you can ship: instruction-following down to the details, designer-grade typography and layout, and multi-reference editing. Four price tiers from $0.03 — 1K/2K at low or medium quality.
Alibaba
Alibaba Qwen Image 3.0 — text-to-image and image editing in one model. Up to 4.5k-token prompts, dense in-image text layout, 10px small-text rendering, 12 languages and 20+ fonts. 1K/2K, flat $0.035 per image; each reference image +$0.004.
Alibaba
Alibaba Qwen Image 3.0 Pro — the higher-fidelity tier. Up to 4.5k-token prompts, dense in-image typesetting for posters / storyboards / menus, 10px small text, micro-expressions, pores and hair strands near photographic realism, 12 languages and 20+ fonts. 1K $0.037 / 2K $0.075; each reference image +$0.004.
ByteDance
ByteDance Seedance 2.5 on the direct official Volcengine Ark API — the long-form and editing tier: 4-30s in a single pass, up to 50 reference assets (30 images + 10 videos + 10 audios), plus video editing and video extension. 480p / 720p (no 1080p — that stays on Seedance 2.0), mp4 or mov, native synced audio and real-person video via the asset library. From $0.134/s.
OpenAI
OpenAI gpt-image-2. Text-to-image and multi-image editing (up to 16 reference images), aspect-ratio control, native 1K / 2K / 4K — $0.025 / $0.03 / $0.05 per image.
OpenAI
OpenAI gpt-image-2, OpenAI-compatible sync API — drop-in for Codex / Cursor / the OpenAI SDK. Full 1K / 2K / 4K × low / medium / high quality, text-to-image and multi-image editing, priced per resolution × quality from $0.01.
Premium image generation powered by Gemini 3 Pro. 99% success rate. Best quality and reliability.
Gemini 3 Pro Image via a budget channel. Professional asset creation with advanced reasoning and high-fidelity text rendering.
Google's gemini-3-pro-image (GA release). Top-quality, high-fidelity image generation and editing with advanced reasoning. Priced by resolution: 1K/2K $0.10 (25% cheaper than official), 4K $0.15 (37.5% cheaper than official).
Fast image generation powered by Gemini 3.1 Flash. Supports text-to-image and image editing — 1K/2K $0.05, 4K $0.08 per image.
Gemini 3.1 Flash Image via a budget channel. High-performance image generation optimized for speed and high-volume use.
The cheapest Gemini 3.1 Flash Image tier — flat $0.025 per image, no resolution tiers. Text-to-image and editing (up to 10 reference images), 1K output only.
The newest Flash generation — same input price as 3.5 Flash with cheaper output, for agentic and long-horizon work where the answer, not the prompt, dominates the bill.
Google's gemini-3.1-flash-image (GA). 2K images cost $0.06 versus Google's $0.101 — 40% less — and 1K is the same $0.06, so 2K is a free upgrade here. 512 $0.04, 4K $0.10. Half of requests return within 14s.
Google's Gemini 2.5 Flash Image at a flat $0.02 per image — the cheapest way to generate or edit an image here. Text-to-image and reference-image editing, one API key, no Google Cloud setup.
Doubao
Doubao Seedream 5.0 Pro (Volcano Ark official) — flagship tier with interactive editing (coordinates / box-select / arrows / hand-drawn marks), layer separation, and strong multilingual in-image text rendering. Text-to-image, single- and multi-image (2–10) fusion editing, 1K / 2K single-image output, no watermark.
Anthropic
Anthropic's newest Opus — a step change on deep reasoning, agentic coding and long-horizon work, at 40% off official pricing. 1M context, adaptive thinking on by default, native tool use. Works with Claude Code, Cursor and the OpenAI SDK alike.
Doubao
Doubao Seedream 5.0 Lite via ByteDance Volcano Ark official API. Unified text-to-image and image-to-image (pass image for I2I, omit for T2I). 2K / 4K output, no watermark, PNG.
Doubao
High quality Doubao Seedream 4.5 image generation. Supports text-to-image and image editing with 2K/4K resolution.
Kling
Kling V3 image generation. Text-to-image and single-reference image-to-image, 1K/2K resolution. $0.05 per image.
Kling
Kling V3 Omni image generation. Multi-image reference & fusion, element consistency, single/series output, 1K/2K/4K — 1K/2K $0.05, 4K $0.10 per image.
Kling
AI image generation and editing by Kling (omni-image, model kling-image-o1). Supports 1K/2K resolution and multi-image input. $0.05 per image.
xAI
Multimodal AI image generation by X platform. Generates high-quality images from text descriptions.
xAI
Upgraded multimodal AI model by X platform with stronger understanding and finer detail generation for higher precision images.
SparkPix
Sub 1 second text-to-image model built for production use cases. State-of-the-art speed, quality, and text rendering.
SparkPix
Sub 1 second multi-image editing model. Fast, affordable AI image editing with precise prompt adherence and multi-image support.
Real-ESRGAN
Real-ESRGAN AI upscaler — best for illustration, anime and hand-drawn artwork: enlarge to 4x HD, 8K and 10K print quality with crisp lines and clean flats. Optional face enhancement. $0.004 per image.
MiniMax
MiniMax Hailuo H3 — native 2K video with synchronized audio from one multimodal request. Prompt alone for text-to-video, or attach up to 9 reference images, 3 videos and 3 audio clips to steer character, motion and voice. Any length from 5 to 15 seconds, billed per second.
xAI
Grok Video 3. Per-second pricing $0.02/s, 6 / 10 / 15 second output. Both T2V (omit images) and I2V (1-7 reference images) supported.
xAI
Grok Video 3. Fixed 10-second clip at a flat $0.20 each. Supports text-to-video and image-to-video (up to 7 reference images).
ByteDance
ByteDance Seedance 2.0 on the direct official Volcengine Ark API — full multimodal generation from text, a first / first+last frame, up to 9 reference images, reference video, reference audio and web search, with native synced audio. Real-person video via the asset library. 480p / 720p / 1080p, 4-15s, per-second pricing from $0.092/s.
ByteDance
The speed- and cost-optimised tier of Seedance 2.0 on the direct official Volcengine Ark API — same full multimodal capability (text / first+last frame / up to 9 reference images / reference video / reference audio / web search / native audio), real person via the asset library. 480p / 720p, 4-15s, from $0.071/s.
OpenAI
GPT-5.6 Sol is the frontier model in the GPT-5.6 family — OpenAI's highest-intelligence tier (the gpt-5.6 alias routes to Sol), roughly the unsuffixed top tier of earlier GPT-5 families. Reasoning-token support, text + image input, 1.05M-token context, 128K max output, Feb 2026 knowledge cutoff. Call it on /v1/chat/completions or /v1/responses — one key, ~78% below OpenAI list, reachable from China.
OpenAI
GPT-5.6 Terra balances intelligence and cost — the mini tier of the GPT-5.6 family, priced at exactly half of Sol. Higher reasoning, fast, text + image input, text output, 1.05M-token context, 128K max output, Feb 2026 knowledge cutoff, reasoning-token support. Call it on /v1/chat/completions or /v1/responses — one key, ~78% below OpenAI list, reachable from China.
OpenAI
GPT-5.6 Luna is optimized for cost-sensitive, high-volume workloads — the nano tier of the GPT-5.6 family, priced at 20% of Sol. High reasoning, fast, text + image input, text output, 1.05M-token context, 128K max output, Feb 2026 knowledge cutoff, reasoning-token support. Served as gpt-5.6-luna-max on /v1/chat/completions or /v1/responses — one key, pay-as-you-go, reachable from China.
ByteDance
The cheapest Seedance 2.0 tier on the direct official Volcengine Ark API — same multimodal inputs (text / first+last frame / reference images, video and audio) and native synced audio at 480p / 720p, 4-15s, from $0.044/s. Best for high-volume drafts and cost-sensitive batch generation.
xAI
xAI's newest flagship model — leads the industry in coding, non-hallucination rate, agentic tool calling, and instruction following. Supports both non-reasoning and reasoning modes, 500K context. Cheaper than xAI official ($2 / $6 per 1M tokens): 25% off input, ~17% off output.
ByteDance
Seedance 2.0 sd-A — full multimodal HD video (720p / 1080p) with an in-playground tier switch (2.0 / Fast / Mini): text, first/last frame, up to 9 reference images + video + audio, web search, native audio, and asset-library characters.
ByteDance
Film-grade edition of Seedance 2.0 — cinematic lighting, mood and camera motion, and it ACCEPTS real-person / realistic human reference images (unlike Ark-direct Seedance 2.0, which rejects real faces). Up to 4 reference images for identity-locked image-to-video — ideal for film-grade portrait and character work. Quality tier sits above the Ark variants; generation takes longer (typically 60-180s). Use only with consented subjects.
Suno
AI music generation — full songs with vocals + lyrics from a one-line idea (inspiration), your own lyrics (custom), or instrumental only. Also sound effects, continue, cover, and voice personas. Each run returns 2 variants.
ByteDance
Seedance 2.0, international line — true native resolution up to 4K (10-bit), no upscaling. Full multimodal: text, first/last frame, up to 9 reference images + video + audio, native synced audio. 480p / 720p / 1080p / 4K, 4-15s, per-second pricing.
ByteDance
The faster, cheaper Seedance 2.0 tier — same full multimodal capability (text / first+last frame / up to 9 reference images / video / audio / native audio) at native 480p or 720p. 4-15s, per-second pricing.
ByteDance
The most cost-effective Seedance 2.0 tier — full multimodal (text / first+last frame / up to 9 reference images / video / audio / native audio) at native 480p or 720p. 4-15s, per-second pricing.
Google VEO 3.1 (standard tier) at full 1080p — the premium, highest-quality VEO. Text-to-video and reference-image support, 8s clip. For a budget 1080p option use VEO 3.1 Fast Full HD ($0.07).
VEO 3.1 Fast HD (720p) video generation. 8s fixed duration, 16:9 aspect ratio, reference image support.
Anthropic
Anthropic's latest and most capable public LLM — about 50% cheaper than official. Works with Claude Code & Cursor via the Anthropic API. Ultra-long context, multimodal, complex reasoning, strict safety.
Anthropic
Anthropic's newest Sonnet — 1M-token context (default & max), 128K max output, adaptive thinking, and the same tools & platform features as Sonnet 4.6 (Priority Tier not supported). Works with Claude Code & Cursor via the Anthropic API.
VEO 3.1 Fast Full HD (1080p) video generation. 8s fixed duration, 16:9 aspect ratio, reference image support.
Smooth cinematic transitions between a required first frame and required last frame. Outputs 720p or 1080p with native audio. Official stable channel — pricier than V3.1-fast but reliable, ideal for production.
APIMODELS
Turn a portrait image + an audio clip into a talking-head video. The output length follows the audio; billed per second.
ElevenLabs
Ultra low latency model in 32 languages. Ideal for real-time conversational use cases.
ElevenLabs
High quality, low latency model in 32 languages. Best for developer use cases where speed matters.
ElevenLabs
Most life-like, emotionally rich mode in 29 languages. Best for voice overs, audiobooks, post-production.
ElevenLabs
Most expressive model with 70+ languages. Supports audio tags like [laughs], [whispers] for emotional control.
ElevenLabs
Multi-speaker dialogue generation with natural conversation flow. Perfect for podcasts and audiobooks.
ElevenLabs
Extract speech from background noise, music and ambient sounds. Clean audio extraction.
ElevenLabs
Translate audio/video while preserving emotion, timing and tone. Automatic lip-sync.
APIMODELS
Remove hardcoded subtitles, burned-in on-screen text, watermarks and corner logos from any video, rebuilding a clean background where they used to be. High-frame-rate OCR detection plus AIGC inpainting for near-invisible results. Let it find the text automatically, or switch to region mode and box exactly what should go — the way to remove a watermark or station bug. Output resolution matches the input (up to 1080p). Billed $0.015 per second of video.
Lightricks
LTX-2.3 unified text-to-video and image-to-video. Send a prompt for T2V, or add one reference image for I2V — fast, with 480p / 720p / 1080p output. Billed per second by resolution.
Omni Flash (Stable) — lower-cost, full-suite Gemini Omni video. Text / image (up to 7 refs) / video-to-video, plus reusable voices and consistent characters. 720p / 1080p / 4k, 6 / 8 / 10s, 16:9 or 9:16, optional seed.
Gemini Omni Flash — unified video generator for text-to-video, image-to-video (1 or 3 reference images) and video editing. 720p / 1080p / 4k, 6 / 8 / 10s, optional 16:9 or 9:16 framing. Per-second pricing: 720p $0.06/s, 1080p $0.065/s, 4k $0.12/s (video edit is billed on the reference clip length).
ByteDance
ByteDance DreamActor V2 motion transfer. Drive any character image with reference video motion, supporting multi-person, anime and pets.
DeepSeek
DeepSeek V4 Flash — the lightweight, high-throughput, cost-effective member of the DeepSeek V4 family for general chat and basic text. 1M-token context, tool calling, streaming; OpenAI-compatible.
DeepSeek
DeepSeek V4 Pro — DeepSeek’s high-performance model with top-tier reasoning and agent capabilities, 1M-token context, and full thinking (reasoning_content) output. Tool calling, streaming; OpenAI-compatible.
Alibaba
Alibaba Qwen3.7 Max — the most capable Qwen3.7 model: top-tier reasoning and agent ability, 1M-token context, hybrid thinking. Tool calling, streaming; OpenAI-compatible.
Alibaba
Alibaba Qwen3.7 Plus — the best-value Qwen3.7 model: strong reasoning at a fraction of Max’s price, 1M-token context, hybrid thinking. Tool calling, streaming; OpenAI-compatible.
Kling
Kling AI lip-sync video generation. Frame-level lip synchronization with audio for real humans, 3D and 2D characters.
Kling
Kling text-to-speech synthesis with multi-language support, voice cloning, speed control and emotion styles.
OpenAI
Frontier model for complex professional work and agentic coding. Served via the Responses API with adjustable reasoning effort (low–xhigh), web search, and function calling.
Zhipu
Zhipu GLM-5.2 — a reasoning model with strong function-calling / tool-use, served via the OpenAI-compatible chat-completions endpoint.
Kling
Generate videos with character motion control. Provide a reference image and motion video to create animated content.
OpenAI
OpenAI’s advanced reasoning model for agentic coding, knowledge work, scientific research, and complex multi-step task execution. Served via the Responses API with adjustable reasoning effort (low–xhigh), web search, and function calling.
Kling
Kling V3 Omni-Video with extended duration and keep-original-sound support for video editing. Flat $0.15/s billing.
Kling
Latest Kling V3 video generation. Supports 3-15s flexible duration, text-to-video and image-to-video with optional audio.
Pruna AI
Fast video generation in ~10 seconds. Text/image/audio-to-video with draft mode for 4x faster previews. Built-in audio generation, up to 1080p 48FPS.
Minimax
Latest high-fidelity TTS by MiniMax (海螺). Predicts emotion and intonation from context for ultra-natural, expressive, personalized speech. Supports voice clone and voice design.
Minimax
Latest fast, cost-effective async TTS by MiniMax (海螺). Great quality-to-price for high-volume synthesis. Supports voice clone and voice design.
Minimax
High-fidelity TTS by MiniMax (海螺). Predicts emotion and intonation from context to produce ultra-natural, expressive, personalized speech — built for social, podcasts, audiobooks, news, education and digital humans. Supports voice clone and voice design.
Anthropic
Anthropic's most capable model yet — built to autonomously carry long, complex work end to end. Ideal for big projects, building agents, and high-stakes scenarios demanding top quality and autonomy.
Anthropic
Latest Opus model with 1M context, 128K max output, and adaptive thinking — same tools and platform features as Opus 4.6.
Anthropic
Claude Opus 4.7 with extended thinking explicitly enabled for the most complex reasoning tasks.
GA release. Our most intelligent Flash model — consistent leadership on agentic execution, coding, and long-horizon tasks at scale.
Most cost-effective multimodal model with fastest performance for high-frequency lightweight tasks.
Latest Pro model with enhanced reasoning and multimodal capabilities.
Kling
Create custom voice profiles from audio samples. Upload .mp3/.wav/.mp4/.mov (5-30s) or reference a video ID.
Kling
Identify faces in a video and return a session ID and face IDs for Kling lip-sync video generation.
Anthropic
Latest Opus model with ultimate performance and reasoning capabilities.
Anthropic
Claude Opus 4.6 with extended thinking capability for the most complex reasoning tasks.
Anthropic
Latest Sonnet model with best performance and efficiency.
Anthropic
Claude Sonnet 4.6 with extended thinking capability for complex reasoning tasks.
Kling
Generate sound effects from text descriptions. 3-10 second audio with natural quality.
Kling
Auto-generate sound effects and background music for videos. Supports ASMR mode for immersive content.
Kling
Text-to-speech with multiple voice options. Adjustable speed and multi-language support.
Minimax
High-definition async TTS by Minimax (海螺). Rich expressiveness with natural prosody. Supports voice clone and voice design.
Minimax
Fast and cost-effective async TTS by Minimax (海螺). Supports voice clone, voice design, and pronunciation dictionaries.
xAI
The budget tier for Grok image generation — $0.0075 per image, a quarter of the standard Grok Imagine Image price, on the same model and the same endpoint. Text-to-image plus single-reference editing, 19s median. Switch tiers by changing one field.
Fast and efficient multimodal model. Great for quick responses and simple tasks.
Advanced multimodal reasoning model with superior capabilities.
Gemini 3 Pro with extended thinking capability for complex reasoning tasks.
Anthropic
Latest Opus model with enhanced capabilities and improved reasoning.
Anthropic
Claude Opus 4.5 with extended thinking capability for the most complex reasoning tasks.
Anthropic
Fast and affordable model for lightweight tasks. Best for simple queries and quick responses.
Anthropic
Claude Haiku 4.5 with extended thinking capability for complex reasoning tasks.
Anthropic
Latest Sonnet model with improved performance and efficiency.
Anthropic
Claude Sonnet 4.5 with extended thinking capability for complex reasoning tasks.
Anthropic
Most capable model with superior reasoning and analysis capabilities.
Anthropic
Claude Opus 4 with extended thinking capability for the most complex reasoning tasks.
Anthropic
Balanced model with excellent performance and cost efficiency. Great for most tasks.
Anthropic
Claude Sonnet 4 with extended thinking capability for complex reasoning tasks.
OpenAI
Small embedding model, efficient and cost-effective for most use cases.
OpenAI
Large embedding model for higher accuracy and flexible dimensions.