Every model below is callable through one unified API and a single key. Click a model for its full playground and docs.
| Model | model (API) | Price | Details |
|---|---|---|---|
| Gemini 2.5 Flash Image | gemini-2.5-flash-image | $0.02 | Google's Gemini 2.5 Flash Image at a flat $0.02 per image — the cheapest way to generate or edit an image here. Text-to-image and reference-image editing, one API key, no Google Cloud setup. |
| GPT Image 2 | gpt-image-2 | $0.025+ | OpenAI gpt-image-2. Text-to-image and multi-image editing (up to 16 reference images), aspect-ratio control, native 1K / 2K / 4K — $0.025 / $0.03 / $0.05 per image. |
| Qwen Image 3.0 | qwen3-image | $0.035 | Alibaba Qwen Image 3.0 — text-to-image and image editing in one model. Up to 4.5k-token prompts, dense in-image text layout, 10px small-text rendering, 12 languages and 20+ fonts. 1K/2K, flat $0.035 per image; each reference image +$0.004. |
| Qwen Image 3.0 Pro | qwen3-image-pro | $0.037+ | Alibaba Qwen Image 3.0 Pro — the higher-fidelity tier. Up to 4.5k-token prompts, dense in-image typesetting for posters / storyboards / menus, 10px small text, micro-expressions, pores and hair strands near photographic realism, 12 languages and 20+ fonts. 1K $0.037 / 2K $0.075; each reference image +$0.004. |
| Real-ESRGAN Upscaler | real-esrgan | $0.004 | Real-ESRGAN AI upscaler — best for illustration, anime and hand-drawn artwork: enlarge to 4x HD, 8K and 10K print quality with crisp lines and clean flats. Optional face enhancement. $0.004 per image. |
| GPT Image 2 All | gpt-image-2-all | $0.01+ | OpenAI gpt-image-2, OpenAI-compatible sync API — drop-in for Codex / Cursor / the OpenAI SDK. Full 1K / 2K / 4K × low / medium / high quality, text-to-image and multi-image editing, priced per resolution × quality from $0.01. |
| Nanobananapro-gemini | nanobananapro-gemini | $0.03 | Gemini 3 Pro Image via a budget channel. Professional asset creation with advanced reasoning and high-fidelity text rendering. |
| Nanobanana2-gemini | nanobanana2-gemini | $0.025 | Gemini 3.1 Flash Image via a budget channel. High-performance image generation optimized for speed and high-volume use. |
| SparkPix Image | sparkpix-image | $0.008 | Sub 1 second text-to-image model built for production use cases. State-of-the-art speed, quality, and text rendering. |
| SparkPix Image Edit | sparkpix-image-edit | $0.013 | Sub 1 second multi-image editing model. Fast, affordable AI image editing with precise prompt adherence and multi-image support. |
| Kling V3 Image | kling-v3-image | $0.05 | Kling V3 image generation. Text-to-image and single-reference image-to-image, 1K/2K resolution. $0.05 per image. |
| Kling V3 Omni | kling-v2-new | $0.05+ | Kling V3 Omni image generation. Multi-image reference & fusion, element consistency, single/series output, 1K/2K/4K — 1K/2K $0.05, 4K $0.10 per image. |
| Kling Omni-Image | kling-image-o1 | $0.05 | AI image generation and editing by Kling (omni-image, model kling-image-o1). Supports 1K/2K resolution and multi-image input. $0.05 per image. |
| Doubao Seedream 5.0 Lite (Official) | doubao-seedream-5-0-260128 | $0.055 | Doubao Seedream 5.0 Lite via ByteDance Volcano Ark official API. Unified text-to-image and image-to-image (pass image for I2I, omit for T2I). 2K / 4K output, no watermark, PNG. |
| Doubao Seedream 5.0 Pro (Official) | doubao-seedream-5-0-pro | $0.076 / $0.15 | Doubao Seedream 5.0 Pro (Volcano Ark official) — flagship tier with interactive editing (coordinates / box-select / arrows / hand-drawn marks), layer separation, and strong multilingual in-image text rendering. Text-to-image, single- and multi-image (2–10) fusion editing, 1K / 2K single-image output, no watermark. |
| Doubao Seedream 4.5 | doubao-seedream-4-5-251128 | $0.05 | High quality Doubao Seedream 4.5 image generation. Supports text-to-image and image editing with 2K/4K resolution. |
| Grok 4.2 Image | grok-4.2-image | $0.0075 | The budget tier for Grok image generation — $0.0075 per image, a quarter of the standard Grok Imagine Image price, on the same model and the same endpoint. Text-to-image plus single-reference editing, 19s median. Switch tiers by changing one field. |
| Grok Imagine Image | grok-imagine-image | $0.03 | Multimodal AI image generation by X platform. Generates high-quality images from text descriptions. |
| Grok Imagine Image Pro | grok-imagine-image-pro | from $0.08 | Upgraded multimodal AI model by X platform with stronger understanding and finer detail generation for higher precision images. |
| Grok Imagine Image 2.0 | grok-imagine-image-2 | from $0.03 | xAI's Image 2.0, built for images you can ship: instruction-following down to the details, designer-grade typography and layout, and multi-reference editing. Four price tiers from $0.03 — 1K/2K at low or medium quality. |
| Nanobanana2 | nanobanana2 | $0.05+ | Fast image generation powered by Gemini 3.1 Flash. Supports text-to-image and image editing — 1K/2K $0.05, 4K $0.08 per image. |
| gemini-3.1-flash-image | gemini-3.1-flash-image | $0.04+ | Google's gemini-3.1-flash-image (GA). 2K images cost $0.06 versus Google's $0.101 — 40% less — and 1K is the same $0.06, so 2K is a free upgrade here. 512 $0.04, 4K $0.10. Half of requests return within 14s. |
| gemini-3-pro-image | gemini-3-pro-image | $0.10+ | Google's gemini-3-pro-image (GA release). Top-quality, high-fidelity image generation and editing with advanced reasoning. Priced by resolution: 1K/2K $0.10 (25% cheaper than official), 4K $0.15 (37.5% cheaper than official). |
| Nanobanana-2-lite | nanobanana-2-lite | $0.025 | The cheapest Gemini 3.1 Flash Image tier — flat $0.025 per image, no resolution tiers. Text-to-image and editing (up to 10 reference images), 1K output only. |
| Gemini 3 Pro Image (Pro) | nanobananapro | $0.06+ | Premium image generation powered by Gemini 3 Pro. 99% success rate. Best quality and reliability. |
Per image and per model: Nano Banana from ~$0.02, GPT Image 2 $0.025 (1K) / $0.04 (2K) / $0.06 (4K), Nano Banana Pro from $0.03. You are charged only on success; new accounts start with $0.1 free.
POST /api/v1/images/generations with a prompt plus image (single) or images[] (up to 16 for GPT Image 2). Async: poll GET ?task_id= until completed. Full cURL/Python/Node examples are on this page.
Yes. Every image model uses the same endpoint and key — switch models by the model parameter. No per-provider signup. Cheaper than each vendor’s official API.