# API Models > API Models (apimodels.app) is an AI model aggregation API that lets developers call image, video, audio, and LLM models from many providers through one unified API and a single balance — usually cheaper than going direct, and usable from China and worldwide. ## Key Pages - [Models](https://apimodels.app/models): Browse 135 image / video / audio / LLM models, each with pricing and a one-line API call. - [Pricing](https://apimodels.app/pricing): Per-call and per-token pricing for every model. - [Docs](https://apimodels.app/docs): API reference, code samples (cURL / Python / Node), and CLI integration guides (OpenAI Codex, Claude Code). - [Unified API](https://apimodels.app/unified-api): One key and one endpoint for image, video, and LLM models. - [Cost calculator](https://apimodels.app/tools): Estimate what a workload costs here versus calling each provider directly. ## Key Facts - Category: AI model aggregator / unified AI API - Audience: developers, indie hackers, and startups building AI features - Main use cases: image generation & editing, video generation, text-to-speech / audio, LLM chat & agents (OpenAI / Claude / Gemini / Grok / GLM compatible) - Models: 135 across image (33), video (40), LLM (42), audio (18) and embeddings (2) - Billing: a single USD balance, pay-per-use; prices typically below official; $0.10 of trial credit on signup; failed calls are never charged - Access: one API key for all models; works from China and globally - Endpoints: LLM POST /api/v1/messages (+ /chat/completions, /responses); Image POST /api/v1/images/generations; Video POST /api/v1/video/generations; Audio POST /api/v1/tts/stream; Embeddings POST /api/v1/embeddings - Result files are kept 7 days, then deleted — download or re-host anything you need to keep - Last updated: 2026-09-16 ## Newest Models September 2026 flagships. Prices are USD per 1M tokens; reasoning tokens count as output. - [GPT-6 Astra](https://apimodels.app/models/gpt-6-astra): OpenAI's September 2026 flagship, added 2026-09-05. $2.40 input / $12.00 output / $0.24 cached input — 76% below OpenAI's own $10 / $50 list. Built for computer use and long-horizon agents: 72.6% on OSWorld 2.0 and 57.7% on Terminal-Bench 4.0, against GPT-5.6 Sol's 65.7% and 37.3%. 1.05M context, 128K output, text + vision in. Requests above 272,000 input tokens are billed at 2x input / 1.5x output. - [Claude Fable 5.1](https://apimodels.app/models/claude-fable-5-1): Anthropic's September 2026 flagship, added 2026-09-03. $5.00 input / $25.00 output / $0.22 cached read — exactly half Anthropic's $10 / $50 list. More than doubles Fable 5 on agentic science (52.6% vs 24.7% on Terminal-Bench-Science). 1M context, served on the native Anthropic Messages API, so Claude Code, the Anthropic SDK and Cursor work unchanged — only the base URL and key change. - [Gemini 3.8 Flash](https://apimodels.app/models/gemini-3.8-flash): Google's September 2026 Flash, added 2026-09-03. $0.450 input / $2.250 output / $0.172 cached input — 40% below Google's own $0.75 / $3.75 list. Tops DeepSWE v1.1 and scores 89.4% on Terminal-Bench 2.1; beats Claude Opus 5 on HLE-Verified, Vals Finance Agent v2 and Harvey's Legal Agent. 1M context, built for coding and agents. - [GLM-5.3](https://apimodels.app/models/glm-5.3): Z.ai's newest reasoning model, added 2026-08-26. $1.33 input / $4.18 output per 1M tokens. 1M-token context, strong function calling, built for long-horizon agent work and complex software engineering. OpenAI-compatible chat-completions endpoint. The GPT-5.6 family remains the volume tier, callable on both /api/v1/chat/completions and /api/v1/responses. Reasoning depth is a model-name suffix (-low / -medium / -high / -xhigh / -max / -ultra), all at the same price. - [GPT-5.6 Sol](https://apimodels.app/models/gpt-5-6-sol): frontier tier. $1.324 input / $6.618 output / $0.132 cached input. The `gpt-5.6` alias routes here. About 67% below OpenAI's own $4 / $20 list (OpenAI cut GPT-5.6 list prices on 2026-08-21). - [GPT-5.6 Terra](https://apimodels.app/models/gpt-5-6-terra): balanced mini tier, exactly half of Sol. $0.551 input / $3.309 output / $0.055 cached input — 72% below OpenAI's own $2 / $12 list. - [GPT-5.6 Luna](https://apimodels.app/models/gpt-5-6-luna): cheapest nano tier for high-volume work. $0.16 input / $0.96 output / $0.016 cached input — 20% below OpenAI's own $0.20 / $1.20 list. Call it as `gpt-5.6-luna-max`. - [Grok 4.6](https://apimodels.app/models/grok-4.6): xAI's frontier LLM, built for long-running agents. $1.765 input / $5.294 output / $0.441 cached input — 12% below xAI's own $2 / $6. 500K context, text and image input, four reasoning-effort levels (low / medium / high / xhigh), default is high. Two things worth knowing: billed output tokens include reasoning tokens, and the default effort measures ~53s, so set client timeouts above 60s. Live web search works on /api/v1/responses, not on /chat/completions. - [Qwen3.8 Max](https://apimodels.app/models/qwen3.8-max): Alibaba's newest flagship. $2.00 input / $6.00 output / $0.25 cached input per 1M tokens — the same as Alibaba's own list price, so the reason to call it here is the single key and balance, not a discount. OpenAI-format /api/v1/chat/completions, function calling, JSON output and streaming. Reachable directly from mainland China with no Alibaba Cloud account. - [Qwen3.8 Flash](https://apimodels.app/models/qwen3.8-flash): the high-volume tier of the same generation. $0.15 input / $0.47 output / $0.016 cached input per 1M tokens, also at Alibaba's list price. - [Qwen3.7 Max](https://apimodels.app/models/qwen3.7-max): previous flagship, 1M-token context and hybrid thinking (returns reasoning_content). $2.25 input / $6.75 output / $0.45 cached input — 10% below Alibaba's $2.5 / $7.5. Thinking is on by default and its tokens are billed inside completion_tokens at the output rate; send enable_thinking:false for simple work. - [Qwen3.7 Plus](https://apimodels.app/models/qwen3.7-plus): the value tier of that generation — reasoning close to Max at roughly a sixth of the price. $0.36 input / $1.44 output / $0.072 cached input, 10% below Alibaba's list. Above 256K input tokens the whole request bills at triple rate ($1.08 / $4.32), a provider tier we pass through without a markup. Strongest on Chinese and long documents. ## Newest Video Models Video is billed per output second unless stated otherwise. Failed generations are never charged. - [Gemini Omni 1.1 Flash](https://apimodels.app/models/gemini-omni-1.1-flash): Google's next Omni video model, added 2026-09-04. Four modes behind one model id — text-to-video, first/last-frame image-to-video, reference-to-video (up to 7 reference images, reusable voices and characters) and video editing from a source clip. 720p $0.07/s, 1080p $0.10/s, 4K $0.20/s — at least 30% below the official rate; generations that take a video as input are a flat $0.70 (720p/1080p) or $1.05 (4K). 4 / 6 / 8 / 10 seconds, 16:9 or 9:16, synced audio, no visible watermark. - [Seedance 2.5 Global](https://apimodels.app/models/seedance-2.5-global): the fixed-shape, flat-price tier of Seedance 2.5, added 2026-09-16. Every request returns one 30-second 720p clip for a flat $2.00 — the same 30-second 720p clip on seedance-2.5 bills at $0.270/s, about $8.10, so this is roughly a quarter of the price. Text-to-video, or text plus up to 9 reference images, in six aspect ratios (16:9, 9:16, 1:1, 3:4, 4:3, 21:9). The trade-offs are hard limits, not defaults: duration and resolution are not parameters (asking for another length or resolution returns a 400 rather than being rewritten and billed in full), there is no video editing and no video extension (a reference video returns a 400), real people are not supported, audio is always generated with no switch to disable it, and it is not built for high concurrency — send jobs one at a time rather than in a burst. About 6 minutes per clip. Use seedance-2.5 instead when you need a specific length, 480p, editing, extension, real people, or parallel throughput. - [Wan 3.0 Video](https://apimodels.app/models/wan-3.0-video): Alibaba's all-in-one video model, added 2026-08-24. One model name covers text-to-video, first/last-frame, and omni-reference (up to 10 images + 5 videos + 5 audio clips, or a document / web page). Native 30-second single-shot output at 30fps with a synced audio track. Pass `mode:"prime"` for the high-speed tier — measured 58s versus 356s on the same job, for 50% more per second. Current per-second rates are on the model page. - [MiniMax H3 Max Turbo](https://apimodels.app/models/minimax-h3-max-turbo): a post-trained, throughput-optimized build of MiniMax H3, added 2026-09-03 — clips in seconds instead of minutes, with stronger prompt adherence. Any whole length from 5 to 15 seconds, text-to-video in six aspect ratios, or image-to-video from a first frame with an optional last frame. No multi-image references, reference video or audio. Current per-second rates for 480P and 768P are on the model page. - [Grok Imagine Video 1.5](https://apimodels.app/models/grok-imagine-video-1.5): xAI's video model, added 2026-08-22. Text-to-video and image-to-video, billed per output second from $0.0294/s. - [Digital Human](https://apimodels.app/models/digital-human): photo-plus-audio talking-avatar video at $0.01 per output second, added 2026-08-18. ## Newest Image Models - [Grok Imagine Image 2.0](https://apimodels.app/models/grok-imagine-image-2): xAI's newest image model, model id `grok-imagine-image-2.0` on /api/v1/images/generations. Priced by resolution × quality: 1K low $0.03, 2K low $0.04, 1K medium $0.045, 2K medium $0.06 per image — all four 25% below xAI list, and 2K low 33% below. `quality` takes low or medium and defaults to medium. Fuses up to 3 reference images in one call (a 4th is refused explicitly, not dropped silently) and reference images cost nothing extra. 14 aspect ratios including ultra-wide 19.5:9, 9:19.5, 20:9 and 9:20. Strong at in-image typography — an exact headline, date line and small-print URL all came back character-correct on a 2K poster. Two verified limits: editing re-renders the frame rather than editing pixels in place, so crop and camera angle shift between passes; and it cannot produce transparent backgrounds — asked for alpha it paints the checkerboard as ordinary pixels and returns an RGB PNG. - [FLUX.2 Klein 4B](https://apimodels.app/models/flux-2-klein-4b): fast (1-2s) photoreal image generation, added 2026-08-21. Text-to-image and image-to-image with up to 3 reference images, fixed 2MP output, 12 aspect ratios. From $0.006 per image, plus $0.0015 per reference image. - [Qwen Image 3.0](https://apimodels.app/models/qwen3-image): Alibaba's text-to-image and image-editing model in one model id, on /api/v1/images/generations. A flat $0.035 per image at both 1K and 2K, plus $0.004 per reference image. Prompts up to about 4,500 tokens, dense in-image text layout, small text readable down to 10px, 12 languages and 20+ fonts. - [Qwen Image 3.0 Pro](https://apimodels.app/models/qwen3-image-pro): the higher-fidelity tier of the same model — micro-expressions, pores and individual hair strands close to photographic. 1K $0.037, 2K $0.075, plus $0.004 per reference image. Same 4,500-token prompts, 10px small text, 12 languages and 20+ fonts; built for posters, storyboards and menus where the typesetting has to survive. - [GPT-Image-2](https://apimodels.app/models/gpt-image-2): OpenAI's image model without an OpenAI organization verification step. Priced by native resolution: $0.025 at 1K, $0.03 at 2K, $0.05 at 4K per image. [GPT-Image-2 All](https://apimodels.app/models/gpt-image-2-all) puts text-to-image and editing behind one model id. - [GPT-Image-2.5 Flare](https://apimodels.app/models/gpt-image-2.5-flare): OpenAI's GPT Image 2.5, default tier (medium quality by default) — sharper output, region edits that stay local, native transparent PNG. Priced per image by resolution × quality across fifteen cells, $0.008–$0.58: 1K $0.008–$0.18, 2K $0.012–$0.35, 4K $0.02–$0.58 (low → max); medium $0.025/$0.025/$0.045 and high $0.045/$0.06/$0.10 at 1K/2K/4K. [GPT-Image-2.5 Sunburst](https://apimodels.app/models/gpt-image-2.5-sunburst) is the highest-fidelity tier (high quality by default) on the same grid. ## In Chinese / 中文 apimodels.app 提供完整中文站,内容与英文站等价,URL 加 `/zh` 前缀: - [模型目录](https://apimodels.app/zh/models):135 个图片 / 视频 / 音频 / 大语言模型,每个都带价格和一行调用示例。 - [价格](https://apimodels.app/zh/pricing):按次和按 token 的完整价目。 - [文档](https://apimodels.app/zh/docs):API 参考、cURL / Python / Node 示例,以及 Claude Code、OpenAI Codex 等命令行工具的接入方式。 - 特点:一把 API Key 调用所有模型,统一美元余额,按量付费,国内可直连,调用失败不计费,注册赠送 $0.10 试用额度。生成结果文件保留 7 天。 ## Contact - Website: https://apimodels.app - Brand: API Models