Every model below is callable through one unified API and a single key. Click a model for its full playground and docs.
| Model | model (API) | Price | Details |
|---|---|---|---|
| Claude Opus 5 | claude-opus-5 | $3 / $15 | Anthropic's newest Opus — a step change on deep reasoning, agentic coding and long-horizon work, at 40% off official pricing. 1M context, adaptive thinking on by default, native tool use. Works with Claude Code, Cursor and the OpenAI SDK alike. |
| Claude Fable 5 | claude-fable-5 | $5 / $25 | Anthropic's latest and most capable public LLM — about 50% cheaper than official. Works with Claude Code & Cursor via the Anthropic API. Ultra-long context, multimodal, complex reasoning, strict safety. |
| Claude Opus 4.8 | claude-opus-4-8 | $3 / $13 | Anthropic's most capable model yet — built to autonomously carry long, complex work end to end. Ideal for big projects, building agents, and high-stakes scenarios demanding top quality and autonomy. |
| Claude Opus 4.7 | claude-opus-4-7 | $3.676 / $18.382 | Latest Opus model with 1M context, 128K max output, and adaptive thinking — same tools and platform features as Opus 4.6. |
| Claude Opus 4.7 (Thinking) | claude-opus-4-7-thinking | $3.676 / $18.382 | Claude Opus 4.7 with extended thinking explicitly enabled for the most complex reasoning tasks. |
| Claude Opus 4.6 | claude-opus-4-6 | $1.765 / $8.824 | Latest Opus model with ultimate performance and reasoning capabilities. |
| Claude Opus 4.6 (Thinking) | claude-opus-4-6-thinking | $1.765 / $8.824 | Claude Opus 4.6 with extended thinking capability for the most complex reasoning tasks. |
| Claude Sonnet 5 | claude-sonnet-5 | $1 / $4.5 | Anthropic's newest Sonnet — 1M-token context (default & max), 128K max output, adaptive thinking, and the same tools & platform features as Sonnet 4.6 (Priority Tier not supported). Works with Claude Code & Cursor via the Anthropic API. |
| Claude Sonnet 4.6 | claude-sonnet-4-6 | $1.059 / $5.295 | Latest Sonnet model with best performance and efficiency. |
| Claude Sonnet 4.6 (Thinking) | claude-sonnet-4-6-thinking | $1.059 / $5.295 | Claude Sonnet 4.6 with extended thinking capability for complex reasoning tasks. |
| Claude Opus 4.5 | claude-opus-4-5-20251101 | $1.765 / $8.824 | Latest Opus model with enhanced capabilities and improved reasoning. |
| Claude Opus 4.5 (Thinking) | claude-opus-4-5-20251101-thinking | $1.765 / $8.824 | Claude Opus 4.5 with extended thinking capability for the most complex reasoning tasks. |
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929 | $1.059 / $5.295 | Latest Sonnet model with improved performance and efficiency. |
| Claude Sonnet 4.5 (Thinking) | claude-sonnet-4-5-20250929-thinking | $1.059 / $5.295 | Claude Sonnet 4.5 with extended thinking capability for complex reasoning tasks. |
| Claude Haiku 4.5 | claude-haiku-4-5-20251001 | $0.353 / $1.765 | Fast and affordable model for lightweight tasks. Best for simple queries and quick responses. |
| Claude Haiku 4.5 (Thinking) | claude-haiku-4-5-20251001-thinking | $0.353 / $1.765 | Claude Haiku 4.5 with extended thinking capability for complex reasoning tasks. |
| Claude Opus 4 | claude-opus-4-20250514 | $5.295 / $26.471 | Most capable model with superior reasoning and analysis capabilities. |
| Claude Opus 4 (Thinking) | claude-opus-4-20250514-thinking | $5.295 / $26.471 | Claude Opus 4 with extended thinking capability for the most complex reasoning tasks. |
| Claude Sonnet 4 | claude-sonnet-4-20250514 | $1.059 / $5.295 | Balanced model with excellent performance and cost efficiency. Great for most tasks. |
| Claude Sonnet 4 (Thinking) | claude-sonnet-4-20250514-thinking | $1.059 / $5.295 | Claude Sonnet 4 with extended thinking capability for complex reasoning tasks. |
| Gemini 3.6 Flash | gemini-3.6-flash | $0.662 / $3.302 | The newest Flash generation — same input price as 3.5 Flash with cheaper output, for agentic and long-horizon work where the answer, not the prompt, dominates the bill. |
| Gemini 3.5 Flash | gemini-3.5-flash | $0.662 / $3.971 | GA release. Our most intelligent Flash model — consistent leadership on agentic execution, coding, and long-horizon tasks at scale. |
| Gemini 3.1 Flash Lite | gemini-3.1-flash-lite | $0.375 / $2.25 | Most cost-effective multimodal model with fastest performance for high-frequency lightweight tasks. |
| Gemini 3.1 Pro Preview | gemini-3.1-pro-preview | $0.824 / $4.941 | Latest Pro model with enhanced reasoning and multimodal capabilities. |
| Gemini 3 Flash Preview | gemini-3-flash-preview | $0.2 / $1.2 | Fast and efficient multimodal model. Great for quick responses and simple tasks. |
| Gemini 3 Pro Preview | gemini-3-pro-preview | $0.442 / $2.648 | Advanced multimodal reasoning model with superior capabilities. |
| Gemini 3 Pro (Thinking) | gemini-3-pro-preview-thinking | $0.442 / $2.648 | Gemini 3 Pro with extended thinking capability for complex reasoning tasks. |
| GPT-5.6 Sol | gpt-5-6-sol | $1.103 / $6.618 | GPT-5.6 Sol is the frontier model in the GPT-5.6 family — OpenAI's highest-intelligence tier (the gpt-5.6 alias routes to Sol), roughly the unsuffixed top tier of earlier GPT-5 families. Reasoning-token support, text + image input, 1.05M-token context, 128K max output, Feb 2026 knowledge cutoff. Call it on /v1/chat/completions or /v1/responses — one key, ~78% below OpenAI list, reachable from China. |
| GPT-5.6 Terra | gpt-5-6-terra | $0.551 / $3.309 | GPT-5.6 Terra balances intelligence and cost — the mini tier of the GPT-5.6 family, priced at exactly half of Sol. Higher reasoning, fast, text + image input, text output, 1.05M-token context, 128K max output, Feb 2026 knowledge cutoff, reasoning-token support. Call it on /v1/chat/completions or /v1/responses — one key, ~78% below OpenAI list, reachable from China. |
| GPT-5.6 Luna | gpt-5-6-luna | $0.221 / $1.324 | GPT-5.6 Luna is optimized for cost-sensitive, high-volume workloads — the nano tier of the GPT-5.6 family, priced at 20% of Sol. High reasoning, fast, text + image input, text output, 1.05M-token context, 128K max output, Feb 2026 knowledge cutoff, reasoning-token support. Served as gpt-5.6-luna-max on /v1/chat/completions or /v1/responses — one key, pay-as-you-go, reachable from China. |
| Grok 4.6 | grok-4.6 | $1.765 / $5.294 | xAI's newest frontier model, built for long-running agents and multi-step work — it matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index. 500K context, text + image input, four reasoning-effort levels, and function calling verified at 12/12 under concurrency with nested schemas. 12% below xAI official ($2 / $6 per 1M). |
| Grok 4.5 | grok-4.5 | $1.5 / $5 | xAI's newest flagship model — leads the industry in coding, non-hallucination rate, agentic tool calling, and instruction following. Supports both non-reasoning and reasoning modes, 500K context. Cheaper than xAI official ($2 / $6 per 1M tokens): 25% off input, ~17% off output. |
| GPT-5.5 | gpt-5-5 | $3 / $18 | OpenAI’s advanced reasoning model for agentic coding, knowledge work, scientific research, and complex multi-step task execution. Served via the Responses API with adjustable reasoning effort (low–xhigh), web search, and function calling. |
| GPT-5.4 | gpt-5-4 | $1.8 / $10.8 | Frontier model for complex professional work and agentic coding. Served via the Responses API with adjustable reasoning effort (low–xhigh), web search, and function calling. |
| GLM-5.2 | glm-5.2 | $0.9 / $3.15 | Zhipu GLM-5.2 — a reasoning model with strong function-calling / tool-use, served via the OpenAI-compatible chat-completions endpoint. |
| DeepSeek V4 Flash | deepseek-v4-flash | $0.12 / $0.24 | DeepSeek V4 Flash — the lightweight, high-throughput, cost-effective member of the DeepSeek V4 family for general chat and basic text. 1M-token context, tool calling, streaming; OpenAI-compatible. |
| DeepSeek V4 Pro | deepseek-v4-pro | $0.37 / $0.74 | DeepSeek V4 Pro — DeepSeek’s high-performance model with top-tier reasoning and agent capabilities, 1M-token context, and full thinking (reasoning_content) output. Tool calling, streaming; OpenAI-compatible. |
| Qwen3.7 Max | qwen3.7-max | $1.76 / $5.29 | Alibaba Qwen3.7 Max — the most capable Qwen3.7 model: top-tier reasoning and agent ability, 1M-token context, hybrid thinking. Tool calling, streaming; OpenAI-compatible. |
| Qwen3.7 Plus | qwen3.7-plus | $0.88 / $1.18 | Alibaba Qwen3.7 Plus — the best-value Qwen3.7 model: strong reasoning at a fraction of Max’s price, 1M-token context, hybrid thinking. Tool calling, streaming; OpenAI-compatible. |
| Text Embedding 3 Small | text-embedding-3-small | $0.018 / 1M tokens | Small embedding model, efficient and cost-effective for most use cases. |
| Text Embedding 3 Large | text-embedding-3-large | $0.059 / 1M tokens | Large embedding model for higher accuracy and flexible dimensions. |
POST /api/v1/messages with a model like claude-opus-4-8 or gemini-3-pro-preview; routing is automatic by model prefix. OpenAI-format models (gpt-, grok-, glm-) also work via /v1/chat/completions.
Yes — LLM calls are billed per token below official list price, with no monthly fee. One key covers Claude, Gemini, GPT, Grok, GLM and more, so you avoid multiple vendor accounts.
Yes. Set stream: true for SSE streaming; tool/function calling passes through to the upstream model. Works with the Anthropic and OpenAI SDKs by changing the base URL.