
gemini-3.8-flashGemini 3.8 Flash is Google's Flash release of 2 September 2026, and on APIMODELS it costs $0.450 per 1M input tokens and $2.250 per 1M output tokens — 40% below Google's own $0.75 / $3.75 list price, and 70% below the $1.50 / $7.50 that list price becomes on 1 January 2027. Google built this one for coding and agents rather than frontier reasoning, and the published numbers follow: top spot on DeepSWE v1.1, 89.4% on Terminal-Bench 2.1, and wins over Claude Opus 5 on HLE-Verified (54.9% vs 54.4%), Vals Finance Agent v2 (61.4% vs 58.6%) and Harvey's Legal Agent (10.0% vs 6.7%). Context is 1M tokens with 64K output; input accepts text, images, audio, video and PDF. Two honest caveats, both measured rather than assumed: it spends noticeably more thinking tokens than earlier Flash models — our own probe burned 83 thinking tokens to answer a two-word prompt, and thinking tokens bill at the output rate — and Artificial Analysis clocks time-to-first-token at 13.3s against a 2.99s median, so it suits batch and agent work far better than interactive chat. For short, high-volume calls where latency and raw price decide, Gemini 3.1 Flash Lite ($0.375 / $2.25) remains the cheaper pick.
$0.450 / $2.250 against $0.75 / $3.75
89.4% on Terminal-Bench 2.1
Text, image, audio, video, PDF in
Thinking tokens bill at the output rate
View complete API reference with all parameters and examples.
View complete API reference with streaming, thinking, and more.
Billing: Cost = (input_tokens * input_price + output_tokens * output_price) / 1,000,000
Gemini 3.8 Flash is a Large Language Model API provided by Google. Gemini 3.8 Flash is Google's Flash release of 2 September 2026, and on APIMODELS it costs $0.450 per 1M input tokens and $2.250 per 1M output tokens — 40% below Google's own $0.75 / $3.75 list price, and 70% below the $1.50 / $7.50 that list price becomes on 1 January 2027. Google built this one for coding and agents rather than frontier reasoning, and the published numbers follow: top spot on DeepSWE v1.1, 89.4% on Terminal-Bench 2.1, and wins over Claude Opus 5 on HLE-Verified (54.9% vs 54.4%), Vals Finance Agent v2 (61.4% vs 58.6%) and Harvey's Legal Agent (10.0% vs 6.7%). Context is 1M tokens with 64K output; input accepts text, images, audio, video and PDF. Two honest caveats, both measured rather than assumed: it spends noticeably more thinking tokens than earlier Flash models — our own probe burned 83 thinking tokens to answer a two-word prompt, and thinking tokens bill at the output rate — and Artificial Analysis clocks time-to-first-token at 13.3s against a 2.99s median, so it suits batch and agent work far better than interactive chat. For short, high-volume calls where latency and raw price decide, Gemini 3.1 Flash Lite ($0.375 / $2.25) remains the cheaper pick. Through APIMODELS platform, you can access this model via a unified API with transparent pay-as-you-go pricing. Current pricing: Input: $0.450, Output: $2.250 per 1M tokens.
Build intelligent conversational systems to automatically answer user queries and improve service efficiency.
Automatically write articles, emails, ad copy, and other text content to boost productivity.
Assist with code writing, debugging, and code review to accelerate software development.
Understand and analyze unstructured data, extract key insights, and generate summary reports.
Gemini 3.8 Flash is available through APIMODELS at: Input: $0.450, Output: $2.250 per 1M tokens. Billing is pay-as-you-go — you only pay for what you generate.
Sign up at APIMODELS, get your API key, and call our unified API endpoint. We provide detailed API documentation with code examples in cURL, Python, and Node.js.
APIMODELS offers the same Gemini 3.8 Flash model through our aggregation platform. We provide a unified API interface so you do not need separate accounts for each provider - one API key to access all models.
Gemini 3.8 Flash is the 2 September 2026 release, and Google tuned this one for coding and agents: top spot on DeepSWE v1.1 and 89.4% on Terminal-Bench 2.1. On APIMODELS it is $0.450 / $2.250, about 32% below 3.6 Flash ($0.662 / $3.302), so it is both the newer and the cheaper of the two. If your workload is short requests at high volume that must answer immediately, Gemini 3.1 Flash Lite ($0.375 / $2.25) is the better fit — cheaper on input and far quicker to first token, where 3.8 Flash measures about 13.3s.
Thinking tokens bill at the output rate, and this generation spends noticeably more of them than earlier Flash models. In our own probe it burned 83 thinking tokens to produce a two-word answer — one visible output token. Artificial Analysis measured 120M output tokens across its benchmark suite against a 71M median. So estimate cost from candidatesTokenCount + thoughtsTokenCount in usageMetadata, not from the text you can see. If the task does not need deep reasoning, cap maxOutputTokens or use 3.1 Flash Lite instead.
Google's own list is $0.75 / $3.75 per 1M tokens; here it is $0.450 / $2.250, which is 40% below. Google's figure is introductory — their announcement puts the standard price at $1.50 / $7.50 from 1 January 2027, at which point the same gap widens to 70%. Failed requests are never billed here, there is no minimum top-up or subscription, and the full 1M context bills at the same input rate with no long-context surcharge.
Text, image, audio, video and PDF in; text out. The context window is 1M tokens with up to 64K output, and the knowledge cutoff is March 2026. Function calling, structured output and streaming all work. Call it either natively — POST /v1beta/models/gemini-3.8-flash:generateContent — or through the OpenAI-compatible /v1/chat/completions. Both endpoints take the same API key.
Use it for multi-turn agent runs, cross-file coding work, long-context document and codebase analysis, and anything you can batch — those are exactly where its scorecard is strong, and a 13-second first token disappears inside a multi-minute task. Avoid it for user-facing real-time chat, autocomplete and voice assistants, and for high-frequency classification of short prompts; for that last case Gemini 3.1 Flash Lite is cheaper per token, much faster to first token, and does not spend extra thinking tokens.
On APIMODELS, Gemini 3.8 Flash runs alongside 60+ models on one API key and one balance, so choosing is about fit, not lock-in. It supports 40% Below Official, Agentic Coding, 1M Context, Multimodal Input, and you can weigh it on price and capability against other Large Language Model models, then switch by changing a single model-name string — no new account or integration. Browse every Large Language Model option with live pricing at apimodels.app/models.
Gemini 3.8 Flash supports: 40% Below Official, Agentic Coding, 1M Context, Multimodal Input. See the APIMODELS docs for full parameters and call examples.
Yes. APIMODELS exposes Gemini 3.8 Flash through a single unified API and one key — no separate provider accounts, and no need to handle each provider's regional network access yourself.
We support Stripe (Visa, Mastercard, and other international cards) and Alipay. Credits are available instantly after payment.