
deepseek-v4.1-flashDeepSeek V4.1 Flash is the V4.1 generation of DeepSeek’s fast, high-throughput model for general chat, summarisation and agent tool use. It has a 1M-token context window, tool calling, JSON output and streaming, and is served via the OpenAI-compatible /v1/chat/completions endpoint. Pricing follows DeepSeek’s peak / off-peak schedule: peak (01:00–04:00 and 06:00–10:00 UTC) $0.405 input / $1.62 output / $0.0081 cached per 1M tokens, off-peak $0.2025 / $0.81 / $0.00405. Prompt-cache hits are billed at the cached rate.
View complete API reference with all parameters and examples.
View complete API reference with streaming, thinking, and more.
Billing: Cost = (input_tokens * input_price + output_tokens * output_price) / 1,000,000
Newer Flash release in the DeepSeek V4 family
Million-token context window for long inputs
Function calling + JSON output, OpenAI-compatible
Off-peak rates are half of peak; cache hits billed at the cached rate
DeepSeek V4.1 Flash is a Large Language Model API provided by DeepSeek. DeepSeek V4.1 Flash is the V4.1 generation of DeepSeek’s fast, high-throughput model for general chat, summarisation and agent tool use. It has a 1M-token context window, tool calling, JSON output and streaming, and is served via the OpenAI-compatible /v1/chat/completions endpoint. Pricing follows DeepSeek’s peak / off-peak schedule: peak (01:00–04:00 and 06:00–10:00 UTC) $0.405 input / $1.62 output / $0.0081 cached per 1M tokens, off-peak $0.2025 / $0.81 / $0.00405. Prompt-cache hits are billed at the cached rate. Through APIMODELS platform, you can access this model via a unified API with transparent pay-as-you-go pricing. Current pricing: Input: $0.405, Output: $1.62 per 1M tokens.
Build intelligent conversational systems to automatically answer user queries and improve service efficiency.
Automatically write articles, emails, ad copy, and other text content to boost productivity.
Assist with code writing, debugging, and code review to accelerate software development.
Understand and analyze unstructured data, extract key insights, and generate summary reports.
DeepSeek V4.1 Flash is available through APIMODELS at: Input: $0.405, Output: $1.62 per 1M tokens. Billing is pay-as-you-go — you only pay for what you generate.
Sign up at APIMODELS, get your API key, and call our unified API endpoint. We provide detailed API documentation with code examples in cURL, Python, and Node.js.
APIMODELS offers the same DeepSeek V4.1 Flash model through our aggregation platform. We provide a unified API interface so you do not need separate accounts for each provider - one API key to access all models.
On APIMODELS, DeepSeek V4.1 Flash runs alongside 60+ models on one API key and one balance, so choosing is about fit, not lock-in. It supports V4.1, Tool Calling, 1M Context, Streaming, OpenAI-compatible, and you can weigh it on price and capability against other Large Language Model models, then switch by changing a single model-name string — no new account or integration. Browse every Large Language Model option with live pricing at apimodels.app/models.
DeepSeek V4.1 Flash supports: V4.1, Tool Calling, 1M Context, Streaming, OpenAI-compatible. See the APIMODELS docs for full parameters and call examples.
Yes. APIMODELS exposes DeepSeek V4.1 Flash through a single unified API and one key — no separate provider accounts, and no need to handle each provider's regional network access yourself.
We support Stripe (Visa, Mastercard, and other international cards) and Alipay. Credits are available instantly after payment.