DeepSeek bills by the clock. On the official API, V4.1 Flash costs $0.30 input / $1.20 output per 1M tokens during Beijing working hours (Monday to Friday, 09:00–12:00 and 14:00–18:00, which is 01:00–04:00 and 06:00–10:00 UTC) and exactly half, $0.15 / $0.60, at every other hour, all weekend and on Chinese public holidays. Cached input is $0.006 at peak and $0.003 off-peak. Most people comparing prices miss this, because most resellers do not pass the off-peak rate through.
On APIMODELS the same model is 10% below the official price in both windows: $0.27 / $1.08 / $0.0054 cached at peak and $0.135 / $0.54 / $0.0027 off-peak, with the same weekday-only peak schedule and the same holiday exemption. The model id is deepseek-v4.1-flash; the older id deepseek-v4-flash still works and bills the same, just as DeepSeek itself now serves the retired V4 Flash name with V4.1.
Reliability is measured, not claimed: over the last 30 days of external production traffic (internal accounts excluded), deepseek-v4-flash served 20,011 calls at a 99.6% success rate, and deepseek-v4.1-flash 493 calls with none failing. Failed calls are never charged.
Prices per 1M tokens, checked on 2026-09-30 on each provider’s pricing page or public API. "All day" means the provider charges one flat rate regardless of the Beijing clock. Several OpenRouter endpoints advertise lower quantised (fp4) variants; they are left out because they are not the same weights.
| Provider | Input | Cached input | Output | When |
|---|---|---|---|---|
| DeepSeek official | $0.30 | $0.006 | $1.20 | peak (weekday Beijing hours) |
| DeepSeek official | $0.15 | $0.003 | $0.60 | off-peak, weekends, holidays |
| APIMODELS | $0.27 | $0.0054 | $1.08 | peak |
| APIMODELS | $0.135 | $0.0027 | $0.54 | off-peak, weekends, holidays |
| Together AI | $0.30 | $0.006 | $1.20 | all day |
| Fireworks | $0.30 | $0.006 | $1.20 | all day (priority tier $0.375 / $1.50) |
| OpenRouter, most providers | $0.30 | — | $1.20 | all day |
| SiliconFlow (own site) | $0.15 | $0.003 | $0.60 | listed rate |
Take a job of 10 million input tokens, half of them cache hits, and 2 million output tokens. Run it entirely at peak on a flat-rate provider at $0.30 / $0.006 / $1.20 and it costs $3.93. The same job on APIMODELS costs $3.54 at peak and $1.77 off-peak. Scheduling batch work outside Beijing working hours saves more than any choice of provider; the 10% on top applies at every hour.
If you already hold a DeepSeek platform account funded in CNY and only need DeepSeek, the official API is fine; our 10% is the whole difference. The case for a gateway is everything around it: no Chinese phone number or real-name verification, one USD balance paid by Stripe, PayPal or Alipay, and the same key for Claude, GPT, Gemini, Qwen and 130+ other models through one OpenAI-compatible base URL.
cURL
curl https://api.apimodels.app/v1/chat/completions \
-H "Authorization: Bearer $APIMODELS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4.1-flash",
"messages": [{"role": "user", "content": "Summarise this log in three bullet points."}]
}'Python
from openai import OpenAI
client = OpenAI(api_key="YOUR_APIMODELS_KEY", base_url="https://api.apimodels.app/v1")
r = client.chat.completions.create(
model="deepseek-v4.1-flash",
messages=[{"role": "user", "content": "Classify this ticket: refund, bug or question?"}],
)
print(r.choices[0].message.content)
print(r.usage) # cache hits appear here and bill at the cached rateOfficially $0.30 input / $1.20 output per 1M tokens at peak (weekday Beijing working hours) and $0.15 / $0.60 at all other times, with cached input at $0.006 and $0.003. On APIMODELS it is 10% less in both windows: $0.27 / $1.08 at peak and $0.135 / $0.54 off-peak.
Peak is Monday to Friday, 09:00–12:00 and 14:00–18:00 Beijing time (01:00–04:00 and 06:00–10:00 UTC). Everything else is off-peak at half price: nights, the lunch break, all weekend and Chinese public holidays such as the October 1–7 National Day week. APIMODELS follows the same calendar.
DeepSeek retired the V4 Flash name and now serves it with V4.1 Flash at the V4.1 price. On APIMODELS both ids are accepted and bill identically, so existing code keeps working.
Those endpoints serve a quantised (fp4) build or an older snapshot such as V4 Flash 0731, which are different weights from V4.1 Flash. Compare like with like: the full-precision V4.1 Flash is $0.30 / $1.20 at peak almost everywhere.