Claude 的价目是一条干净的阶梯 —— 顶上是 Fable 5.1,往下依次 Opus 5、Sonnet 5、Haiku 4.5,首尾相差 50 倍。这一页把 Anthropic 官方挂牌价与 APIMODELS 的实收价逐档摆在一起,包括大多数比价页会略过的缓存两列,也直说哪些情况下直接走官方更合适。
开头先更正一件很多比价页至今写错的事:Claude Sonnet 5 的挂牌价是每百万输入/输出 token $2 / $10。流传很广的 $3 / $15 是原定 2026-09-01 生效的涨价,而 Anthropic **明确取消了**它 —— 官方定价页现在写着 $2 / $10「is now the standard price」、那次涨价「will not occur」。凡是拿 $3 / $15 做基准的对比,都会高估自己对 Sonnet 5 声称的折扣 —— 包括我们自己,直到发现为止。
在 APIMODELS 上:Fable 5.1 $5 / $25,恰为 Anthropic $10 / $50 的一半;Opus 5 $3 / $15,对应 $5 / $25 的挂牌价,低 40%;Sonnet 5 $1.60 / $8.00,对应 $2 / $10,低 20%;Haiku 4.5 $0.353 / $1.765,对应 $1 / $5,低约 65%。整条阶梯上的折扣并不统一,我们也不假装统一 —— 它反映的是每一档在上游的真实成本。
提示词缓存是这条阶梯上最有意思、也最容易被通用折扣误导的地方。Anthropic 对多数模型按基础输入价的 0.1 倍收缓存命中,但对 Fable 5.1 按 0.025 倍 —— 于是它的官方缓存读价只有每百万 $0.25。我们是 $0.22,确实更便宜,但只低约 12%,而不是 Fable 5.1 基础价上那 50%。到了 Sonnet 5 则反过来:官方缓存读 $0.20,我们 $0.10,整整低 50%。如果你的负载重度依赖缓存,请专门比这一列,别看通用折扣。
可靠性是实测出来的,不是宣称的:近 30 天外部生产流量(已剔除内部账号)里,Sonnet 5 处理 1,509 次调用,成功率 97.7%、耗时中位数 8.2 秒;Opus 5 处理 500 次,成功率 97.4%、中位数 8.8 秒。Opus 5 的 p95 是 283 秒 —— 长 Agent 回合本来就是这个量级,客户端超时要按此设置。请求失败一律不计费。Fable 5.1 于 2026-09-03 上线,样本还不足以引用,所以我们不引用。
这些模型跑在原生 Anthropic Messages API(/v1/messages)上,不是套一层 OpenAI 形状的转译。也就是说 Claude Code、Anthropic SDK 和 Cursor 只改 base URL 和 key 就能用 —— 工具调用、thinking 块、流式事件都保持原生结构。不需要 Anthropic 账号,中国大陆可直连,平台上所有模型共用一个美元余额。
Price per 1M tokens — Anthropic list vs APIMODELS, verified 2026-09-05
input output cache read vs list
Claude Fable 5.1
Anthropic (list) $10.00 $50.00 $0.25
APIMODELS $5.00 $25.00 $0.22 -50%
Claude Opus 5
Anthropic (list) $5.00 $25.00 $0.50
APIMODELS $3.00 $15.00 $0.391 -40%
Claude Sonnet 5
Anthropic (list) $2.00 $10.00 $0.20
APIMODELS $1.60 $8.00 $0.10 -20%
Claude Haiku 4.5
Anthropic (list) $1.00 $5.00 $0.10
APIMODELS $0.353 $1.765 -- -65%
Note on the cache column: Anthropic reads cache hits at 0.1x base input,
except on Fable 5.1 where it is 0.025x. That is why Fable 5.1's official
cache read ($0.25) is already low and our edge there is ~12%, not 50%.
Sonnet 5 lists at $2/$10. The $3/$15 figure still quoted in many places
was a 2026-09-01 increase that Anthropic cancelled.cURL
# Native Anthropic Messages API — only the base URL and key change
curl https://api.apimodels.app/v1/messages \
-H "Authorization: Bearer $APIMODELS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-5",
"max_tokens": 512,
"messages": [{"role": "user", "content": "Refactor this function for readability."}]
}'Python
import anthropic
client = anthropic.Anthropic(
api_key="YOUR_APIMODELS_KEY",
base_url="https://api.apimodels.app",
)
msg = client.messages.create(
model="claude-fable-5-1", # or claude-opus-5 / claude-sonnet-5
max_tokens=1024,
messages=[{"role": "user", "content": "Plan a migration from Postgres 14 to 17."}],
)
print(msg.content[0].text)
print(msg.usage) # cache_read_input_tokens bills at the cache rate aboveAnthropic 官方挂牌价(每百万输入/输出 token):Fable 5.1 $10 / $50、Opus 5 $5 / $25、Sonnet 5 $2 / $10、Haiku 4.5 $1 / $5。在 APIMODELS 上同样这些模型是 $5 / $25、$3 / $15、$1.60 / $8.00、$0.353 / $1.765,较官方分别低约 50%、40%、20%、65%。
是 $2 / $10。$2 / $10 最初是截至 2026-08-31 的introductory 价,原定 2026-09-01 涨到 $3 / $15 —— 但 Anthropic 取消了,官方定价页现在写着 $2 / $10「is now the standard price」、那次涨价「will not occur」。不少比价内容仍在引用 $3 / $15,那会把该页声称的折扣算高。
可以。这些模型跑在原生 Anthropic Messages API(/v1/messages)上,不是 OpenAI 形状的转译,所以 Claude Code、Anthropic SDK 和 Cursor 只要把 base URL 指向 https://api.apimodels.app 并换成我们的 key 就能用。工具调用、thinking 块和流式事件结构都不变。
缓存命中的计费:Fable 5.1 每百万 $0.22、Opus 5 $0.391、Sonnet 5 $0.10。有一点值得知道:Anthropic 对多数模型按基础输入价的 0.1 倍收缓存命中,但对 Fable 5.1 按 0.025 倍,所以它的官方缓存读本来就只有 $0.25,我们在这一列的优势约 12%,而不是基础价上那 50%。Sonnet 5 则相反:官方缓存读 $0.20、我们 $0.10,整整低一半。缓存 token 数会出现在每次响应的 usage 对象里。
三种情况建议走官方,如实说。其一,Batch API:Anthropic 对异步批处理打五折,我们不转售这一档,大批量离线任务直接走官方可能更便宜。其二,我们没有的账户级产品 —— Opus 5 的 Fast mode、优先档、数据驻留(inference_geo)承诺。其三,DPA、SOC 2 链条这类企业合规文件,只有与 Anthropic、Bedrock 或 Vertex 直签才能拿到。
近 30 天外部流量(已剔除内部账号):Sonnet 5 处理 1,509 次调用,成功率 97.7%(中位 8.2 秒、p95 85.6 秒);Opus 5 处理 500 次,成功率 97.4%(中位 8.8 秒、p95 283.6 秒)。Opus 5 的 p95 是长 Agent 回合的正常量级,不是卡死 —— Agent 类负载请把客户端超时设到 300 秒以上。请求失败一律不计费。Fable 5.1 于 2026-09-03 上线,样本量足够有意义时我们再公布它的数字。