
claude-haiku-5-5Claude Haiku 5.5 is the Haiku tier of Anthropic's Claude 5.5 family, released on 7 October 2026, and on APIMODELS it costs $0.09 per 1M input tokens and $0.45 per 1M output — 10% below Anthropic's $0.10 / $0.50 list — with prompt-cache reads at $0.009 and cache writes at $0.1125. Anthropic positions it for high-volume, cost-sensitive work: summaries, context compaction, database queries, classification, sub-agents inside larger agent systems, live customer support and browser use. The Claude Haiku 5.5 benchmarks Anthropic published (max effort, averaged over five trials) are a large step up from Haiku 4.5: GDPval-AA knowledge work 1620 against 735, OSWorld 2.1 computer use 72.4% against 15.7%, Terminal-Bench 4.0 39.2% against 0.0%, Humanity's Last Exam without tools 45.9% against 10.2%, and Chartography 46.4% against 6.4%. Haiku 5.5 vs GPT-6 Luna, OpenAI's high-volume model at the same price here: Haiku 5.5 leads on every benchmark both have a score for — GDPval-AA 1620 vs 1437, AA-Briefcase 1578 vs 1336, OSWorld 2.1 72.4% vs 48.9%, Terminal-Bench 4.0 39.2% vs 16.4%, FrontierCode 1.1 46.4% vs 42.4% and Chartography 46.4% vs 29.1% — though the Luna figures come from Anthropic's comparison and third-party runs rather than from OpenAI. Anthropic is also direct that Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding. Specifications: a 1M-token context window, 128K max output, text and image input, a knowledge cutoff of June 2026 and a minimum cacheable prompt of 512 tokens; the tokenizer produces about 30% more tokens than Haiku 4.5's for the same text. As on Anthropic's own price list, a request whose prompt exceeds 100,000 tokens bills every rate at 5× ($0.45 input, $2.25 output). Adaptive thinking is on by default with effort medium (low to max available); thinking can be disabled only at effort high or below, and budget_tokens is rejected. Call it through /v1/messages or the OpenAI-compatible /v1/chat/completions with one key; failed requests are not charged.
View complete API reference with all parameters and examples.
Enable real-time streaming responses with Server-Sent Events.
{
"model": "claude-haiku-5-5",
"stream": true,
"max_tokens": 1024,
"messages": [...]
}Enable Claude to use tools and call functions.
{
"model": "claude-haiku-5-5",
"max_tokens": 1024,
"tools": [{
"name": "get_weather",
"description": "Get current weather for a location",
"input_schema": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"}
},
"required": ["location"]
}
}],
"tool_choice": {"type": "auto"},
"messages": [{"role": "user", "content": "What's the weather in Tokyo?"}]
}Analyze PDF documents by sending them as base64 encoded content.
{
"model": "claude-haiku-5-5",
"max_tokens": 1024,
"messages": [{
"role": "user",
"content": [{
"type": "document",
"source": {
"type": "base64",
"media_type": "application/pdf",
"data": "<base64_encoded_pdf>"
}
}, {
"type": "text",
"text": "Summarize this document."
}]
}]
}Get structured JSON responses that match your schema.
{
"model": "claude-haiku-5-5",
"max_tokens": 1024,
"output_format": {
"type": "json_schema",
"schema": {
"type": "object",
"properties": {
"name": {"type": "string"},
"age": {"type": "integer"}
},
"required": ["name", "age"]
}
},
"messages": [{"role": "user", "content": "Extract info: John is 30 years old"}]
}Enable Claude to search the web for up-to-date information.
{
"model": "claude-haiku-5-5",
"max_tokens": 1024,
"tools": [{
"type": "web_search_20250305",
"name": "web_search",
"max_uses": 5
}],
"messages": [{"role": "user", "content": "What's the latest news about AI?"}]
}| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | Yes | Model identifier (e.g., claude-haiku-5-5) |
| messages | array | Yes | Array of message objects with role and content |
| max_tokens | integer | Yes | Maximum tokens in the response (1 - 128000) |
| system | string | No | System prompt to set context |
| stream | boolean | No | Enable streaming responses (SSE) |
| temperature | number | No | Sampling temperature (0.0 - 1.0) |
| top_p | number | No | Nucleus sampling threshold (0.0 - 1.0) |
| top_k | integer | No | Top-k sampling (0 - infinity) |
| stop_sequences | array | No | Sequences that stop generation |
| tools | array | No | Function calling tools definition |
| tool_choice | object | No | Tool selection strategy (auto/any/tool) |
| thinking | object | No | Enable extended thinking mode |
| output_format | object | No | Structured output with JSON schema |
View complete API reference with streaming, thinking, and more.
Billing: Cost = (input_tokens * input_price + output_tokens * output_price) / 1,000,000
$0.09 / $0.45 per 1M tokens against Anthropic's $0.10 / $0.50 list; cache reads $0.009, writes $0.1125
Anthropic, max effort: GDPval-AA 1620 (Haiku 4.5: 735), OSWorld 2.1 72.4% (15.7%), Terminal-Bench 4.0 39.2% (0.0%)
Ahead on every shared benchmark: GDPval-AA 1620 vs 1437, OSWorld 2.1 72.4% vs 48.9%, Terminal-Bench 4.0 39.2% vs 16.4%
Anthropic: summaries, compaction, classification, database queries, sub-agents, live support and browser use
Text and image input; knowledge cutoff June 2026; prompts over 100K tokens bill at 5× on every rate
Effort low to max; thinking can be disabled at effort high or below; budget_tokens returns an error
Claude Haiku 5.5 is a Large Language Model API provided by Anthropic. Claude Haiku 5.5 is the Haiku tier of Anthropic's Claude 5.5 family, released on 7 October 2026, and on APIMODELS it costs $0.09 per 1M input tokens and $0.45 per 1M output — 10% below Anthropic's $0.10 / $0.50 list — with prompt-cache reads at $0.009 and cache writes at $0.1125. Anthropic positions it for high-volume, cost-sensitive work: summaries, context compaction, database queries, classification, sub-agents inside larger agent systems, live customer support and browser use. The Claude Haiku 5.5 benchmarks Anthropic published (max effort, averaged over five trials) are a large step up from Haiku 4.5: GDPval-AA knowledge work 1620 against 735, OSWorld 2.1 computer use 72.4% against 15.7%, Terminal-Bench 4.0 39.2% against 0.0%, Humanity's Last Exam without tools 45.9% against 10.2%, and Chartography 46.4% against 6.4%. Haiku 5.5 vs GPT-6 Luna, OpenAI's high-volume model at the same price here: Haiku 5.5 leads on every benchmark both have a score for — GDPval-AA 1620 vs 1437, AA-Briefcase 1578 vs 1336, OSWorld 2.1 72.4% vs 48.9%, Terminal-Bench 4.0 39.2% vs 16.4%, FrontierCode 1.1 46.4% vs 42.4% and Chartography 46.4% vs 29.1% — though the Luna figures come from Anthropic's comparison and third-party runs rather than from OpenAI. Anthropic is also direct that Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding. Specifications: a 1M-token context window, 128K max output, text and image input, a knowledge cutoff of June 2026 and a minimum cacheable prompt of 512 tokens; the tokenizer produces about 30% more tokens than Haiku 4.5's for the same text. As on Anthropic's own price list, a request whose prompt exceeds 100,000 tokens bills every rate at 5× ($0.45 input, $2.25 output). Adaptive thinking is on by default with effort medium (low to max available); thinking can be disabled only at effort high or below, and budget_tokens is rejected. Call it through /v1/messages or the OpenAI-compatible /v1/chat/completions with one key; failed requests are not charged. Through APIMODELS platform, you can access this model via a unified API with transparent pay-as-you-go pricing. Current pricing: Input: $0.09, Output: $0.45 per 1M tokens.
Build intelligent conversational systems to automatically answer user queries and improve service efficiency.
Automatically write articles, emails, ad copy, and other text content to boost productivity.
Assist with code writing, debugging, and code review to accelerate software development.
Understand and analyze unstructured data, extract key insights, and generate summary reports.
Claude Haiku 5.5 is available through APIMODELS at: Input: $0.09, Output: $0.45 per 1M tokens. Billing is pay-as-you-go — you only pay for what you generate.
Sign up at APIMODELS, get your API key, and call our unified API endpoint. We provide detailed API documentation with code examples in cURL, Python, and Node.js.
APIMODELS offers the same Claude Haiku 5.5 model through our aggregation platform. We provide a unified API interface so you do not need separate accounts for each provider - one API key to access all models.
$0.09 per 1M input tokens and $0.45 per 1M output, with cache reads at $0.009 and cache writes at $0.1125 — 10% below Anthropic's $0.10 / $0.50, $0.01 and $0.125. As on Anthropic's own price list, a request whose prompt (new input plus cache reads and writes) exceeds 100,000 tokens bills every rate at 5×: $0.45 input, $2.25 output, $0.045 cache read and $0.5625 cache write, against $0.50 / $2.50 / $0.05 / $0.625 on the list. Billing is per actual token, failed requests are free, and there is no monthly minimum.
Anthropic's launch table, run at max effort and averaged over five trials: GDPval-AA v2.1 knowledge work 1620 (Haiku 4.5: 735), AA-Briefcase v1.1 1578 (614), OSWorld 2.1 offline computer use 72.4% (15.7%), Terminal-Bench 4.0 39.2% (0.0%), Humanity's Last Exam 45.9% without tools and 57.4% with tools (10.2% and 18.7%), FrontierCode 1.1 46.4%, and Chartography without tools 46.4% (6.4%). The system card adds SWE-Bench Pro 64.8. Two caveats: the API default is medium effort, where Anthropic reports GDPval-AA 1277 rather than 1620, and on agentic coding Sonnet 5.5 (Terminal-Bench 4.0 70.6%) is still well ahead.
Haiku 5.5 is ahead on every benchmark where both have a published score: GDPval-AA 1620 vs 1437, AA-Briefcase 1578 vs 1336, OSWorld 2.1 72.4% vs 48.9%, Terminal-Bench 4.0 39.2% vs 16.4%, FrontierCode 1.1 46.4% vs 42.4%, Chartography 46.4% vs 29.1% and PhysicianBench 43.0% vs 33.6%. Read the Luna side with care: none of these Luna numbers come from OpenAI. Artificial Analysis ran GDPval-AA and AA-Briefcase, Anthropic ran OSWorld and PhysicianBench through OpenAI's API, Terminal-Bench is the public leaderboard, Cognition ran FrontierCode and Surge AI published Chartography. OpenAI's own launch post gives Luna one figure, DeepSWE v1.1 66.6%, a test Haiku 5.5 has no score for.
On APIMODELS both cost $0.09 / $0.45 per 1M tokens with cache reads at $0.009, so the choice is about the workload, not the list price. Pick Haiku 5.5 for agent sub-tasks, computer and browser use, and knowledge work, where its benchmark lead is largest. Pick GPT-6 Luna for prompts between 100K and 272K tokens: Haiku 5.5 bills those at 5×, while Luna's long-prompt step starts only above 272K (2× input, 1.5× output). Luna also accepts reasoning effort none for no reasoning at all, and the two tokenizers differ, so the same text is not the same token count — measure on your own prompts. For complex agentic coding, Anthropic itself points to Sonnet 5.5 ($1.20 / $6.00 here).
Both work: the native Anthropic POST https://api.apimodels.app/v1/messages, or the OpenAI-compatible /v1/chat/completions, with model claude-haiku-5-5 and Authorization: Bearer your apimodels key. Streaming, tool calling, image input and prompt caching (cache_control, minimum 512 tokens) are supported; Claude Code and the Anthropic SDK only need the base URL changed to https://api.apimodels.app and your apimodels key. Coming from Haiku 4.5 ($0.353 / $1.765 here), change the model name to claude-haiku-5-5; reasoning_effort on the OpenAI endpoint maps to Anthropic's effort setting.
A 1M-token context window, up to 128K output tokens per call, text and image input, and a knowledge cutoff of June 2026. Anthropic's rules for Haiku 5.5: adaptive thinking is on by default with effort defaulting to medium; thinking can be sent as disabled only at effort high or below, and budget_tokens returns an error; temperature is fixed at 1 and top_p at 0.99, so other values return an error on Anthropic's API — our OpenAI-compatible endpoint leaves them out so standard clients keep working; and the tokenizer yields about 30% more tokens than Haiku 4.5 for the same text.
On APIMODELS, Claude Haiku 5.5 runs alongside 60+ models on one API key and one balance, so choosing is about fit, not lock-in. It supports 10% Below Official, High Volume, 1M Context, Adaptive Thinking, and you can weigh it on price and capability against other Large Language Model models, then switch by changing a single model-name string — no new account or integration. Browse every Large Language Model option with live pricing at apimodels.app/models.
Claude Haiku 5.5 supports: 10% Below Official, High Volume, 1M Context, Adaptive Thinking. See the APIMODELS docs for full parameters and call examples.
Yes. APIMODELS exposes Claude Haiku 5.5 through a single unified API and one key — no separate provider accounts, and no need to handle each provider's regional network access yourself.
We support Stripe (Visa, Mastercard, and other international cards) and Alipay. Credits are available instantly after payment.
How to get access, regional availability, and how this model compares with its alternatives.