Direct
Direct upstream connection — best when you need native behavior and the full context window.
| Input context | Input | Output | Cache read | Cache write |
|---|---|---|---|---|
| ≤ 100K | 0.10/M | 0.50/M | 0.01/M | 0.13/M |
| > 100K | 0.50/M | 2.50/M | 0.05/M | 0.63/M |

claude-haiku-5-5The new-generation model in the Haiku series, and the fastest and most economical one in the Claude 5.5 family.
The same model is available through multiple service channels — choose based on latency, reliability and cost.
Prices in $ / 1M tokensprovider field to the request body, for example "provider": { "channel": "direct" }. Valid values are direct / stable / economical; omit it to use the default channel.Direct upstream connection — best when you need native behavior and the full context window.
| Input context | Input | Output | Cache read | Cache write |
|---|---|---|---|---|
| ≤ 100K | 0.10/M | 0.50/M | 0.01/M | 0.13/M |
| > 100K | 0.50/M | 2.50/M | 0.05/M | 0.63/M |
Claude Haiku 5.5 is a new-generation model in Anthropic's Haiku series released on October 7, 2026, and the fastest and most economical model in the Claude 5.5 family, nearly a year after the previous-generation Haiku 4.5 was released. It is officially positioned for high-concurrency, latency-sensitive work: classification, extraction, routing, summarization, real-time customer service, and sub-agents that run subtasks under large-model planning.
This generation is a big leap. The context window expanded from 200K tokens to 1M, max output doubled from 64K to 128K, and it is the first time adaptive thinking and Effort control are supported in the Haiku tier. In evaluations, GDPval-AA v2.1 knowledge work rose from 735 on Haiku 4.5 to 1620, OSWorld 2.1 computer use rose from 15.7% to 72.4%, and Terminal-Bench 4.0 rose from 0.0% to 39.2%, all three higher than GPT-6 Luna.
Pricing was sharply cut in the opposite direction: for requests within 100K tokens, the unit price is 90% lower than Haiku 4.5. Considering that the new tokenizer counts somewhat more tokens, Anthropic's estimate for average running cost is about 75% lower than Haiku 4.5.
SeaWhale AI provides Claude Haiku 5.5 through an OpenAI-compatible interface and the Anthropic native Messages API, supporting reasoning Effort control, tool calling, streaming output, and image and PDF input.
Get API Key · Model ID:
claude-haiku-5-5
low to max, trading off cost and intelligence depending on the taskClassification, extraction, routing, summarization, and context compression are Haiku 5.5's home turf. AA-Briefcase v1.1 scores 1578 (Haiku 4.5 614, GPT-6 Luna 1336), Humanity's Last Exam with tools 57.4% (Haiku 4.5 18.7%). AlphaSense measured 0.84 on 400-document QA vs 0.76 for Haiku 4.5, with about 8 million calls per week for this feature; HubSpot's CRM evaluation averaged 92.8% over three runs, the best result it has seen on a small model.
Latency is the core selling point of this tier. In Box's early tests, it scored 11 points higher than Haiku 4.5 while latency was about half; Asana reports task completion latency down more than 30%, with reasoning up to 2.5x faster per agent turn. Artificial Analysis measured output speed at about 173 tokens/sec at the high level.
low / medium levels for latency-sensitive scenarioshigh and below for minimum time to first tokenAnthropic's recommended usage is to let Opus 5.5 or Sonnet 5.5 handle planning and Haiku 5.5 handle clearly defined subtasks. FrontierCode 1.1 scores 46.4% (GPT-6 Luna 42.4%), Terminal-Bench 4.0 scores 39.2% (GPT-6 Luna 16.4%). In Devin Fusion, Cognition uses Opus 5.5 as the lead and Haiku 5.5 as the sidekick, scoring 66.2 on FrontierCode while reducing cost and latency.
OSWorld 2.1 (offline subset) 72.4%, more than four times Haiku 4.5; Sonnet 5.5 scores 83.9% on the same subset. For repetitive operations such as form filling, data entry, and moving information across applications, running Haiku 5.5 at scale has the lowest cost. Chartography chart recognition without tools 46.4% (Haiku 4.5 6.4%, GPT-6 Luna 29.1%).
| Scenario | Description |
|---|---|
| Batch classification and extraction | High concurrency, low unit price; most cost-effective for requests within 100K tokens |
| Real-time customer service and voice | Currently the fastest Claude model, latency controllable at low levels |
| Sub-agents | Large model plans, Haiku 5.5 executes, lowering total cost of agent systems |
| Summarization and context compression | 1M token window, can process an entire long document |
| Browser and desktop automation | OSWorld 2.1 72.4%, suitable for scaled runs of repetitive operations |
| Upgrading from Haiku 4.5 | Capabilities, context, and output ceiling all improved, while price is lower |
| Capability | Claude Haiku 5.5 | Claude Haiku 4.5 | Claude Sonnet 5.5 |
|---|---|---|---|
| Model ID | claude-haiku-5-5 |
claude-haiku-4-5 |
claude-sonnet-5-5 |
| Release date | October 7, 2026 | October 15, 2025 | September 28, 2026 |
| GDPval-AA v2.1 | 1620 | 735 | 1840 |
| AA-Briefcase v1.1 | 1578 | 614 | 1824 |
| OSWorld 2.1 (offline subset) | 72.4% | 15.7% | 83.9% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 70.6% |
| Humanity's Last Exam (with tools) | 57.4% | 18.7% | 64.5% |
| Chartography (no tools) | 46.4% | 6.4% | 61.6% |
| Context window | 1M tokens | 200K tokens | 1M tokens |
| Max output | 128K tokens | 64K tokens | 128K tokens |
| Thinking mode | Adaptive, enabled by default, can be disabled at high and below |
Manual extended thinking | Adaptive, enabled by default |
| Effort control | Supported, default medium |
Not supported | Supported, default high |
| Assistant message prefill | Not supported, returns 400 | Supported | — |
| Billing tiers | Two tiers by prompt length, split at 100K tokens | Single price | Single price |
| Cache read relative unit price | 0.1x base input price | 0.1x base input price | 0.05x base input price |
| Relative latency | Fastest | — | Fast |
| Positioning | High-concurrency, latency-sensitive tasks | Previous-generation Haiku | Best combination of speed and intelligence |
Evaluation data are all values published by Anthropic; the Sonnet 5.5 column is taken from the Haiku 5.5 release page. Specific billing is subject to the real-time price card at the top of the page.
When was Claude Haiku 5.5 released? October 7, 2026. It was available on Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry on release day. The previous generation Haiku 4.5 was released on October 15, 2025; there was no Haiku 5 in between.
What has been upgraded compared with Haiku 4.5? Four areas. First, capability: GDPval-AA from 735 to 1620, OSWorld 2.1 from 15.7% to 72.4%, Terminal-Bench 4.0 from 0.0% to 39.2%. Second, specs: context from 200K to 1M tokens, max output from 64K to 128K. Third, control: added adaptive thinking and Effort levels, and added browser operation tools. Fourth, cost: requests within 100K tokens are 90% lower in unit price, above 100K are 50% lower, and average running cost is about 75% lower.
Why are there two billing tiers? Haiku 5.5 bills by prompt length: prompts within 100K tokens are billed at the lower tier; requests over 100K tokens have input, output, and cache unit prices all 5 times the lower tier. Anthropic says about 90% of Haiku 4.5 requests fall in the lower tier. Long-context tasks need to factor this into cost, and compress or split first when necessary.
Do I need to change code when migrating from Haiku 4.5?
There are five breaking changes. First, manual extended thinking budget_tokens returns 400; switch to adaptive thinking plus Effort. Second, non-default temperature, top_p, and top_k return 400; just remove them. Third, assistant message prefill returns 400; messages must end with a user message; when a fixed format is needed, switch to structured output. Fourth, for computer use on Claude API and Google Cloud, replace computer_20250124 with computer_toolset_20260801. Fifth, when passing back thinking blocks, the session must remain append-only; modifying earlier history invalidates the thinking blocks. In addition, thinking text is not returned by default; when a summary is needed, set thinking.display to summarized; thinking tokens count toward max_tokens, so setting the limit too small may cause it to stop after outputting only thinking blocks.
Why do token counts increase after migration?
Haiku 5.5 uses the new tokenizer from Claude 4.7 and later models; the same text produces about 30% more tokens than Haiku 4.5, depending on content. The structure of requests and responses is unchanged, but anything estimated by tokens—prompt length, max_tokens, cost budget—must be recalculated.
How do I choose an Effort level?
Five levels: low, medium, high, xhigh, max; API default is medium. For tasks like classification, extraction, and routing, low is enough; for real-time conversation, use low or medium; for sub-agent coding and computer use, you can raise to high or above. The level greatly affects hard tasks: according to VentureBeat, Terminal-Bench 4.0 is about 39% at the highest level, and only about 20% at medium. Thinking can be disabled at high and below with thinking: {"type": "disabled"}, but the official recommendation is to simply lower Effort.
Is there anything lagging or worth noting?
Yes. It is lower than Sonnet 5.5 on all published evaluations; Anthropic explicitly states that complex agentic coding should still use Sonnet 5.5 or Opus 5.5. Requests over 100K tokens have unit prices 5 times the lower tier. Artificial Analysis measured time to first token at about 26 seconds at high, far above the median for its class—for low latency you need to lower Effort or disable thinking. On safety filtering, cybersecurity restrictions are stricter than Haiku 4.5, and penetration testing requests are refused; refused requests return stop_reason: "refusal", and Haiku 5.5 does not support server-side automatic fallback, so clients must handle it themselves.
How do I choose between it and Sonnet 5.5? Look at task difficulty and volume. Use Haiku 5.5 for high-volume, clearly defined, latency-sensitive tasks; use Sonnet 5.5 when you need long-horizon autonomous execution, complex coding, or higher one-shot delivery quality. The most common pairing is both: Sonnet 5.5 or Opus 5.5 for planning and acceptance, Haiku 5.5 for parallel subtask execution.
claude-haiku-5-5https://api.haijingai.com/v2/"provider": { "channel": "direct" }SeaWhale AI is compatible with the OpenAI API protocol, so you can call it with the OpenAI SDK or plain HTTP requests. Streaming is enabled by default.
curl https://api.haijingai.com/v2/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <API_KEY>" \
-d '{
"model": "claude-haiku-5-5",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
],
"provider": { "channel": "direct" },
"stream": true
}'
# provider is optional — remove this line to use the default channelfrom openai import OpenAI
client = OpenAI(
base_url="https://api.haijingai.com/v2",
api_key="<API_KEY>",
)
stream = client.chat.completions.create(
model="claude-haiku-5-5",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"},
],
stream=True,
# Optional: pick a service channel; omit to use the default
extra_body={"provider": {"channel": "direct"}},
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)import OpenAI from 'openai'
const client = new OpenAI({
baseURL: 'https://api.haijingai.com/v2',
apiKey: '<API_KEY>',
})
const stream = await client.chat.completions.create({
model: 'claude-haiku-5-5',
messages: [
{ role: 'system', content: 'You are a helpful assistant.' },
{ role: 'user', content: 'Hello!' },
],
stream: true,
// Optional: pick a service channel; omit to use the default
// @ts-expect-error provider is a SeaWhale AI extension, not in the OpenAI SDK types
provider: { channel: 'direct' },
})
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? '')
}