Skip to content
Sign in

Claude Haiku 5.5

claude-haiku-5-5

The new-generation model in the Haiku series, and the fastest and most economical one in the Claude 5.5 family.

Context window1.0M
ProviderClaude
Released2026/10/08

Playground

Pricing

The same model is available through multiple service channels — choose based on latency, reliability and cost.

Prices in $ / 1M tokens
To pick a channel, add a provider field to the request body, for example "provider": { "channel": "direct" }. Valid values are direct / stable / economical; omit it to use the default channel.

Direct

Direct upstream connection — best when you need native behavior and the full context window.

Input contextInputOutputCache readCache write
≤ 100K0.10/M0.50/M0.01/M0.13/M
> 100K0.50/M2.50/M0.05/M0.63/M

Overview

Input
Text Image
Output
Text

Claude Haiku 5.5 API: The Fastest Claude Model for High Concurrency and Low Latency

Claude Haiku 5.5 is a new-generation model in Anthropic's Haiku series released on October 7, 2026, and the fastest and most economical model in the Claude 5.5 family, nearly a year after the previous-generation Haiku 4.5 was released. It is officially positioned for high-concurrency, latency-sensitive work: classification, extraction, routing, summarization, real-time customer service, and sub-agents that run subtasks under large-model planning.

This generation is a big leap. The context window expanded from 200K tokens to 1M, max output doubled from 64K to 128K, and it is the first time adaptive thinking and Effort control are supported in the Haiku tier. In evaluations, GDPval-AA v2.1 knowledge work rose from 735 on Haiku 4.5 to 1620, OSWorld 2.1 computer use rose from 15.7% to 72.4%, and Terminal-Bench 4.0 rose from 0.0% to 39.2%, all three higher than GPT-6 Luna.

Pricing was sharply cut in the opposite direction: for requests within 100K tokens, the unit price is 90% lower than Haiku 4.5. Considering that the new tokenizer counts somewhat more tokens, Anthropic's estimate for average running cost is about 75% lower than Haiku 4.5.

SeaWhale AI provides Claude Haiku 5.5 through an OpenAI-compatible interface and the Anthropic native Messages API, supporting reasoning Effort control, tool calling, streaming output, and image and PDF input.

Get API Key · Model ID: claude-haiku-5-5


Why Choose Claude Haiku 5.5

  • Anthropic's fastest model — at standard speed, it is the lowest-latency tier in the current Claude lineup
  • Significantly lower cost — requests within 100K tokens have a unit price 90% lower than Haiku 4.5, with average running cost about 75% lower
  • Knowledge work doubled — GDPval-AA v2.1 scores 1620, vs 735 for Haiku 4.5 and 1437 for GPT-6 Luna
  • Computer use that works — OSWorld 2.1 scores 72.4%, vs 15.7% for Haiku 4.5 and 48.9% for GPT-6 Luna
  • First Haiku to support Effort — five levels from low to max, trading off cost and intelligence depending on the task
  • 1M token context + 128K output — the previous generation was 200K and 64K; under Batch API it can expand to 300K output

Core Capabilities

01 High-concurrency text tasks

Classification, extraction, routing, summarization, and context compression are Haiku 5.5's home turf. AA-Briefcase v1.1 scores 1578 (Haiku 4.5 614, GPT-6 Luna 1336), Humanity's Last Exam with tools 57.4% (Haiku 4.5 18.7%). AlphaSense measured 0.84 on 400-document QA vs 0.76 for Haiku 4.5, with about 8 million calls per week for this feature; HubSpot's CRM evaluation averaged 92.8% over three runs, the best result it has seen on a small model.

  • Large-batch classification, tagging, information extraction, and structured output
  • Long-document summarization and context compression within agent systems
  • Database queries, ticket triage, intent routing

02 Real-time conversation and low-latency responses

Latency is the core selling point of this tier. In Box's early tests, it scored 11 points higher than Haiku 4.5 while latency was about half; Asana reports task completion latency down more than 30%, with reasoning up to 2.5x faster per agent turn. Artificial Analysis measured output speed at about 173 tokens/sec at the high level.

  • Online customer service, voice agents, in-app assistants
  • Use low / medium levels for latency-sensitive scenarios
  • Thinking can be disabled at high and below for minimum time to first token

03 Sub-agents and lightweight coding

Anthropic's recommended usage is to let Opus 5.5 or Sonnet 5.5 handle planning and Haiku 5.5 handle clearly defined subtasks. FrontierCode 1.1 scores 46.4% (GPT-6 Luna 42.4%), Terminal-Bench 4.0 scores 39.2% (GPT-6 Luna 16.4%). In Devin Fusion, Cognition uses Opus 5.5 as the lead and Haiku 5.5 as the sidekick, scoring 66.2 on FrontierCode while reducing cost and latency.

  • Focused single-point edits and small multi-file changes
  • Multi-step tool calling and parallel subtasks
  • Complex agentic coding is still recommended for Sonnet 5.5 or Opus 5.5

04 Browser and desktop automation

OSWorld 2.1 (offline subset) 72.4%, more than four times Haiku 4.5; Sonnet 5.5 scores 83.9% on the same subset. For repetitive operations such as form filling, data entry, and moving information across applications, running Haiku 5.5 at scale has the lowest cost. Chartography chart recognition without tools 46.4% (Haiku 4.5 6.4%, GPT-6 Luna 29.1%).

  • Form filling, data entry, cross-application information migration
  • Supports browser operation tools and the new computer use toolset
  • Basic understanding of UI screenshots and charts

Best Use Cases

Scenario Description
Batch classification and extraction High concurrency, low unit price; most cost-effective for requests within 100K tokens
Real-time customer service and voice Currently the fastest Claude model, latency controllable at low levels
Sub-agents Large model plans, Haiku 5.5 executes, lowering total cost of agent systems
Summarization and context compression 1M token window, can process an entire long document
Browser and desktop automation OSWorld 2.1 72.4%, suitable for scaled runs of repetitive operations
Upgrading from Haiku 4.5 Capabilities, context, and output ceiling all improved, while price is lower

Differences Between Claude Haiku 5.5, Haiku 4.5, and Sonnet 5.5

Capability Claude Haiku 5.5 Claude Haiku 4.5 Claude Sonnet 5.5
Model ID claude-haiku-5-5 claude-haiku-4-5 claude-sonnet-5-5
Release date October 7, 2026 October 15, 2025 September 28, 2026
GDPval-AA v2.1 1620 735 1840
AA-Briefcase v1.1 1578 614 1824
OSWorld 2.1 (offline subset) 72.4% 15.7% 83.9%
Terminal-Bench 4.0 39.2% 0.0% 70.6%
Humanity's Last Exam (with tools) 57.4% 18.7% 64.5%
Chartography (no tools) 46.4% 6.4% 61.6%
Context window 1M tokens 200K tokens 1M tokens
Max output 128K tokens 64K tokens 128K tokens
Thinking mode Adaptive, enabled by default, can be disabled at high and below Manual extended thinking Adaptive, enabled by default
Effort control Supported, default medium Not supported Supported, default high
Assistant message prefill Not supported, returns 400 Supported —
Billing tiers Two tiers by prompt length, split at 100K tokens Single price Single price
Cache read relative unit price 0.1x base input price 0.1x base input price 0.05x base input price
Relative latency Fastest — Fast
Positioning High-concurrency, latency-sensitive tasks Previous-generation Haiku Best combination of speed and intelligence

Evaluation data are all values published by Anthropic; the Sonnet 5.5 column is taken from the Haiku 5.5 release page. Specific billing is subject to the real-time price card at the top of the page.


FAQ

When was Claude Haiku 5.5 released? October 7, 2026. It was available on Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry on release day. The previous generation Haiku 4.5 was released on October 15, 2025; there was no Haiku 5 in between.

What has been upgraded compared with Haiku 4.5? Four areas. First, capability: GDPval-AA from 735 to 1620, OSWorld 2.1 from 15.7% to 72.4%, Terminal-Bench 4.0 from 0.0% to 39.2%. Second, specs: context from 200K to 1M tokens, max output from 64K to 128K. Third, control: added adaptive thinking and Effort levels, and added browser operation tools. Fourth, cost: requests within 100K tokens are 90% lower in unit price, above 100K are 50% lower, and average running cost is about 75% lower.

Why are there two billing tiers? Haiku 5.5 bills by prompt length: prompts within 100K tokens are billed at the lower tier; requests over 100K tokens have input, output, and cache unit prices all 5 times the lower tier. Anthropic says about 90% of Haiku 4.5 requests fall in the lower tier. Long-context tasks need to factor this into cost, and compress or split first when necessary.

Do I need to change code when migrating from Haiku 4.5? There are five breaking changes. First, manual extended thinking budget_tokens returns 400; switch to adaptive thinking plus Effort. Second, non-default temperature, top_p, and top_k return 400; just remove them. Third, assistant message prefill returns 400; messages must end with a user message; when a fixed format is needed, switch to structured output. Fourth, for computer use on Claude API and Google Cloud, replace computer_20250124 with computer_toolset_20260801. Fifth, when passing back thinking blocks, the session must remain append-only; modifying earlier history invalidates the thinking blocks. In addition, thinking text is not returned by default; when a summary is needed, set thinking.display to summarized; thinking tokens count toward max_tokens, so setting the limit too small may cause it to stop after outputting only thinking blocks.

Why do token counts increase after migration? Haiku 5.5 uses the new tokenizer from Claude 4.7 and later models; the same text produces about 30% more tokens than Haiku 4.5, depending on content. The structure of requests and responses is unchanged, but anything estimated by tokens—prompt length, max_tokens, cost budget—must be recalculated.

How do I choose an Effort level? Five levels: low, medium, high, xhigh, max; API default is medium. For tasks like classification, extraction, and routing, low is enough; for real-time conversation, use low or medium; for sub-agent coding and computer use, you can raise to high or above. The level greatly affects hard tasks: according to VentureBeat, Terminal-Bench 4.0 is about 39% at the highest level, and only about 20% at medium. Thinking can be disabled at high and below with thinking: {"type": "disabled"}, but the official recommendation is to simply lower Effort.

Is there anything lagging or worth noting? Yes. It is lower than Sonnet 5.5 on all published evaluations; Anthropic explicitly states that complex agentic coding should still use Sonnet 5.5 or Opus 5.5. Requests over 100K tokens have unit prices 5 times the lower tier. Artificial Analysis measured time to first token at about 26 seconds at high, far above the median for its class—for low latency you need to lower Effort or disable thinking. On safety filtering, cybersecurity restrictions are stricter than Haiku 4.5, and penetration testing requests are refused; refused requests return stop_reason: "refusal", and Haiku 5.5 does not support server-side automatic fallback, so clients must handle it themselves.

How do I choose between it and Sonnet 5.5? Look at task difficulty and volume. Use Haiku 5.5 for high-volume, clearly defined, latency-sensitive tasks; use Sonnet 5.5 when you need long-horizon autonomous execution, complex coding, or higher one-shot delivery quality. The most common pairing is both: Sonnet 5.5 or Opus 5.5 for planning and acceptance, Haiku 5.5 for parallel subtask execution.


Why Choose SeaWhale AI to Use the Claude Haiku 5.5 API

  • Full native Messages API passthrough — Effort control, adaptive thinking, prompt caching, and beta features are all directly available
  • Dual-interface access — choose either native Messages API or OpenAI-compatible interface, with low migration cost
  • Direct access in China — no overseas account or self-built proxy needed; stability and latency are guaranteed by the platform side
  • One key for multiple models — use Haiku 5.5 for high traffic, switch to Sonnet 5.5 for daily use, and Opus 5.5 for hard tasks, with unified billing under the same account

API

API integration ​

Model IDUse this value as the model in inference requests
claude-haiku-5-5
API KeyBearer token used to authenticate inference requests
Base URLOpenAI compatible · /chat/completions
OpenAIhttps://api.haijingai.com/v2/
provider OptionalSelects a service channel; omit it and the system picks the default
"provider": { "channel": "direct" }

claude-haiku-5-5 usage examples ​

SeaWhale AI is compatible with the OpenAI API protocol, so you can call it with the OpenAI SDK or plain HTTP requests. Streaming is enabled by default.

js
curl https://api.haijingai.com/v2/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <API_KEY>" \
  -d '{
    "model": "claude-haiku-5-5",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Hello!"}
    ],
    "provider": { "channel": "direct" },
    "stream": true
  }'
# provider is optional — remove this line to use the default channel
js
from openai import OpenAI

client = OpenAI(
    base_url="https://api.haijingai.com/v2",
    api_key="<API_KEY>",
)

stream = client.chat.completions.create(
    model="claude-haiku-5-5",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Hello!"},
    ],
    stream=True,
    # Optional: pick a service channel; omit to use the default
    extra_body={"provider": {"channel": "direct"}},
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
js
import OpenAI from 'openai'

const client = new OpenAI({
  baseURL: 'https://api.haijingai.com/v2',
  apiKey: '<API_KEY>',
})

const stream = await client.chat.completions.create({
  model: 'claude-haiku-5-5',
  messages: [
    { role: 'system', content: 'You are a helpful assistant.' },
    { role: 'user', content: 'Hello!' },
  ],
  stream: true,
  // Optional: pick a service channel; omit to use the default
  // @ts-expect-error provider is a SeaWhale AI extension, not in the OpenAI SDK types
  provider: { channel: 'direct' },
})

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? '')
}
Contact support