Skip to content
Sign in

Claude Sonnet 5.5

claude-sonnet-5-5

Output speed is more than 30% faster than Sonnet 5, while using fewer tokens and tool calls to complete the same task.

Context window1.0M
ProviderClaude
Released2026/10/08

Playground

Pricing

The same model is available through multiple service channels — choose based on latency, reliability and cost.

Prices in $ / 1M tokens
To pick a channel, add a provider field to the request body, for example "provider": { "channel": "direct" }. Valid values are direct / stable / economical; omit it to use the default channel.

Direct

Direct upstream connection — best when you need native behavior and the full context window.

InputOutputCache readCache write
2.00/M10.00/M0.10/M2.50/M

Overview

Input
Text Image
Output
Text

Claude Sonnet 5.5 API: The Speed-and-Intelligence Balance Tier Approaching Opus 5.5

Claude Sonnet 5.5 is the new generation model in Anthropic's Sonnet series, released on September 28, 2026. It launched six days after Opus 5.5 and is the second member of the Claude 5.5 family. Its official positioning is "the best combination of speed and intelligence": unit price is on par with Sonnet 5 and only half of Opus 5.5, output speed is more than 30% faster than Sonnet 5, and because it uses fewer tokens and tool calls to complete the same task, Anthropic estimates per-task cost can drop by up to another 30%.

Its most striking capability is agentic coding: it scores 70.6% on Terminal-Bench 4.0, while Sonnet 5 scores only 10.3%, and it even exceeds Opus 5.5's 66.4% at the xhigh tier. On the knowledge work benchmark GDPval-AA v2.1, it scores 1844, almost tied with Opus 5.5's 1846 and far above Sonnet 5's 1449 and GPT-6 Sol's 1487.

It supports text and image input, a 1 million token context window, 128K maximum output, adaptive thinking enabled by default, and five Effort tiers from low to max to control reasoning depth.

SeaWhale AI provides Claude Sonnet 5.5 through an OpenAI-compatible interface and the Anthropic native Messages API, supporting reasoning Effort control, tool calling, streaming output, and image and PDF input.

Get API key · Model ID: claude-sonnet-5-5


Why choose Claude Sonnet 5.5

  • Agentic coding surpasses Opus 5.5 — Terminal-Bench 4.0 scores 70.6%, Opus 5.5 66.4% (xhigh), Sonnet 5 10.3%
  • Knowledge work close to Opus 5.5 — GDPval-AA v2.1 1844 vs 1846, AA-Briefcase v1.1 1811 vs 1822
  • Fastest Sonnet to date — output speed more than 30% faster than Sonnet 5
  • Same price, more savings — unit price same as Sonnet 5 and half of Opus 5.5, cache reads down to half of Sonnet 5, per-task cost up to 30% lower
  • Much improved image understanding — Chartography chart recognition 61.6%, Sonnet 5 15.6%, GPT-6 Sol 53.6%
  • 1 million token context + 128K output — can extend to 300K output under Batch API

Core capabilities

01 Agentic coding

The biggest improvement over Sonnet 5 in this generation is coding. Terminal-Bench 4.0 rose from 10.3% to 70.6%, CursorBench 4.0 from 34.1% to 55.5% (Opus 5.5 57.8%), and FrontierCode 1.1 scores 52.1% at xhigh (Sonnet 5 42.4%, GPT-6 Sol 49.3%, Opus 5.5 54.4%). Anthropic says it surpasses Sonnet 5's best result on Terminal-Bench using medium, while per-task cost is less than one-tenth of the latter.

  • Multi-file feature development and repository-level refactoring
  • Completes the same tasks with fewer tool calls and shell executions
  • Sustained progress on long tasks lasting hours

02 Knowledge work and document generation

GDPval-AA v2.1 1844, AA-Briefcase v1.1 1811, both close to Opus 5.5; Humanity's Last Exam with tools 64.5% (Sonnet 5 54.9%, Opus 5.5 67.7%). In Anthropic's internal test, it generated a 10-page performance review slide deck from financial materials and templates, and two professional reviewers thought the first draft could be sent without modification.

  • Finance, research, legal document-intensive analysis
  • End-to-end production of reports, spreadsheets, and slides
  • Clearer writing than the previous generation, more natural dialogue

03 Visual understanding and computer use

Chartography tool-free chart recognition 61.6%, nearly four times Sonnet 5 (15.6%), close to Opus 5.5's 64.4%. OSWorld 2.1 computer use 80.1% (partial score), Sonnet 5 57.0%, Opus 5.5 81.8%. It is also the first Sonnet model to complete Pokémon Red using only screenshots.

  • Parsing dense charts, UI screenshots, and tables embedded in PDFs
  • Multi-step autonomous operation of browser and desktop applications
  • Visual preprocessing scaffolding built for older models can be re-evaluated for necessity

04 Faster output and lower per-task cost

Unit price is unchanged; savings come from three places: fewer tokens used, fewer tool calls, and cache reads down to half of Sonnet 5 (0.05x base input price). Minimum cacheable prompt length also dropped from 1,024 tokens to 512 tokens, so short system prompts can also benefit from caching.

  • Output speed more than 30% faster than Sonnet 5
  • The longer and more reused the cache prefix, the greater the cost advantage
  • Supports switching Effort per message mid-conversation (beta); raise the tier for difficult steps and lower it for routine steps

Best use cases

Scenario Description
Production-grade coding agents Terminal-Bench 4.0 higher than Opus 5.5, half the unit price
Enterprise knowledge work GDPval-AA and AA-Briefcase both close to Opus 5.5; documents, spreadsheets, and slides generated directly
Customer service and conversational products Faster output; low / medium tiers have controllable latency
Chart and screenshot understanding Chartography 61.6%, no cropping tools needed
Computer-use agents OSWorld 2.1 80.1%, within less than 2 percentage points of Opus 5.5
Cost reduction from Opus tier Similar performance on most coding and knowledge work, price tier 0.5x Opus 5.5

Differences between Claude Sonnet 5.5, Sonnet 5, and Opus 5.5

Capability Claude Sonnet 5.5 Claude Sonnet 5 Claude Opus 5.5
Model ID claude-sonnet-5-5 claude-sonnet-5 claude-opus-5-5
Release date September 28, 2026 — September 22, 2026
Terminal-Bench 4.0 70.6% 10.3% 66.4% (xhigh)
CursorBench 4.0 55.5% 34.1% 57.8%
GDPval-AA v2.1 1844 1449 1846
AA-Briefcase v1.1 1811 1359 1822
Humanity's Last Exam (with tools) 64.5% 54.9% 67.7%
OSWorld 2.1 (partial score) 80.1% 57.0% 81.8%
Chartography (no tools) 61.6% 15.6% 64.4%
Context window 1 million tokens 1 million tokens 1 million tokens
Max output 128K tokens 128K tokens 128K tokens
Thinking mode Adaptive, on by default; minimum is between_tools Adaptive, can be disabled Always on, cannot be disabled
Default Effort high high medium
Forced tool call tool_choice: any/tool Not supported, returns 400 Supported Not supported, returns 400
Minimum cacheable prompt 512 tokens 1,024 tokens —
Cache read relative unit price 0.05x base input price 0.1x base input price 0.05x base input price
Relative price tier 0.5x Opus 5.5 Same as Sonnet 5.5 1
Relative latency Fast — Medium
Positioning Best combination of speed and intelligence Previous-generation Sonnet Long-horizon coding and knowledge work

All benchmark data are values published by Anthropic. For specific billing, refer to the live pricing card at the top of the page.


FAQ

When was Claude Sonnet 5.5 released? September 28, 2026. It launched on Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry the same day, while Sonnet 5 remains available.

What has been upgraded compared with Sonnet 5? Three areas. First, capability: Terminal-Bench 4.0 from 10.3% to 70.6%, GDPval-AA from 1449 to 1844, OSWorld 2.1 from 57.0% to 80.1%, and Chartography from 15.6% to 61.6%. Second, speed and cost: output more than 30% faster, unit price unchanged but token usage lower, and cache reads halved. Third, new features: per-message Effort switching, mid-conversation system messages and tool changes, and on-demand compaction (beta). The tokenizer is the same as Sonnet 5, so token counts remain unchanged after migration.

Do I need to change code to migrate from Sonnet 5? There are five breaking changes. First, thinking: disabled returns 400; use between_tools instead, and it is only available at low, medium, and high tiers; xhigh and max must use adaptive thinking. Second, tool_choice values any and tool return 400; use auto plus strict: true or structured outputs. Third, thinking blocks are bound to the model and session that produced them: Sonnet 5.5 can read thinking blocks from Sonnet 5, Opus 4.8, and earlier models, but cannot read those from Opus 5, Opus 5.5, Fable, and Mythos; for accounts created after August 31, 2026, replaying thinking blocks after modifying earlier history returns 400, so sessions must remain append-only. Fourth, computer use on Claude API and Google Cloud only accepts computer_toolset_20260801; the older computer_20251124 returns 400. Fifth, the advisor tool no longer accepts Opus 4.8, Opus 4.7, and Sonnet 5 as advisors.

Why is there no output in the frontend between tool calls? Longer explanatory text between tool calls is now returned as thinking blocks; with the default display: "omitted", the text is empty, no error is raised, but the interface goes "silent." When using adaptive thinking, set thinking.display to retrieve the text, or switch to between_tools, and the text will be returned directly.

How should I choose Effort tiers? Five tiers: low, medium, high, xhigh, max; the API default is high. The tiers in this generation have been recalibrated, and the thinking amount at the same tier differs from Sonnet 5, so do not directly reuse old settings. Anthropic's recommendation: start with medium for well-defined agentic coding and multi-step tool calling, move to high for harder and longer tasks; use medium or low for latency-sensitive scenarios such as chat.

Are there any downsides or things to note? Yes. FrontierCode 1.1 at max is only 46.2%, lower than xhigh's 52.1%; Anthropic explains that the highest tier is more prone to out-of-scope modifications and timeouts—higher tier is not always better. OSWorld, Chartography, and Humanity's Last Exam are still slightly below Opus 5.5. On safety filtering, it is the first Sonnet with cybersecurity protection; high-risk cybersecurity requests will be rejected or fall back to Sonnet 5; some microbiology and virology requests may be misjudged; requests attempting to make the model reproduce internal reasoning in the body text will be refused under the reasoning_extraction category. Rejected requests return HTTP 200 and stop_reason: "refusal", which clients need to handle.

How do I choose between it and Opus 5.5? For most coding and knowledge work the two are similar; Sonnet 5.5 is faster and half the price, making it the default choice. Opus 5.5's advantages are in the longest-horizon autonomous tasks, the last few percentage points in image understanding and computer use, and the hardest coding benchmarks such as FrontierCode. When upgrading a session from Sonnet 5.5 to Opus 5.5, thinking blocks can be retained; the reverse will be discarded.


Why choose SeaWhale AI to use the Claude Sonnet 5.5 API

  • Full passthrough of native Messages API — Effort control, per-message Effort switching, thinking blocks, prompt caching, and beta features are all directly available
  • Dual-interface access — choose either native Messages API or OpenAI-compatible interface, with low migration cost
  • Direct connection in China — no overseas account or self-built proxy needed; stability and latency are guaranteed by the platform side
  • One key for multiple models — use Sonnet 5.5 for daily work, switch to Opus 5.5 for difficult tasks, switch to Haiku 5.5 for high traffic, with unified billing under the same account

API

API integration ​

Model IDUse this value as the model in inference requests
claude-sonnet-5-5
API KeyBearer token used to authenticate inference requests
Base URLOpenAI compatible · /chat/completions
OpenAIhttps://api.haijingai.com/v2/
provider OptionalSelects a service channel; omit it and the system picks the default
"provider": { "channel": "direct" }

claude-sonnet-5-5 usage examples ​

SeaWhale AI is compatible with the OpenAI API protocol, so you can call it with the OpenAI SDK or plain HTTP requests. Streaming is enabled by default.

js
curl https://api.haijingai.com/v2/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <API_KEY>" \
  -d '{
    "model": "claude-sonnet-5-5",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Hello!"}
    ],
    "provider": { "channel": "direct" },
    "stream": true
  }'
# provider is optional — remove this line to use the default channel
js
from openai import OpenAI

client = OpenAI(
    base_url="https://api.haijingai.com/v2",
    api_key="<API_KEY>",
)

stream = client.chat.completions.create(
    model="claude-sonnet-5-5",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Hello!"},
    ],
    stream=True,
    # Optional: pick a service channel; omit to use the default
    extra_body={"provider": {"channel": "direct"}},
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
js
import OpenAI from 'openai'

const client = new OpenAI({
  baseURL: 'https://api.haijingai.com/v2',
  apiKey: '<API_KEY>',
})

const stream = await client.chat.completions.create({
  model: 'claude-sonnet-5-5',
  messages: [
    { role: 'system', content: 'You are a helpful assistant.' },
    { role: 'user', content: 'Hello!' },
  ],
  stream: true,
  // Optional: pick a service channel; omit to use the default
  // @ts-expect-error provider is a SeaWhale AI extension, not in the OpenAI SDK types
  provider: { channel: 'direct' },
})

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? '')
}
Contact support