Direct
Direct upstream connection — best when you need native behavior and the full context window.
| Input | Output | Cache read | Cache write |
|---|---|---|---|
| 2.00/M | 10.00/M | 0.10/M | 2.50/M |

claude-sonnet-5-5Output speed is more than 30% faster than Sonnet 5, while using fewer tokens and tool calls to complete the same task.
The same model is available through multiple service channels — choose based on latency, reliability and cost.
Prices in $ / 1M tokensprovider field to the request body, for example "provider": { "channel": "direct" }. Valid values are direct / stable / economical; omit it to use the default channel.Direct upstream connection — best when you need native behavior and the full context window.
| Input | Output | Cache read | Cache write |
|---|---|---|---|
| 2.00/M | 10.00/M | 0.10/M | 2.50/M |
Claude Sonnet 5.5 is the new generation model in Anthropic's Sonnet series, released on September 28, 2026. It launched six days after Opus 5.5 and is the second member of the Claude 5.5 family. Its official positioning is "the best combination of speed and intelligence": unit price is on par with Sonnet 5 and only half of Opus 5.5, output speed is more than 30% faster than Sonnet 5, and because it uses fewer tokens and tool calls to complete the same task, Anthropic estimates per-task cost can drop by up to another 30%.
Its most striking capability is agentic coding: it scores 70.6% on Terminal-Bench 4.0, while Sonnet 5 scores only 10.3%, and it even exceeds Opus 5.5's 66.4% at the xhigh tier. On the knowledge work benchmark GDPval-AA v2.1, it scores 1844, almost tied with Opus 5.5's 1846 and far above Sonnet 5's 1449 and GPT-6 Sol's 1487.
It supports text and image input, a 1 million token context window, 128K maximum output, adaptive thinking enabled by default, and five Effort tiers from low to max to control reasoning depth.
SeaWhale AI provides Claude Sonnet 5.5 through an OpenAI-compatible interface and the Anthropic native Messages API, supporting reasoning Effort control, tool calling, streaming output, and image and PDF input.
Get API key · Model ID:
claude-sonnet-5-5
xhigh), Sonnet 5 10.3%The biggest improvement over Sonnet 5 in this generation is coding. Terminal-Bench 4.0 rose from 10.3% to 70.6%, CursorBench 4.0 from 34.1% to 55.5% (Opus 5.5 57.8%), and FrontierCode 1.1 scores 52.1% at xhigh (Sonnet 5 42.4%, GPT-6 Sol 49.3%, Opus 5.5 54.4%). Anthropic says it surpasses Sonnet 5's best result on Terminal-Bench using medium, while per-task cost is less than one-tenth of the latter.
GDPval-AA v2.1 1844, AA-Briefcase v1.1 1811, both close to Opus 5.5; Humanity's Last Exam with tools 64.5% (Sonnet 5 54.9%, Opus 5.5 67.7%). In Anthropic's internal test, it generated a 10-page performance review slide deck from financial materials and templates, and two professional reviewers thought the first draft could be sent without modification.
Chartography tool-free chart recognition 61.6%, nearly four times Sonnet 5 (15.6%), close to Opus 5.5's 64.4%. OSWorld 2.1 computer use 80.1% (partial score), Sonnet 5 57.0%, Opus 5.5 81.8%. It is also the first Sonnet model to complete Pokémon Red using only screenshots.
Unit price is unchanged; savings come from three places: fewer tokens used, fewer tool calls, and cache reads down to half of Sonnet 5 (0.05x base input price). Minimum cacheable prompt length also dropped from 1,024 tokens to 512 tokens, so short system prompts can also benefit from caching.
| Scenario | Description |
|---|---|
| Production-grade coding agents | Terminal-Bench 4.0 higher than Opus 5.5, half the unit price |
| Enterprise knowledge work | GDPval-AA and AA-Briefcase both close to Opus 5.5; documents, spreadsheets, and slides generated directly |
| Customer service and conversational products | Faster output; low / medium tiers have controllable latency |
| Chart and screenshot understanding | Chartography 61.6%, no cropping tools needed |
| Computer-use agents | OSWorld 2.1 80.1%, within less than 2 percentage points of Opus 5.5 |
| Cost reduction from Opus tier | Similar performance on most coding and knowledge work, price tier 0.5x Opus 5.5 |
| Capability | Claude Sonnet 5.5 | Claude Sonnet 5 | Claude Opus 5.5 |
|---|---|---|---|
| Model ID | claude-sonnet-5-5 |
claude-sonnet-5 |
claude-opus-5-5 |
| Release date | September 28, 2026 | — | September 22, 2026 |
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% (xhigh) |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% |
| GDPval-AA v2.1 | 1844 | 1449 | 1846 |
| AA-Briefcase v1.1 | 1811 | 1359 | 1822 |
| Humanity's Last Exam (with tools) | 64.5% | 54.9% | 67.7% |
| OSWorld 2.1 (partial score) | 80.1% | 57.0% | 81.8% |
| Chartography (no tools) | 61.6% | 15.6% | 64.4% |
| Context window | 1 million tokens | 1 million tokens | 1 million tokens |
| Max output | 128K tokens | 128K tokens | 128K tokens |
| Thinking mode | Adaptive, on by default; minimum is between_tools |
Adaptive, can be disabled | Always on, cannot be disabled |
| Default Effort | high |
high |
medium |
Forced tool call tool_choice: any/tool |
Not supported, returns 400 | Supported | Not supported, returns 400 |
| Minimum cacheable prompt | 512 tokens | 1,024 tokens | — |
| Cache read relative unit price | 0.05x base input price | 0.1x base input price | 0.05x base input price |
| Relative price tier | 0.5x Opus 5.5 | Same as Sonnet 5.5 | 1 |
| Relative latency | Fast | — | Medium |
| Positioning | Best combination of speed and intelligence | Previous-generation Sonnet | Long-horizon coding and knowledge work |
All benchmark data are values published by Anthropic. For specific billing, refer to the live pricing card at the top of the page.
When was Claude Sonnet 5.5 released? September 28, 2026. It launched on Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry the same day, while Sonnet 5 remains available.
What has been upgraded compared with Sonnet 5? Three areas. First, capability: Terminal-Bench 4.0 from 10.3% to 70.6%, GDPval-AA from 1449 to 1844, OSWorld 2.1 from 57.0% to 80.1%, and Chartography from 15.6% to 61.6%. Second, speed and cost: output more than 30% faster, unit price unchanged but token usage lower, and cache reads halved. Third, new features: per-message Effort switching, mid-conversation system messages and tool changes, and on-demand compaction (beta). The tokenizer is the same as Sonnet 5, so token counts remain unchanged after migration.
Do I need to change code to migrate from Sonnet 5?
There are five breaking changes. First, thinking: disabled returns 400; use between_tools instead, and it is only available at low, medium, and high tiers; xhigh and max must use adaptive thinking. Second, tool_choice values any and tool return 400; use auto plus strict: true or structured outputs. Third, thinking blocks are bound to the model and session that produced them: Sonnet 5.5 can read thinking blocks from Sonnet 5, Opus 4.8, and earlier models, but cannot read those from Opus 5, Opus 5.5, Fable, and Mythos; for accounts created after August 31, 2026, replaying thinking blocks after modifying earlier history returns 400, so sessions must remain append-only. Fourth, computer use on Claude API and Google Cloud only accepts computer_toolset_20260801; the older computer_20251124 returns 400. Fifth, the advisor tool no longer accepts Opus 4.8, Opus 4.7, and Sonnet 5 as advisors.
Why is there no output in the frontend between tool calls?
Longer explanatory text between tool calls is now returned as thinking blocks; with the default display: "omitted", the text is empty, no error is raised, but the interface goes "silent." When using adaptive thinking, set thinking.display to retrieve the text, or switch to between_tools, and the text will be returned directly.
How should I choose Effort tiers?
Five tiers: low, medium, high, xhigh, max; the API default is high. The tiers in this generation have been recalibrated, and the thinking amount at the same tier differs from Sonnet 5, so do not directly reuse old settings. Anthropic's recommendation: start with medium for well-defined agentic coding and multi-step tool calling, move to high for harder and longer tasks; use medium or low for latency-sensitive scenarios such as chat.
Are there any downsides or things to note?
Yes. FrontierCode 1.1 at max is only 46.2%, lower than xhigh's 52.1%; Anthropic explains that the highest tier is more prone to out-of-scope modifications and timeouts—higher tier is not always better. OSWorld, Chartography, and Humanity's Last Exam are still slightly below Opus 5.5. On safety filtering, it is the first Sonnet with cybersecurity protection; high-risk cybersecurity requests will be rejected or fall back to Sonnet 5; some microbiology and virology requests may be misjudged; requests attempting to make the model reproduce internal reasoning in the body text will be refused under the reasoning_extraction category. Rejected requests return HTTP 200 and stop_reason: "refusal", which clients need to handle.
How do I choose between it and Opus 5.5? For most coding and knowledge work the two are similar; Sonnet 5.5 is faster and half the price, making it the default choice. Opus 5.5's advantages are in the longest-horizon autonomous tasks, the last few percentage points in image understanding and computer use, and the hardest coding benchmarks such as FrontierCode. When upgrading a session from Sonnet 5.5 to Opus 5.5, thinking blocks can be retained; the reverse will be discarded.
claude-sonnet-5-5https://api.haijingai.com/v2/"provider": { "channel": "direct" }SeaWhale AI is compatible with the OpenAI API protocol, so you can call it with the OpenAI SDK or plain HTTP requests. Streaming is enabled by default.
curl https://api.haijingai.com/v2/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <API_KEY>" \
-d '{
"model": "claude-sonnet-5-5",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
],
"provider": { "channel": "direct" },
"stream": true
}'
# provider is optional — remove this line to use the default channelfrom openai import OpenAI
client = OpenAI(
base_url="https://api.haijingai.com/v2",
api_key="<API_KEY>",
)
stream = client.chat.completions.create(
model="claude-sonnet-5-5",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"},
],
stream=True,
# Optional: pick a service channel; omit to use the default
extra_body={"provider": {"channel": "direct"}},
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)import OpenAI from 'openai'
const client = new OpenAI({
baseURL: 'https://api.haijingai.com/v2',
apiKey: '<API_KEY>',
})
const stream = await client.chat.completions.create({
model: 'claude-sonnet-5-5',
messages: [
{ role: 'system', content: 'You are a helpful assistant.' },
{ role: 'user', content: 'Hello!' },
],
stream: true,
// Optional: pick a service channel; omit to use the default
// @ts-expect-error provider is a SeaWhale AI extension, not in the OpenAI SDK types
provider: { channel: 'direct' },
})
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? '')
}