Skip to main content
Coming soon

Claude Sonnet 5.5 API

Anthropic Sonnet model that lands within two points of Claude Opus 5.5 on knowledge work and coding, with a 1M-token context window and 30%+ faster output than Claude Sonnet 5. We'll open it on Unifically the day the API goes live.

What is Claude Sonnet 5.5?

Claude Sonnet 5.5 is Anthropic's September 2026 Sonnet model for everyday coding, bug fixing, and document, slide, and spreadsheet work. Released September 28, 2026, it is the second model of the Claude 5.5 family and replaces Claude Sonnet 5 in the Sonnet line. It takes text and image input, outputs text, and carries a 1M token context window with up to 128k output tokens. It is not callable on Unifically yet; at launch it will run as anthropic/claude-sonnet-5-5. The headline is how close it gets to Claude Opus 5.5: within two points on knowledge work and coding, while generating output more than 30% faster than Claude Sonnet 5.

What's new in Claude Sonnet 5.5

Agentic coding: 70.6% on Terminal-Bench 4.0

Claude Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, up from 10.3% for Claude Sonnet 5 and above the 66.4% Claude Opus 5.5 reaches at xhigh effort. On CursorBench 4.0, built from real Cursor coding sessions, it scores 55.5% against 34.1% for Sonnet 5 and 57.8% for Opus 5.5.

Knowledge work close to Opus 5.5

On GDPval-AA v2.1, real tasks across 44 occupations scored by Elo, Claude Sonnet 5.5 reaches 1844 against 1846 for Claude Opus 5.5, 1487 for GPT-6 Sol, and 1449 for Claude Sonnet 5. OSWorld 2.1 computer use lands at 80.1% (Sonnet 5: 57.0%) and Humanity's Last Exam with tools at 64.5% (Sonnet 5: 54.9%).

#3 on the Artificial Analysis Intelligence Index

The independent Artificial Analysis index scores Claude Sonnet 5.5 at 56 at max effort, third of 216 models as of September 28, 2026. Only Claude Opus 5.5 at max (58) and xhigh (56) sit at or above it; Claude Fable 5.1 and GPT-6 Astra trail at 53.

Faster, with fewer tokens per task

Output runs more than 30% faster than Claude Sonnet 5, and a typical task costs up to 30% less because it needs fewer tokens and tool calls. Balyasny measured about 121k tokens per answer on 2,441 finance tasks where Sonnet 5 used 497k. Base44 saw app builds finish in 3.6 iterations on average where Claude Opus 5 took 7.7.

Best for

Everyday coding

Bug fixes, scoped features, and code review where speed matters more than deep planning.

Coding sub-agents

Implements plans set by Claude Opus 5.5, with fewer tool calls per task.

Documents and decks

Reports, spreadsheets, and slides that follow a template with little editing.

Computer use

Browser and desktop automation, with 80.1% on OSWorld 2.1.

Support and chat agents

Fast replies and escalation calls on high-volume ticket queues.

Charts and screenshots

Reads charts and UI screenshots from image input, 61.6% on Chartography.

Use cases

Put it behind a coding agent that handles the day-to-day queue: failing tests, small features, and pull request reviews, with Claude Opus 5.5 kept for architecture and the hard cases. Use it for support automation, where Zendesk reports tickets processed 20% faster than with the Claude models it runs in production. Build finance and research tools that read long filings with charts inline and return structured answers at a fraction of the tokens Sonnet 5 needed. It also fits presentation and report generators: in Anthropic's own test, a 10-slide operating review built from earnings materials and a template was judged ready to send on the first draft.

Limitations

Claude Sonnet 5.5 is not live on Unifically yet, so everything here is from Anthropic's release and independent boards, not from our own runs. Opus 5.5 remains clearly stronger at complex, open-ended work that needs sustained judgment, and still leads on Humanity's Last Exam (67.7% against 64.5%) and Chartography (64.4% against 61.6%). FrontierCode also drops at max effort compared with xhigh, so max is not always the best setting for coding.

It is the first Sonnet model with cybersecurity safeguards: higher-risk security tasks visibly fall back to Claude Sonnet 5, while routine bug fixing is unaffected. Classifiers against reasoning extraction are also new, and thinking is tied to the account that created it, so conversations moved between accounts need care. Output is text only.

Claude Sonnet 5.5 vs Claude Opus 5.5

Same 1M token context, same 128k output limit, both with adaptive thinking. Sonnet 5.5 edges Opus 5.5 on Terminal-Bench 4.0 (70.6% against 66.4%) and sits within two points on GDPval-AA v2.1 (1844 against 1846), CursorBench 4.0 (55.5% against 57.8%), and OSWorld 2.1 (80.1% against 81.8%). Opus 5.5 leads on Humanity's Last Exam, AA-Briefcase (1822 against 1811), and FrontierCode, and it scores 58 to Sonnet's 56 on the Artificial Analysis Intelligence Index. Arena.ai has no Sonnet 5.5 votes yet as of September 28, 2026. Our pick: run Sonnet 5.5 at low or medium effort for the bulk of agent and coding traffic, and route open-ended planning and hard debugging to Opus 5.5.

When to use Claude Sonnet 5.5

Use Claude Sonnet 5.5 as the upgrade for anything running on Claude Sonnet 5 today, and as the fast worker next to Claude Opus 5.5 in a planner-and-implementer setup. Set effort to fit the job: low or medium for routine calls, where Anthropic reports it already beats Sonnet 5's best scores at about a tenth of the cost per task, and high or xhigh for longer agent runs. Until it opens on Unifically, Claude Sonnet 5 is the closest callable model.

FAQs

People also ask

Not yet. The page is live ahead of the API, and the model will run as anthropic/claude-sonnet-5-5 on the OpenAI-compatible POST /v1/chat/completions endpoint once it opens. Claude Sonnet 5 and Claude Opus 5.5 are callable today.

Anthropic released Claude Sonnet 5.5 on September 28, 2026, as the second model of the Claude 5.5 family after Claude Opus 5.5. Claude Haiku 5.5 is set to follow in the coming weeks.

A 1M token context window with up to 128k output tokens per response. It takes text and image input and returns text. Its reliable knowledge cutoff is June 2026.

It lands close on most published scores. GDPval-AA v2.1 is 1844 against 1846, CursorBench 4.0 is 55.5% against 57.8%, and OSWorld 2.1 is 80.1% against 81.8%. On Terminal-Bench 4.0 it scores 70.6%, above the 66.4% Opus 5.5 posts at xhigh effort. Opus 5.5 stays stronger on open-ended work that needs sustained judgment.

Yes, by a wide margin. Terminal-Bench 4.0 goes from 10.3% to 70.6%, CursorBench 4.0 from 34.1% to 55.5%, GDPval-AA from 1449 to 1844, and OSWorld 2.1 from 57.0% to 80.1%. Output is 30%+ faster and a typical task uses fewer tokens.

It uses adaptive thinking, where the model decides how much to think and the effort setting steers it from low to max. The default effort on the API is high. If you ran Claude Sonnet 5 with thinking off, switch to the new between_tools setting, which keeps up-front thinking off.

It scores 56 at max effort, third of 216 models as of September 28, 2026, behind two Claude Opus 5.5 settings and ahead of Claude Fable 5.1 and GPT-6 Astra at 53.

Guides and comparisons

Review pricing, limits, and tested outputs before running this model.