Skip to main content
Anthropic

Claude Haiku 5.5 API

Anthropic

Anthropic's fastest model for classification, extraction, routing and sub-agent work, with a 1M-token context window and adaptive thinking.

anthropic/claude-haiku-5-5

Documentation

Conversation

Anthropic

Start a conversation

Anthropic's fastest model for classification, extraction, routing and sub-agent work, with a 1M-token context window and adaptive thinking.

Enter to send · Shift+Enter for a new line

UnificAlly doesn't store your LLM prompts and replies. We only keep token usage information for billing purposes.

Uses POST /v1/chat/completions with your Unifically API key. Supports system and user prompts, tools, streaming, and thinking when available.

What is Claude Haiku 5.5?

Claude Haiku 5.5 is Anthropic's October 2026 small model for high-volume, latency-sensitive work: classification, extraction, routing and sub-agent steps. Released October 7, 2026, it is the third model of the Claude 5.5 family and the fastest model in Anthropic's current lineup. It takes text and image input, outputs text, and carries a 1M token context window with up to 128k output tokens. On Unifically it runs as anthropic/claude-haiku-5-5. It is also the first Haiku with adaptive thinking and the effort parameter, which is where most of its jump over Claude Haiku 4.5 comes from.

What's new in Claude Haiku 5.5

Computer use: 72.4% on OSWorld 2.1

On the offline subset of OSWorld 2.1, Claude Haiku 5.5 scores 72.4%, up from 15.7% for Claude Haiku 4.5 and ahead of GPT-6 Luna at 48.9%. Claude Sonnet 5.5 sits at 83.9%. The score is partial credit; one independent strict pass-rate run puts it at 37.1%.

Knowledge work close to the mid-tier

On GDPval-AA v2.1, real tasks across 44 occupations scored by Elo, Claude Haiku 5.5 reaches 1620 against 735 for Haiku 4.5, 1437 for GPT-6 Luna and 1840 for Claude Sonnet 5.5. Humanity's Last Exam with tools lands at 57.4% (Haiku 4.5: 18.7%) and Chartography visual reasoning at 46.4% (Haiku 4.5: 6.4%).

43 on the Artificial Analysis Intelligence Index

The independent Artificial Analysis index scores Claude Haiku 5.5 at 43 at max effort as of October 7, 2026, ahead of Gemini 3.8 Flash (41) and GPT-6 Luna (38) and one point behind Kimi K3. The catch is tokens: about 162k output tokens per index task at max effort, roughly three times GPT-6 Luna. At high effort it scores 38 with about 55k.

Agentic coding for scoped tasks

Terminal-Bench 4.0 goes from 0.0% on Haiku 4.5 to 39.2%, and FrontierCode 1.1 Main lands at 46.4% against 42.4% for GPT-6 Luna. Cognition reports Devin Fusion, with Haiku 5.5 as the sidekick model, holding a FrontierCode score of 66.2 while cutting cost and latency.

Best for

Classification and routing

Ticket triage, intent detection and model routing at high request volume.

Extraction

Pulling structured fields out of emails, invoices, filings and web pages.

Sub-agents

Search, read and summarize steps under a Claude Sonnet 5.5 or Opus 5.5 planner.

Computer use

Browser and desktop automation, with 72.4% on OSWorld 2.1.

Chat and support

Low-latency replies where Asana measured over 30% lower task latency.

Charts and screenshots

Reads charts and UI screenshots from image input.

Use cases

Put it at the front of a support or CRM pipeline that classifies, tags and drafts replies: HubSpot reports 92.8% on its simulated-portal CRM suite, the best score it has seen on that suite. Use it as the fast worker in a multi-agent setup, where a larger Claude model plans and Haiku 5.5 runs the many small search, read and extract calls. Build enterprise search and research tools over document stores; Box saw it score 11 points higher than Haiku 4.5 at about half the latency, and AlphaSense measured 0.84 against 0.76 across 400 queries. It also fits browser automation that has to stay cheap per step.

Limitations

Anthropic positions it for narrowly scoped tasks. On Terminal-Bench 4.0 it scores 39.2% against 70.6% for Claude Sonnet 5.5, so complex agentic coding still belongs on Sonnet 5.5 or Opus 5.5. At max effort it spends a lot of output tokens, so the default medium effort is usually the better trade.

Pricing is tiered by prompt length: once a prompt passes 100,000 tokens, the whole request bills at five times the short-prompt rate, so the 1M context is there but long-context work costs much more than short calls. It uses the newer Claude tokenizer, so the same text counts about 30% more tokens than on Haiku 4.5. Artificial Analysis measured factual recall below Gemini 3.8 Flash and GPT-6 Luna (36% accuracy on AA-Omniscience). Cybersecurity safeguards are stricter than Haiku 4.5's: routine defensive work runs, penetration testing does not. Output is text only.

Claude Haiku 5.5 vs GPT-6 Luna

Both are the small, high-volume tier of their families. On Anthropic's published table Haiku 5.5 leads GPT-6 Luna on OSWorld 2.1 (72.4% against 48.9%), Terminal-Bench 4.0 (39.2% against 16.4%), GDPval-AA v2.1 (1620 against 1437) and Chartography (46.4% against 29.1%). Artificial Analysis puts it at 43 against Luna's 38, but at max effort Haiku uses about three times the output tokens, and Luna scores higher on AutomationBench-AA, where Haiku's result was held down by an over-refusal bug Anthropic is fixing. Arena.ai has no Haiku 5.5 votes yet as of October 9, 2026. Our pick: Haiku 5.5 at medium effort for agent steps, computer use and anything with images or charts; test GPT-6 Luna for pure fact lookup and very long prompts.

When to use Claude Haiku 5.5

Use Claude Haiku 5.5 as the upgrade for anything on Claude Haiku 4.5 today, and as the cheap, fast model next to Claude Sonnet 5.5 or Opus 5.5 in a planner-and-worker setup. Keep effort at medium for routine calls, raise it to high for harder extraction or agent runs, and keep prompts under 100k tokens to stay on the lower price tier.

API examples

Call Claude Haiku 5.5 from any language by POSTing to /v1/chat/completions, the OpenAI-compatible endpoint shared by every language model on the platform. Full parameter docs live at docs.unifically.com/models/llm/anthropic/claude-haiku-5-5.

curl -X POST https://api.unifically.com/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer YOUR_API_KEY"   -d '{
    "model": "anthropic/claude-haiku-5-5",
    "messages": [
      { "role": "user", "content": "Classify this support ticket as billing, bug, or feature request and return JSON with a one-line summary: \"I was charged twice for my October top-up.\"" }
    ]
  }'

The response comes back synchronously with the completion. Set "stream": true to receive tokens as they generate.

FAQs

People also ask

anthropic/claude-haiku-5-5, called through the OpenAI-compatible POST /v1/chat/completions endpoint with your Unifically API key.

Anthropic released Claude Haiku 5.5 on October 7, 2026, as the third model of the Claude 5.5 family after Claude Opus 5.5 and Claude Sonnet 5.5. It is live on Unifically since October 9, 2026.

A 1M token context window with up to 128k output tokens per response. It takes text and image input and returns text. Its reliable knowledge cutoff is June 2026.

Haiku 5.5 is priced by prompt length. Requests whose prompt (input plus cached tokens) is over 100,000 tokens bill every token of that request at a higher rate, five times the short-prompt rate. Both tiers are listed on the pricing page. Keep prompts under 100k where you can, or split long documents into chunks.

Yes, by a wide margin on Anthropic's published scores. OSWorld 2.1 goes from 15.7% to 72.4%, Terminal-Bench 4.0 from 0.0% to 39.2%, GDPval-AA v2.1 from 735 to 1620, and Humanity's Last Exam with tools from 18.7% to 57.4%. It also gains a 1M context window, adaptive thinking and the effort parameter.

It uses adaptive thinking, where the model decides how much to think and the effort setting steers it. The default effort on the API is medium. Leave temperature, top_p and top_k unset, since non-default values return a 400 error.

It scores 43 at max effort as of October 7, 2026, ahead of Gemini 3.8 Flash at 41 and GPT-6 Luna at 38, and just behind Kimi K3 at 44. At max effort it used about 162k output tokens per index task; at high effort it scores 38 with about 55k.

Guides and comparisons

Review pricing, limits, and tested outputs before running this model.