Skip to main content
Claude API Pricing & Cost per 1M Tokens: Opus, Sonnet, Haiku (2026)
Guides

Claude API pricing: what every Claude model costs per token, and the cheapest way to run it

Claude API pricing per 1M tokens for Opus, Sonnet, Haiku and Fable: input, output and cache rates, real request costs, and the cheapest way to call Claude.

Unifically Model Research Team
7 min readUpdated October 7, 2026

Claude API pricing is billed per token: you pay one rate for the tokens you send and a higher rate for the tokens the model writes back. Unifically serves the same ten Claude models through OpenAI-compatible and Anthropic-compatible endpoints, below Anthropic's list rates, with the Claude Opus 5 API at $1.25 / $6.25 and Claude Haiku 4.5 from $0.25 per million input tokens. This guide covers every model's rate, how token billing works, and what a real request costs.

TL;DR: On Unifically, Claude costs per 1M tokens (input / output) $6.25 / $31.25 (Fable 5), $1.25 / $6.25 (Opus 5), $0.75 / $3.75 (Sonnet 5), and $0.25 / $1.25 (Haiku 4.5), below Anthropic's list price on every model and on the same /v1/chat/completions and /v1/messages request shapes. New accounts get $0.50 of balance on signup, enough for several hundred thousand Claude tokens. See live rates on the pricing page.

Claude API pricing per model (2026)

Anthropic API pricing is published on Anthropic's pricing page.

One recent change worth knowing: Sonnet 5 launched at an introductory price, with a planned increase on September 1, 2026. Anthropic has since made the introductory price the standard one, so the increase will not happen. That leaves Sonnet 5 cheaper than the older Sonnet 4.6 at Anthropic's list rates.

Unifically serves all ten Claude models, below Anthropic's list rate on each. Claude model pricing on Unifically, per million tokens:

Our rates as of ; refreshed from the live price list on every deploy.
ModelOptionOur rate
Claude Fable 5cache read $0.625/1M tokens, input $6.25/1M tokens, cache creation $7.8125/1M tokens, output$31.25/1M tokens
Claude Opus 5cache read $0.125/1M tokens, input $1.25/1M tokens, cache creation $1.5625/1M tokens, output$6.25/1M tokens
Claude Opus 4.8cache read $0.125/1M tokens, input $1.25/1M tokens, cache creation $1.5625/1M tokens, output$6.25/1M tokens
Claude Opus 4.7cache read $0.125/1M tokens, input $1.25/1M tokens, cache creation $1.5625/1M tokens, output$6.25/1M tokens
Claude Opus 4.6cache read $0.125/1M tokens, input $1.25/1M tokens, cache creation $1.5625/1M tokens, output$6.25/1M tokens
Claude Sonnet 5cache read $0.075/1M tokens, input $0.75/1M tokens, cache creation $0.9375/1M tokens, output$3.75/1M tokens
Claude Sonnet 4.6cache read $0.075/1M tokens, input $0.75/1M tokens, cache creation $0.9375/1M tokens, output$3.75/1M tokens
Claude Haiku 4.5cache read $0.025/1M tokens, input $0.25/1M tokens, cache creation $0.3125/1M tokens, output$1.25/1M tokens

Each model has its own page with a playground: Claude Fable 5, Claude Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Opus 4.5, Claude Sonnet 5, Sonnet 4.6, Sonnet 4.5, and Claude Haiku 4.5.

Which model to pick:

  • Claude Sonnet 5 is the default. Near-Opus quality on coding and agentic work, a 1M-token context window, and the best price-to-capability ratio in the family.
  • Claude Opus 5 is good for hard agentic coding, multi-file refactors, and long autonomous runs.
  • Claude Haiku 4.5 is the cheapest Claude API option. Use it for classification, extraction, and high-volume pipelines where latency and cost matter more than depth.
  • Claude Fable 5 is Anthropic's most capable model, priced well above Opus 5. Reserve it for the hardest reasoning and long-horizon agent work.

How Claude token pricing works (input vs output vs cache)

Every Claude request bills three things:

  1. Input tokens. Everything you send: system prompt, conversation history, documents, tool definitions. Billed at the input rate.
  2. Output tokens. Everything the model writes back. On every current Claude model, output costs several times the input rate. Reasoning ("thinking") tokens also bill at the output rate, even when the response hides them.
  3. Cache tokens. Prompt caching stores a stable prefix (system prompt, tool list, long documents) so repeat requests reread it at a fraction of the cost. Cache reads bill at a small fraction of the input rate; writing the cache costs more than plain input for the 5-minute tier.

Two details change the real Claude token cost more than the headline rates:

  • The tokenizer changed. Claude Opus 4.7 and later models (including Sonnet 5 and Fable 5) use a newer tokenizer that produces roughly 30% more tokens for the same text than Sonnet 4.6 and earlier. Compare models by cost per task, not cost per token.
  • Batch requests are half price on Anthropic. Anthropic's Batch API processes async jobs at 50% of standard rates, with most batches finishing within an hour.

Because history is resent on every turn, a long chat's cost is dominated by input tokens. Prompt caching and a frozen system prompt cut most of that on the cached portion.

Claude API cost examples (what a real request costs)

The Claude API cost of a task follows directly from the token counts. Four worked examples at Unifically rates:

TaskTokens (in / out)Input costOutput cost
One chat turn, Sonnet 51,000 / 500$0.00075$0.001875
Summarize a 50-page PDF, Opus 530,000 / 2,000$0.0375$0.0125
Classify 10,000 support tickets, Haiku 4.53M / 200K$0.75$0.25
Chatbot month, Sonnet 520M / 4M$15$15

The pattern: single requests cost fractions of a cent on Sonnet and Haiku, and even Opus-tier work stays under a dollar for most one-off jobs. Costs only get interesting at volume, which is where the per-token rate and prompt caching decide your bill.

How to access the Claude API

Two routes:

  1. Anthropic directly. Create a key in the Anthropic Console and call api.anthropic.com at Anthropic's list prices.
  2. Unifically. One API key covers Claude alongside GPT, image, video, and audio models, billed pay-per-use from a single balance. No subscription, and balance does not expire. Claude models work on three endpoint formats: OpenAI-compatible /v1/chat/completions and /v1/responses, plus the Anthropic-compatible /v1/messages.

A working request against Claude Sonnet 5:

curl -X POST https://api.unifically.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "anthropic/claude-sonnet-5",
    "messages": [
      {"role": "user", "content": "Explain prompt caching in two sentences."}
    ]
  }'

Swap the model id for any row in the table above (anthropic/claude-opus-5, anthropic/claude-haiku-4-5, and so on). Existing OpenAI SDK code works by pointing base_url at https://api.unifically.com/v1, and existing Anthropic SDK code works by swapping the base URL and using /v1/messages.

FAQ

How much does the Claude API cost?

On Unifically, per million tokens (input / output): $1.25 / $6.25 for Opus 5, $0.75 / $3.75 for Sonnet 5, $0.25 / $1.25 for Haiku 4.5, and $6.25 / $31.25 for Fable 5, all below Anthropic's list prices. A typical single request costs well under a cent on Sonnet or Haiku.

Is the Claude API free?

No. Both Anthropic and Unifically bill per token. Claude is free to try on Unifically though: new accounts get $0.50 of balance on signup, which covers several hundred thousand Haiku 4.5 tokens. There is no subscription, and unused balance does not expire.

Which Claude model is the cheapest?

Claude Haiku 4.5: $0.25 / $1.25 per million tokens on Unifically, below Anthropic's list price. It is built for high-volume, low-latency work like classification and extraction.

What does Claude Opus 5 cost per million tokens?

On Unifically, Claude Opus 5 costs $1.25 input / $6.25 output per million tokens, with cache reads at $0.125, below Anthropic's list price.

Why is Anthropic Claude API pricing different from what I pay on Unifically?

Anthropic Claude API pricing is the first-party list rate. Unifically resells the same models at lower per-token rates, below Anthropic's list price on every Claude model. The request and response formats are the same, so switching is a base-URL change.

Do cached tokens cost less?

Yes. Cache reads bill at a small fraction of the input rate on both platforms. On Unifically, Sonnet 5 cache reads are $0.075 per million tokens against $0.75 for fresh input, so a chatbot with a large stable system prompt saves most of its input spend.

Is Claude Fable 5 the same as Claude Mythos 5?

They share the same underlying model. Claude Fable 5 is the generally available version with additional safety measures for dual-use capabilities; Claude Mythos 5 is the same model offered without those measures to approved organizations only. API buyers get Fable 5, and its pricing above is the number that matters; Mythos 5 is not purchasable through standard API access.


Our rates in this post come from Unifically's live price list and refresh on every deploy. Anthropic's own list prices are on Anthropic's pricing page. We will update this post when new Claude models land on the platform. For every rate at once, check the pricing page.

Last updated: October 7, 2026

Continue reading

More Blogs